A detailed bug report on OpenAI's Codex GitHub repository suggests the company's GPT-5.5 model exhibits unusual reasoning token clustering that may impact performance on complex coding tasks.
The issue, filed by user vguptaa45, analyzed 390,195 response-level token records from February to June 2026. The data shows GPT-5.5 responses disproportionately terminate at exactly 516 reasoning tokens, with additional spikes at 1034 and 1552 tokens.
GPT-5.5 accounts for just 19.3% of all Codex responses but represents 82% of exact-516 token events. The model's exact-516 clustering rate is 44%, compared to 1.3% for other models — a 33.6x difference.
The clustering pattern emerged sharply in recent months. May 2026 saw 53.3% of responses cluster at exactly 516 tokens, up from 0.11% in February. Meanwhile, overall reasoning token intensity declined, with mean tokens dropping from 268 in February to 107 in May.
Performance implications
The bug report links to a previous issue where GPT-5.5 runs ending at exactly 516 reasoning tokens returned incorrect answers on complex tasks. The author suggests this indicates potential "thresholded reasoning-budget behavior" rather than natural variation.
Other OpenAI models show markedly different patterns. GPT-5.2 has a 0.34% exact-516 clustering rate across 247,575 responses, while specialized variants like gpt-5.3-codex and gpt-5.3-codex-spark show no clustering at all.
The issue requests OpenAI investigate whether GPT-5.5 has reasoning budget caps, routing logic, or scheduler behavior causing responses to terminate at fixed token boundaries. The report suggests internal validation checks including quality evaluations comparing exact-516 responses against longer reasoning outputs.
The GitHub issue has attracted significant community attention, with 288 thumbs-up reactions and 60 "eyes" reactions from developers. OpenAI's automated systems tagged it with bug, model-behavior, and rate-limits labels.
The findings raise questions about reasoning consistency across OpenAI's model family and potential trade-offs between computational efficiency and task performance in production AI systems.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.