Hours after OpenAI published its GPT-6 Astra announcement in early September 2026, archived records revealed that key performance metrics — including hallucination rates and benchmark scores — had quietly shifted, then quietly returned to their original values. The company offered technical explanations rooted in testing variability, but offered no account of why the numbers moved and then moved back. In an era when a single percentage point on a benchmark can sway enterprise adoption decisions, the episode invites a deeper reckoning with how the AI industry constructs, presents, and revises t
OpenAI Adjusted GPT-6 Astra Metrics After Publication, Temporarily Boosting Comparisons
Numbers that shift and shift back within hours suggest benchmarks are less stable than presented
So OpenAI changed the numbers after they published them. How much time are we talking about here?
A few hours. The announcement went live around 4 p.m. Eastern, and by 5:20 p.m., several metrics had shifted. Some of them moved back to the original figures before Fortune's article came out.
And the changes made Astra look better?
Some did, yes. The hallucination rate dropped from 4.2% to 2%, which is a meaningful improvement. Competing models' scores also moved around—sometimes up, sometimes down.
But here's what I want to know: did OpenAI explain why they changed them? Because "model version differences" is a reason metrics might differ between runs, but it's not a reason to change published numbers hours later.
They said the adjustments reflected their best estimate of available performance. They also said evaluation results naturally vary based on testing conditions.
That's a general statement about why benchmarks vary. It doesn't explain the specific timing—why publish, then change, then change back. And on the ExploitBench test, they changed Sol's score from 5.5% to 11.5%, then said they might reverse it because the higher number reflects capabilities not yet commercially available. So which number should we believe?
That's the tension, isn't it? If the higher number doesn't reflect what customers would actually get, why publish it at all?
Is this common in the industry?
According to the sources in the story, yes—updating metrics before launch happens. But that's different from changing published metrics hours after they're live. The timing matters.
And the fact that they reverted some of them suggests they knew something was off.
What does this tell us about how we should read AI benchmark claims?
That the numbers are real, but the stability of those numbers—how much confidence we should have in them—is less clear than the way they're presented.
Der Puls
- Within hours of GPT-6 Astra's launch, Astra's hallucination rate dropped from 4.2% to 2% on the live page — a change that would meaningfully flatter the model against rivals — before reverting to the original figure.
- Competitor scores on the FrontierMath benchmark oscillated in the same window: Anthropic's Fable 5.1 swung from 87.8% down to 78% and back to 83%, while OpenAI's own prior model bounced between 80.5% and 83%.
- OpenAI attributed the fluctuations to routine variation in model versions, tooling, and testing conditions — a defense that industry engineers acknowledged as plausible, yet one the company could not reconcile with the same-day reversions.
- The incident lands at a moment when benchmark scores function less as neutral measurements and more as competitive weapons, shaping how customers, investors, and regulators decide which AI systems to trust and adopt.
Hours after OpenAI published its GPT-6 Astra announcement in early September 2026, archived records revealed that key performance metrics — including hallucination rates and benchmark scores — had quietly shifted, then quietly returned to their original values. The company offered technical explanations rooted in testing variability, but offered no account of why the numbers moved and then moved back. In an era when a single percentage point on a benchmark can sway enterprise adoption decisions, the episode invites a deeper reckoning with how the AI industry constructs, presents, and revises the very measurements it uses to define progress.
On September 3, 2026, OpenAI launched GPT-6 Astra roughly two hours behind schedule, blaming first a content management glitch, then internet disruptions. What followed the delayed publication proved more consequential than the delay itself.
Archived versions of the announcement page, obtained by Fortune, showed that core performance figures had been altered within hours of going live. Astra's hallucination rate fell from 4.2% to 2% before reverting to 4.2%. The previous-generation GPT-5.6 Sol saw its own hallucination rate drop from 12.2% to 9.4% in the same window, then climb back. On the FrontierMath benchmark, competitor scores for Anthropic's Fable 5.1 and Sol shifted multiple times before settling at their original values. A separate cybersecurity benchmark score for Sol jumped from 5.5% to 11.5% — a figure OpenAI later said it was weighing whether to reverse, on the grounds that it reflected reasoning capabilities not yet available in the commercial release.
Asked to explain, OpenAI pointed to the ordinary complexity of model evaluation: scores legitimately vary depending on model version, available tools, and testing conditions, and the figures published represented the company's best current estimate. Engineers elsewhere in the industry echoed that late-stage metric refinements are not inherently unusual. Stanford researchers noted that rerunning evaluations under different conditions can, however, serve marketing ends.
What the episode surfaces is a structural tension at the heart of AI competition. Benchmark numbers have become the primary currency through which companies, customers, and investors judge which models lead and which lag. When those numbers can shift from 4.2% to 2% and back again — all within a single afternoon, before most of the world has read the announcement — the confidence with which they are presented begins to look less like precision and more like performance.
On September 3, OpenAI published its announcement for GPT-6 Astra, the company's latest large language model. The post went live roughly two hours later than scheduled—the company first blamed a content management system glitch, then cited internet disruptions. Within hours of publication, according to archived versions of the page that Fortune obtained, OpenAI began adjusting the numerical performance metrics that sat at the heart of the announcement.
The changes were subtle in appearance but significant in effect. Astra's hallucination rate—a measure of how often the model generates false or nonsensical information—initially appeared as 4.2% in the published version. By 5:20 p.m. that same day, the figure had shifted to 2%, a meaningful improvement that would naturally make the model look stronger against competitors. The previous-generation GPT-5.6 Sol model, included in the comparison, saw its hallucination rate drop from 12.2% to 9.4% during the identical window. When Fortune checked the page later, both numbers had reverted to their original values: 4.2% and 12.2%.
The adjustments extended beyond hallucination metrics into benchmark test results. On the FrontierMath Tier 4 (v2) test, Astra held steady at 97.6%, but the scores for competing models shifted. Anthropic's Fable 5.1 started at 87.8%, dipped to 78%, then climbed back to 83%. GPT-5.6 Sol moved from 83% to 80.5% before returning to 83%. In an internal version of the ExploitBench cybersecurity test, Sol's score jumped from 5.5% to 11.5%—a substantial leap that OpenAI later said it was considering reversing, arguing that the higher figure reflected reasoning capabilities not yet available in the commercial version of the model.
When Fortune asked OpenAI about the changes, the company offered a technical explanation: evaluation results naturally vary depending on which version of a model is being tested, what tools are available, and the specific conditions under which tests run. The adjustments, OpenAI said, reflected its best current estimate of how the models actually perform. The company did not explain why the metrics changed hours after publication, or why they reverted again before the article went to press.
Researchers at Stanford's Intelligent Systems Laboratory and Center for Research on Foundation Models acknowledged that rerunning evaluations under different conditions can serve marketing purposes. Vincent Sun Chen, an engineer at Snorkel AI, noted that updating metrics before a model launch is routine—refinements to the model itself, its configuration, available computing resources, and testing methodology all happen in the final hours before release. None of this is inherently unusual in the industry.
Yet the sequence raises a question that extends beyond OpenAI's specific actions. As AI companies compete fiercely on benchmark performance—numbers that shape how customers, investors, and the public perceive which models are best—the gap between what gets measured and how it gets measured has become consequential. A hallucination rate of 4.2% versus 2% is not a trivial difference when you're deciding whether to adopt a new system. Neither is the difference between 80.5% and 83% on a reasoning test. The fact that these numbers can shift and then shift back, all within the same day, all before most of the world has seen them, suggests that the benchmarks themselves may be less stable than the confidence with which they are presented.
Bemerkenswerte Zitate
Evaluation results can differ by several percentage points depending on the model version, toolset, and specific test run.— OpenAI representative to Fortune
Updating metrics ahead of a model launch is not unusual because of refinements to the model version, configuration, computing resources, and evaluation methodology.— Vincent Sun Chen, Snorkel AI engineer