Emily Forlini reports in Fortune that OpenAI has quietly changed how it evaluates GPT-6 Astra, in ways that appear to flatter the model, and that it has continued revising other metrics since launch.
Changing a measurement after you have seen the result it produces is the oldest problem in empirical work. It is not automatically dishonest. Evaluation suites do get saturated, contaminated by training data, or found to be measuring something other than what they claim. Revising them is sometimes the correct call. What makes it a story is doing it quietly, and doing it in a direction that helps.
The structural issue
Frontier labs grade their own models on tests they design, run internally, and report selectively. There is no external auditor, no requirement to publish the harness, and in most cases no way for an outsider to reproduce the number. Under those conditions, a benchmark score is closer to a marketing claim than a measurement, and it should be read that way.
Third-party evaluations exist and are more useful for exactly this reason. So are held-out sets the lab cannot see, and independent replications on public infrastructure. None of them carry the same weight in a launch announcement, which is the point.
What would fix it
Version the evaluations the way you version software. Publish the harness. When a metric changes, say so at the time, state the reason, and show the old and new numbers side by side for the same model. That costs a lab almost nothing and would make quiet revision impossible to do accidentally.
Until something like that becomes standard practice, the reasonable default is to discount published self-reported gains and wait for external results. Buyers evaluating models for production have mostly figured this out already, which is why internal evals matter less every cycle.