Emily Forlini reports in Fortune that OpenAI has quietly changed how it evaluates GPT-6 Astra, in ways that appear to flatter the model, and that it has continued revising other metrics since launch. Changing a measurement after you have seen the result it produces is the oldest problem in empirical work. It is not automatically dishonest. Evaluation suites do get saturated, … [Read more...] about OpenAI Quietly Revised Its GPT-6 Astra Evaluation Metrics After Launch