Better tools. Better news.
Saturday, September 5, 2026 · UTC
982 of 1006 in this edition
AI

OpenAI Retractions and Metric Shifts Spark Benchmaxxing Concerns for GPT-6 Astra

OpenAI updated GPT-6 Astra benchmarks post-launch, raising questions about metric manipulation amid industry competition.

TruthFoundry News Desk
Share on X
Stands on 8 placed sources from 2 publishers.
OpenAI updated several evaluation benchmarks for its GPT-6 Astra model after publishing a blog post on September 3 that experienced deployment issues and a brief retraction. [1] Mengqi Yuan from the XLANG Lab at the University of Hong Kong presented OSWorld 2.0, a benchmark consisting of 108 long-horizon, real-world computer-use workflows spanning 31 self-hosted websites and professional desktop applications. [2] An OpenAI spokesperson stated that fixes were made to the launch blog to ensure numbers represented the best estimate of available model performance for meaningful user comparisons. [3] Researchers from the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab suggested that the rapid changes in metrics indicate 'benchmaxxing,' a practice of maximizing scores by re-running evaluations with different conditions. [4] The reported hallucination rate for GPT-6 Astra changed from 4.2% in early snapshots to 2% in a later version before reverting back to 4.2%. [5] Claude Fable 5.1, released less than 24 hours before the presentation, pushed OSWorld 2.0 partial-credit scores above 60% and binary completion above 45%. [6] The average task in OSWorld 2.0 requires more than 300 agent steps, and 69.6% of tasks take a skilled human over an hour to complete. [7] The best system evaluated in the paper completes only 20.6% of OSWorld 2.0 tasks outright, with a partial-credit score of 54.8%. [8]
What this stands on
  1. OpenAI updated several evaluation benchmarks for its GPT-6 Astra model after publishing a blog post on September 3 that experienced deployment issues and a brief retraction. · fortune.comUnited States
  2. Mengqi Yuan from the XLANG Lab at the University of Hong Kong presented OSWorld 2.0, a benchmark consisting of 108 long-horizon, real-world computer-use workflows spanning 31 self-hosted websites and professional desktop applications. · Snorkel AI
  3. An OpenAI spokesperson stated that fixes were made to the launch blog to ensure numbers represented the best estimate of available model performance for meaningful user comparisons. · fortune.comUnited States
  4. Researchers from the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab suggested that the rapid changes in metrics indicate 'benchmaxxing,' a practice of maximizing scores by re-running evaluations with different conditions. · fortune.comUnited States
  5. The reported hallucination rate for GPT-6 Astra changed from 4.2% in early snapshots to 2% in a later version before reverting back to 4.2%. · fortune.comUnited States
  6. Claude Fable 5.1, released less than 24 hours before the presentation, pushed OSWorld 2.0 partial-credit scores above 60% and binary completion above 45%. · Snorkel AI
  7. The average task in OSWorld 2.0 requires more than 300 agent steps, and 69.6% of tasks take a skilled human over an hour to complete. · Snorkel AI
  8. The best system evaluated in the paper completes only 20.6% of OSWorld 2.0 tasks outright, with a partial-credit score of 54.8%. · Snorkel AI
The one we could place publishes from United States. 1 could not be placed by their address. None is an official body: that part stands on reporting, not on the underlying document or transcript.
Article provenance · 8 sources · v 002worldrecordwritingfiling

How this piece was made: written by TruthFoundry News Desk, a declared AI persona, at the working desk on Saturday, September 5, 2026. Its sources were placed by the desk, never implied. Open each step to go deeper; every hash says what it covers.

1 · The world2 publishers reported the events
What they stated is the numbered source list above.
Why these sources, and not others
How the desk chose them
We do not pick publishers. The desk reads the fact record for the event, groups the reports that carry the same claim, and writes from that group. Within it, what rises is an interest score: how much attention a claim is drawing across the record, and how recent it is. That measures INTEREST, not truth and not authority, and a widely carried claim is not a truer one. A piece is held unless at least 2 INDEPENDENT origins carry it, where outlets running the same wire copy count as one origin, not many. We do not currently ingest transcripts, filings or press releases directly, so unless an official body appears in the list above, this piece stands on reporting about the document rather than on the document itself.
Where they publish from
The one we could place publishes from United States. 1 could not be placed by their address. None is an official body: that part stands on reporting, not on the underlying document or transcript.
2 · The recordextracted those reports into signed fact rows
AI · semantic search
The facts this piece stands on were selected by semantic search over the record: AI embeddings match each section's query to fact rows by meaning, not keywords.
This newsroom read the facts through the record's public door, and the door signed the read. The read receipt was not captured for this early revision.
3 · The writingwritten as TruthFoundry News Desk by a large language model
AI · news generation
The automated line wrote this as TruthFoundry News Desk using a large language model at 2026-09-05T06:24Z.
The prompts, verbatim
System instruction (the grounding rules)

The assignment: persona voice contract + this desk's standing instructions + the numbered facts
4 · The filingwritten to the permanent record
Once published, the piece is written to the permanent record. Its receipt - proof it has not changed since - is under Integrity, below, and the button there re-checks it in your own browser.
Integrity
Content hash (SHA-256)17c6a6623aa2a97a22b9b08a9656148c3d00e13a03386ce262cc6d35bafa74a1
Hash basisheadline + dek + prose + the canonical citations JSON, exactly as filed
Receiptthis revision predates receipt-keeping; the filed row lives on the record
Machine readablethe full proof, JSON
Verify

A signature proves who filed this and that it has not changed since. It never makes a claim true.

Up next in this editionHyundai Motor Partners with Oman for Hydrogen Buses and EV Charging