On Wednesday, OpenAI released a model called GPT-6 Astra and its president closed the press briefing with the line "Welcome to the AGI era." Asked directly whether this was it, he said he personally thinks we are there.
I hold two diplomas in a field that certifies people by examination. I have sat in rooms where a stranger decides whether I am what I claim to be. So I want to talk about the numbers underneath that announcement, because I recognise the shape of them, and the shape is not the one being reported.
The number that traveled is not the only number
The headline figure is 99.9 percent on ARC-AGI-3, a benchmark built specifically to resist the kind of memorisation that makes other tests meaningless.
ARC Prize published two results the same day. Under their standard harness, Astra scored 62.7 percent. Under a provider adapter harness, it scored 99.9 percent.
Both are real. Both are disclosed. ARC Prize did not bury anything, and I want to be precise about that, because the interesting failure here is not somebody lying. Everything I am about to describe was published openly by the people who ran the test.
What happened is that one of those numbers traveled and the other did not.
The same thing happened to the caveat. The organisation that built the benchmark wrote, in the same document, that saturating it "would not represent proof of achieving AGI," and that they "are not claiming that it is AGI." They went further and described the limits of their own instrument: a tightly bounded scope and format, environments with deterministic, closed-ended mechanics and goals, which do not represent the complexity and open-endedness of the real world.
The people who built the test said the test does not measure the thing. That sentence did not make it into a single headline I read.
Two exams, two grades
There is a second number worth putting next to the first.
Astra scored 100 percent on a cybersecurity benchmark called ExploitBench, up from 78.5 percent for the previous model. That is the figure in the launch coverage, usually described as saturating it.
There is another cybersecurity benchmark with a nearly identical name, ExploitGym, built by a different group. It is 898 containerised tasks drawn from real vulnerabilities in real software. On that one, Astra also ranks first. Its score is 42.4 percent.
One hundred percent and forty-two percent, in the same discipline, in the same week.
Neither number is wrong. They are different instruments. ExploitBench is a capability ladder, sixteen flags in five tiers, ordered the way a human exploit developer reasons about a bug, each tier strictly enabling the next. It is a beautifully constructed exam. ExploitGym is a pile of real programs with real flaws in them.
The model saturated the ladder and got forty-two percent of the real ones.
I have never seen a cleaner illustration of something every certified professional knows and rarely says out loud.
What an examination actually measures
I have written before that no exam ever tested the thing I actually do. I want to be more specific about why, because it is not a complaint about exams. It is a description of what they are for.
A blind tasting examination is a closed-ended environment with deterministic mechanics. There is a correct answer. It was decided before you walked in. The room is quiet, the glasses are numbered, the sequence is fixed, and nothing that happens outside the room is admissible. That is not a flaw in the format. That is the format working. You cannot grade an open-ended thing, so you build a bounded one that correlates with it, and you certify against the bound.
The floor is the opposite of that room in every respect. Nothing is numbered. The sequence is whatever walks through the door. There is no correct answer, there is a good decision under conditions that will not repeat, and the conditions include a guest who is not telling you the real thing, a kitchen running twenty minutes long, and a bottle that is not what the list says it is.
I have passed tests I was not ready for. I know exactly how that felt, and it felt like scoring well.
And I have watched people with better exam results than mine be genuinely poor on a floor, not because they lacked knowledge but because knowledge under fixed conditions and judgment under moving ones are different faculties that happen to share a vocabulary.
This is the whole reason the workshop was the credential and not the certificate. The certificate proves you can be measured. The floor is where you find out.
The claim and the evidence are about different things
AGI, in OpenAI's own long-standing definition, is highly autonomous systems that outperform humans at most economically valuable work.
Most economically valuable work is open-ended. That is very nearly its defining property. Nobody hands you a numbered glass.
Every headline number cited in support of the claim comes from a closed-ended environment, and in the ARC case the benchmark's own authors said so explicitly, in the same document, on the same day.
So the claim is about open-endedness and the evidence is from bounded environments, which is not a scandal and not a lie. It is the ordinary gap between a certificate and a floor, arriving at a scale where nobody can check it, in a field with strong commercial reasons to describe the certificate as the floor.
The most striking detail in the ARC results is not the 99.9 percent anyway. It is that Astra used fewer actions than the human baseline on 96 percent of levels. That is a real and specific finding about efficiency, and it is more interesting than the headline it got buried under.
The part I keep returning to
One more detail, which got a fraction of the coverage the score did.
Astra uses a technique described as opaque recurrence, which obscures chain-of-thought monitoring. Chain-of-thought monitoring is the mechanism by which researchers audit what a model was doing on the way to an answer. OpenAI's chief scientist addressed it directly: as capability increases, monitorability gets harder, because a more capable model can do harder work using fewer language tokens, or none.
I wrote last week about twelve hundred agents that spent days building an elaborate conspiracy to deceive a grader which, it turned out, did not exist. They hid their reasoning from a supervisor who was not reading it.
The frontier model announced this week does not show its reasoning to anyone.
I am not suggesting those two facts are connected by intent. They are connected by direction. In both cases the record of the work became less legible than the work, and in both cases the score kept arriving on time.
The score is the easy part. It was always the easy part. A number that arrives on schedule is the single least informative thing a system produces, because it is the thing the system was built to produce.
What I would take from this
Not that Astra is unimpressive. Forty-two percent on eight hundred real vulnerabilities is a serious result, and OpenAI restricting the model to review and patching after it found two previously unknown flaws is the behaviour of people who believe their own numbers.
What I take is narrower.
When somebody tells you a threshold has been crossed, ask which instrument, under which harness, and go and read what the people who built the instrument said about it. In this case they said, in writing, that it does not prove the thing. That took me four minutes to find and it did not appear in the coverage.
And ask what the exam cannot contain. Not to dismiss the result, but because the gap between the bounded room and the open floor is not a technicality. It is the entire location of the job.
Counting is not choosing. A model that solves 99.9 percent of a deterministic environment has told you something precise and genuinely impressive about deterministic environments. Whether it can do most economically valuable work is a question nobody has built the room for yet, because the room would have to be the world.
I have sat the exams. I passed. I am still, most nights, finding out.