Eleven definitions of artificial general intelligence, from lab charters to kitchen tests, each held up against public evidence on the frontier of GPT-6 Astra, Claude Fable/Mythos 5.1, and Gemini 3.7. The evidence is independent measurement, not what the labs claim. The verdicts are one reader's judgment.
The one definition fully met, Norvig and Agüera y Arcas's “AGI is already here”, asks only for generality. DeepMind's “Competent” level comes closest after it, and is a low bar by design. “Partially” usually means jagged: superhuman in some domains, unreliable or absent in others.
The evidence
Four measurements carry most of the weight
Scores come from neutral harnesses where one exists. A lab's own harness can tell a different story: the same Astra model scores 99.9% on ARC-AGI-3 with its provider adapter.
ARC-AGI-3
62.7% vs ~100% human
GPT-6 Astra on the semi-private set under ARC Prize's standard harness, at $26,098. ARC Prize says outright that this does not make Astra AGI.
For tasks it finishes half the time, the length is at METR's 16-hour ceiling. For tasks it finishes 4 times in 5, the length is about 3 hours. Latest published entry: Claude Mythos Preview (early).
Research-level problems written by mathematicians are close to solved. This is the Nobel-laureate-level end of the jagged profile.
Epoch AI · Glazer et al. 2024
The pattern
The definitions split on reliability
Each definition is placed by how much consistency it demands, from “sometimes does it” to “does it as dependably as a person”, and by whether it needs the physical or legal world. Placements are editorial. Click a dot to read the entry.
metpartiallynot methonourable mentionwhere the frontier's reliability reaches
Tolerate 50% success on hard tasks, and it is roughly met. Everything in or left of the band is met or partly met, because a system that is brilliant on its good runs satisfies it. Just past the band, Amodei's definition is met only in the fields where the frontier is strongest.
Demand human-level consistency or a body, and it is not. “Is it AGI?” is less useful than “how reliable, and where?”
The ledger
Eleven definitions, and two more
Change log
What has moved, and when
Verdicts are revisited as new measurements land. Each review is recorded here, including reviews that change nothing.