AGI tracker · reviewed

Is it AGI yet?

Eleven definitions of artificial general intelligence, from lab charters to kitchen tests, each held up against public evidence on the frontier of GPT-6 Astra, Claude Fable/Mythos 5.1, and Gemini 3.7. The evidence is independent measurement, not what the labs claim. The verdicts are one reader's judgment.

The one definition fully met, Norvig and Agüera y Arcas's “AGI is already here”, asks only for generality. DeepMind's “Competent” level comes closest after it, and is a low bar by design. “Partially” usually means jagged: superhuman in some domains, unreliable or absent in others.

The evidence

Four measurements carry most of the weight

Scores come from neutral harnesses where one exists. A lab's own harness can tell a different story: the same Astra model scores 99.9% on ARC-AGI-3 with its provider adapter.

ARC-AGI-3

62.7% vs ~100% human

GPT-6 Astra on the semi-private set under ARC Prize's standard harness, at $26,098. ARC Prize says outright that this does not make Astra AGI.

ARC Prize, 3 Sep 2026 · blog/astra

METR time horizon

5.6× reliability gap

50%17.4 h
80%3.1 h

For tasks it finishes half the time, the length is at METR's 16-hour ceiling. For tasks it finishes 4 times in 5, the length is about 3 hours. Latest published entry: Claude Mythos Preview (early).

METR Horizon v1.1 · time-horizons

OSWorld 2.0

72.6%

Computer use: completing real tasks in desktop apps and operating systems through the screen. Strong, but short of reliable.

as reported · thepcenthusiast

FrontierMath Tier 4

≈ saturated

Research-level problems written by mathematicians are close to solved. This is the Nobel-laureate-level end of the jagged profile.

Epoch AI · Glazer et al. 2024

The pattern

The definitions split on reliability

Each definition is placed by how much consistency it demands, from “sometimes does it” to “does it as dependably as a person”, and by whether it needs the physical or legal world. Placements are editorial. Click a dot to read the entry.

Definitions of AGI placed by reliability demanded and by whether they require the real world
metpartiallynot methonourable mentionwhere the frontier's reliability reaches

Tolerate 50% success on hard tasks, and it is roughly met. Everything in or left of the band is met or partly met, because a system that is brilliant on its good runs satisfies it. Just past the band, Amodei's definition is met only in the fields where the frontier is strongest.

Demand human-level consistency or a body, and it is not. “Is it AGI?” is less useful than “how reliable, and where?”

The ledger

Eleven definitions, and two more

Change log

What has moved, and when

Verdicts are revisited as new measurements land. Each review is recorded here, including reviews that change nothing.