Since ChatGPT · 30 Nov 2022 to today

AI Progress

Five lines of AI since ChatGPT, on one clock: what it achieved, what the models learned, the harness built around them, the research and the compute. An agent is a model using tools in a loop, and each leap came when an improvement on one side met something on the other that could use it.

achievements a model line a harness line research compute dotted until its first real release a technology open weights open source a lab founded where lines met in a breakthrough fed into (hover a station) AGI, by my definition
Five timelines of AI, November 2022 to today The achievements timeline (science, math contests, mathematics, coding, benchmarks saturated) above the model timeline (frontier models, reasoning, tool use, long context, agentic training, screens, open weights) above the harness timeline (agent frameworks, coding agents, personal agents, embedded agents, subagents, agent messaging, skills, tools and MCP, connectors and plugins, context and memory, sandboxes, browser use, computer use) the research timeline (methods and benchmarks), and the compute timeline (chips and clusters, fast inference, inference engines, running locally), each development a station, with the breakthroughs where model and harness met marked as numbered interchanges on a rail between them. Open-weight releases are green squares, open-source ones amber squares, and a lab’s founding a diamond; a strip on the left holds everything from 2015 to ChatGPT. Two panels below share the time axis: METR’s time horizon for frontier models, and the Artificial Analysis Intelligence Index of the best model and the best open-weight model released by each date. A hatched band marks January to June 2026, when by the author’s definition, better than most people on most digital tasks, AGI was crossed. Every station and index record is also listed in the tables below the map.

swipe the map sideways · tap a station for its story

AGI tracker · verdicts reviewed

Is it AGI yet?

My own definition first, then eleven published definitions of artificial general intelligence, from lab charters to kitchen tests, each held up against public evidence on the frontier of GPT-6 Astra, Claude Fable 5.1 and Gemini 3.7. The evidence is independent measurement where it exists, and a lab’s own figure is labelled as one. The verdicts are my judgment.

Four measurements carry most of the weight

The definitions split on reliability

Each definition is placed by how much consistency it demands, from one good run to human-level dependability, and by whether it needs the physical or legal world. Placements are editorial. Click a dot to read the entry.

Definitions of AGI placed by the reliability each demands and by whether each requires the real world

metpartiallynot methonourable mentionwhere the frontier’s reliability reaches

Eleven definitions, and two more

Every development, as a table
DateLineDevelopmentWhoOpenSource
Every record on the Intelligence Index, as a table

ReleasedRecordModel, as Artificial Analysis ran itMakerIndex

Sources

Changelog