Attention Is All You Need
Google researchers introduce the Transformer, an architecture for machine translation.
Every model on this page is a descendant. It parallelises well, so it scales with compute.
In 2018 a language model could finish your sentence, badly. Today it can take a ticket, read the codebase, write the fix, run the tests and open the pull request. This is the log of how we got here, one stop at a time, and how far the road still runs to recursive self-improvement: AI that meaningfully builds its own successor.
Models that predict the next word. Impressive party tricks, no idea what you actually wanted.
Google researchers introduce the Transformer, an architecture for machine translation.
Every model on this page is a descendant. It parallelises well, so it scales with compute.
OpenAI pre-trains a 117M-parameter Transformer on books, then fine-tunes it per task.
The recipe that stuck: learn language first, specialise later.
Google's bidirectional encoder sweeps the language-understanding benchmarks and ends up inside Search.
Proof that pre-trained Transformers were useful for real products, not just papers.
1.5B parameters, released in stages because OpenAI called it too dangerous to publish at once.
It wrote convincing paragraphs that fell apart a page later. The first public argument about release safety.
Kaplan et al. show loss falls predictably as you add parameters, data and compute.
Turned "make it bigger" from a hunch into an investment plan.
175B parameters. Show it a few examples in the prompt and it picks up the task without retraining.
In-context learning: the model became programmable with plain text.
Models tuned to follow instructions and hold a conversation. Suddenly everyone could use one.
Codex, a GPT-3 fine-tuned on code, ships as autocomplete inside the editor.
The first AI tool developers used every day. Code turned out to be the killer domain.
Asking a model to "think step by step" boosts its maths. RLHF makes a small model preferred over GPT-3.
Two levers beyond size: how the model reasons and what it is trained to want.
A prompting pattern that interleaves reasoning with tool calls: think, act, observe, repeat.
The skeleton of every agent loop that followed.
A free chat box on top of GPT-3.5. A million users in five days.
The moment the rest of the world noticed.
Meta releases weights to researchers. They leak within a week and open models take off.
Anyone with a decent GPU could now experiment, not just three labs.
Same day: GPT-4 passes the bar exam in the top tenth, and Anthropic opens up Claude.
The capability jump that made people take "agents" seriously.
GPT-4 in a loop with a goal and a to-do list. Top of GitHub, mostly going in circles.
The idea was right, the models were not ready. A preview, not a product.
Models trained to think before answering, and given hands: a mouse, a keyboard, a terminal.
Claude 3 Opus edges past GPT-4. Cognition demos Devin, billed as the first AI software engineer.
Frontier was now a race, and "AI engineer" became a pitch.
A mid-size model that writes better code than the previous flagships.
The model that made Cursor-style coding mainstream.
OpenAI trains a model with reinforcement learning to produce long hidden reasoning before it answers.
A new scaling axis: spend more compute at answer time, get better answers.
Claude looks at screenshots and drives the mouse and keyboard. Slow and clumsy, but it works.
Any software with a screen became a possible tool.
Anthropic publishes an open standard for connecting models to tools and data.
Plumbing, not magic, and that is why everyone adopted it.
A reasoning model scores 75–88% on a puzzle benchmark built to resist language models.
The "they only memorise" argument lost a lot of ground.
An open-weights reasoning model from China, near o1 level, reportedly trained on a fraction of the budget.
Reasoning was not a moat. Nvidia lost roughly $600B in market value in one day.
Models that work for hours on their own: they plan, use tools, check their work and hand back results.
Agents that browse for half an hour and write a cited report. Claude Code puts an agent in the terminal with your repo.
Agents stopped being demos and became tools people paid for.
METR measures how long a task (in human time) agents can finish reliably: it has doubled roughly every seven months since 2019.
The closest thing this road has to a speedometer.
Opus 4 works on code autonomously for hours. OpenAI's Codex runs parallel coding agents in the cloud.
"Assign a ticket to the AI" became a normal workflow.
Experimental models from OpenAI and Google DeepMind solve five of six IMO problems in natural language.
Real novel reasoning under exam conditions, no formal-proof crutches.
OpenAI merges the chat and reasoning lines into one model that decides how long to think.
Reasoning became the default, not a separate product.
A week apart, both labs push the frontier on coding and long agentic tasks.
The gap between the labs shrank to weeks.
The first version of this site, from DNS records to deploy script, was built by an AI agent in one terminal session.
A small data point. The human still picked the idea and the domain.
The point where AI does most of the work of building better AI, and each generation speeds up the next. Labs already use their own models to write their code. Whether that loop closes, and how fast, is the open question.
The road isn't paved yet. New stops get added as they happen.