The Model Wars · Preface
Preface: It Starts with One Paper
On June 12, 2017, a paper was uploaded to arXiv.
No launch event, no countdown. On a trading floor, nobody looked up for it. It entered a database, filed among the many papers of that day, a pebble dropped into the ocean.
The paper had eight authors and a title: Attention Is All You Need. The problem they were solving was machine translation — at the time a specialized, very concrete problem.
Nobody said they were inventing an era.
The history that matters usually arrives without opening music.
Within a few years, that architecture grew GPT, BERT, and ChatGPT, and grew an enormous industry around chips, capital, energy, copyright, and competition between nations.
Looking back, it’s easy to write the whole thing as a straight line: the Transformer appears, models get bigger, intelligence emerges, the world changes.
History never runs straight.
It behaves more like groundwater — branching in the dark, going around rock, building pressure exactly where you assumed things were solid, and finally breaking the surface somewhere nobody was watching.
This is not a textbook on algorithms.
We’ll talk about attention without asking you to derive anything. We’ll talk about parameters and compute without building walls out of numbers.
The questions actually worth chasing are elsewhere. Why did Google invent the Transformer and then let OpenAI be the one to turn it into a consumer product? Why did a generative approach that looked like the weaker bet end up swallowing the entire field of natural language processing? Why does “open” get more closed every year, and why does closure keep producing new openness? How did a quant fund in Hangzhou spend a reported $5.57 million on one training run and take $589 billion off Wall Street in a single day?
The history of technology looks like it’s about machines. Underneath, it’s still about people.
Some believe in scale; some are afraid of losing control. Some publish the paper; some lock the weights in a server. Some walk out of a large company; some come back with capital. Arguments over direction, organizational inertia, personal ambition, and the accidents of timing together decided what the tools in your hands look like today.
You don’t need to understand AI before reading this. You need to hold on to three questions.
What did the models learn? Whose hands did those capabilities end up in? And when exactly did machines move from answering questions to completing tasks?
Those three questions map to three threads that run through the book.
The first is capability. From GPT-1’s hundred-odd million parameters to Kimi K3’s 2.88 trillion; from GPT-2 being “too dangerous to release” to DeepSeek-R1 shipping in full under an MIT license. This thread follows what machines can actually do, and how the edge of that keeps getting pushed further out.
The second is deployment. However strong the technology, it’s a castle in the air until it lands. This thread follows how capability becomes something ordinary people can touch — from a developer API, to a text box, to that feature on your phone you use daily and no longer bother calling “AI.”
The third is agents. The youngest thread, and possibly the most important. AutoGPT took a hundred thousand GitHub stars in a matter of weeks in 2023 and collapsed at the same speed; by 2025, a command-line tool was doing a billion dollars in annualized revenue. This thread follows when machines stop answering and start doing.
The three run in parallel, cross, and occasionally strike sparks off each other.
Watch them, and you’ll see a picture larger than any single event in it.
The book is in five parts, running from 2017 to 2026.
The first two require no technical background. They read as history because they are history. From the third part on, the narrative takes on more technical comparison and industry analysis. Parts four and five touch some specifics — mixture of experts, GRPO, thinking budgets, MCP, test-time compute.
One promise up front: if a technical passage starts to feel like work, skip it.
Skipping won’t cost you the story. Each chapter holds together on its own, and you can leave the technical passages out without losing the thread.
One last thing, honestly.
We have said wow more times over these nine years than we can count. Wow, ChatGPT is amazing. Wow, DeepSeek wrecked Nvidia’s stock with a paper. Wow, a cabinet secretary switched off a model.
The wow is genuine. The trouble is what comes after it: you close the tab, and the next day looks exactly like the last one.
What this book is trying to do is simple. Put something after the wow — an oh, so that’s how that happened.
We’re standing inside this history right now. The ending is too far off to see, and the water is already over our feet.
Facts in a book like this age. The current errata, and whatever we publish after it, live at pangoralabs.com/books/the-model-wars. If you catch something wrong, corrections@pangoralabs.com reaches us.
The story starts with that paper, the one with no opening music.