On the morning of December 5, 2017, DeepMind published a preprint on arXiv describing something that had seemed impossible. A program called AlphaZero had started from nothing — no opening book, no endgame tablebases, no human games — and learned chess from first principles in four hours. Then it played Stockfish 8, the strongest traditional engine in the world. In 100 games with a one-minute-per-move time control, AlphaZero won 28 and lost none. Stockfish, which searches roughly 70 million positions per second, could not beat it once.
That result did something unusual in computer chess: it surprised everyone, including the researchers who built it.
How Did AlphaZero Learn Chess Without Human Knowledge?
AlphaZero learned chess through pure self-play reinforcement learning, starting from nothing but the rules of the game. No opening theory, no annotated grandmaster games, no hand-crafted evaluation function — just the legal moves and the binary outcome of win, loss, or draw. It played millions of games against itself, adjusted its neural network after each training batch, and gradually improved until it surpassed anything humans had built.
Traditional engines like Stockfish work through brute-force alpha-beta search: they enumerate candidate moves many plies ahead, evaluate leaf positions using a hand-tuned function, and propagate the best values back to the root. Stockfish evaluates roughly 70 million positions per second. AlphaZero evaluates approximately 80,000 per second — but chooses which positions to examine far more selectively, guided by a deep neural network that simultaneously estimates position value and assigns move probabilities. It searches less and understands more.
DeepMind’s training used Google’s custom Tensor Processing Units (TPUs). The AlphaZero paper, published in the journal Science in December 2018 under lead author David Silver and colleagues, reported that after just four hours of self-play training, AlphaZero’s chess performance had surpassed that of Stockfish 8. After nine hours, the team ran the 100-game exhibition match that would make headlines.
The program had developed its entire understanding of chess — openings, middlegame plans, endgame technique — without consulting a single human game.
The December 2017 Match: AlphaZero 28, Stockfish 0
The 100-game match used a time control of one minute per move. AlphaZero played as both White and Black. Final score: 28 wins for AlphaZero, 72 draws, zero wins for Stockfish.
The result ignited immediate debate, and much of that debate was well-founded. Stockfish ran on a commodity CPU server using 64 threads. AlphaZero had access to Google’s specialized TPU hardware. Stockfish was not permitted to use its opening book or endgame tablebases — two resources that are core to how traditional engines operate in practice. The hardware and conditions were not equivalent.
These objections were taken seriously enough that DeepMind eventually published a second, more carefully controlled study. In the 2018 Science paper, the team ran AlphaZero against a later version of Stockfish under conditions intended to be fairer. The results held: across 1,000 games starting from standard positions, AlphaZero won 155, drew 839, and lost 6. In 1,000 games starting from positions favoring Black, AlphaZero still dominated.
The hardware debate has never been fully settled; different configurations produce different margins. What is not in dispute is what the games themselves revealed.
DeepMind released 10 sample games from the original match. Chess players and analysts who studied those games did not primarily argue about hardware. They argued about something more interesting: why AlphaZero played the way it did.
What Made AlphaZero’s Chess Style Unusual?
AlphaZero played like no engine anyone had seen. Its games showed four recurring patterns that distinguished it sharply from Stockfish and every other traditional program.
Piece activity over material count. AlphaZero would sacrifice a pawn — sometimes more — for long-term positional compensation that alpha-beta engines running at shallow depth couldn’t properly evaluate. Where Stockfish would correctly identify an immediate material deficit and defend it, AlphaZero accepted structural weaknesses in exchange for piece mobility and initiative, and converted those advantages over many moves. The positions it sought looked, to human eyes, like the kind of investments a grandmaster makes — not a machine.
Wing pawns as weapons. Moves like a4 and h4, pushing pawns on the flank early in the middlegame, appeared repeatedly across AlphaZero’s games. Human opening theory had generally treated these moves as slow or optionally weakening. AlphaZero played them aggressively, using expanded space to support piece maneuvers, gain control of key squares, and gradually restrict the opponent. The h-pawn push, in particular, became a signature. In the Queen’s Indian game below, AlphaZero advances its h-pawn to h5 and then uses the resulting weakness on g6 as the foundation for a long technical squeeze.
King safety as a long-term investment. Traditional engines treat the uncastled king as a near-term liability to be resolved quickly. AlphaZero sometimes kept its king in the center or used it as an active piece in the endgame far earlier than standard evaluation functions would recommend — and profited from the extra flexibility.
The Berlin Defense rediscovery. One of the most striking findings from a 2022 PNAS study by Thomas et al., which used interpretability tools to analyze AlphaZero’s neural network, was that during training AlphaZero independently converged on the Berlin Defense (1.e4 e5 2.Nf3 Nc6 3.Bb5 Nf6) and began preferring it strongly. For many years, human theory had rated the Berlin as marginally inferior to 3…a6. Elite grandmasters had only recently elevated it to its current status as a theoretical cornerstone. AlphaZero arrived at the same conclusion without being shown it — a sign that its neural network was learning something real about chess structure, not just memorizing patterns from games.
The Queen’s Indian: A Lesson in Patient Pressure
Among the ten games DeepMind published, the Queen’s Indian Defense game is one of the most instructive demonstrations of how AlphaZero converts small advantages through relentless positional pressure.
Playing White with 1.d4 Nf6 2.c4 e6 3.Nf3 b6 4.g3 Bb7 5.Bg2, AlphaZero sets up the fianchetto system favored throughout the match. After 5…Bb4+ 6.Bd2 Be7 7.Nc3, the game takes a sharp turn with 8.e4 d5 9.e5 Ne4. AlphaZero captures Stockfish’s knight and establishes a mobile pawn center.
The key moment comes at move 10…Ba6, where Stockfish lifts its bishop to a6 to contest the c4 pawn. AlphaZero responds with 11.b3 Nxc3 12.Bxc3 dxc4 — accepting a disrupted pawn structure in exchange for open lines — and then plays 13.b4 to launch queenside expansion. AlphaZero’s advantage is not material; it’s structural. Its pawn center controls space, its bishop on g2 eyes the long diagonal, and its knight heads for an outpost.
The endgame reveals AlphaZero’s style most clearly. Rather than seek a quick tactical blow, it maneuvers its pieces into optimal coordination, uses the h-pawn as a battering ram (h4–h5 forcing …g6), and then switches the attack vector once Black’s defenses adapt. By move 44, AlphaZero has traded the queens and entered a technically winning rook-and-bishop endgame. The conversion is methodical: Stockfish defends for 24 more moves, making no obvious blunders, until the position simply collapses under accumulated pressure.
There is no moment where Stockfish makes a visible mistake. AlphaZero makes the position worse, move by move, until resistance is objectively pointless. Stockfish resigns after 68 moves.
AlphaZero vs Stockfish 8, London 2017 — Queen's Indian Defense
| 1. | ||
| 2. | ||
| 3. | ||
| 4. | ||
| 5. | ||
| 6. | ||
| 7. | ||
| 8. | ||
| 9. | ||
| 10. | ||
| 11. | ||
| 12. | ||
| 13. | ||
| 14. | ||
| 15. | ||
| 16. | ||
| 17. | ||
| 18. | ||
| 19. | ||
| 20. | ||
| 21. | ||
| 22. | ||
| 23. | ||
| 24. | ||
| 25. | ||
| 26. | ||
| 27. | ||
| 28. | ||
| 29. | ||
| 30. | ||
| 31. | ||
| 32. | ||
| 33. | ||
| 34. | ||
| 35. | ||
| 36. | ||
| 37. | ||
| 38. | ||
| 39. | ||
| 40. | ||
| 41. | ||
| 42. | ||
| 43. | ||
| 44. | ||
| 45. | ||
| 46. | ||
| 47. | ||
| 48. | ||
| 49. | ||
| 50. | ||
| 51. | ||
| 52. | ||
| 53. | ||
| 54. | ||
| 55. | ||
| 56. | ||
| 57. | ||
| 58. | ||
| 59. | ||
| 60. | ||
| 61. | ||
| 62. | ||
| 63. | ||
| 64. | ||
| 65. | ||
| 66. | ||
| 67. | ||
| 68. |
Import this game into PGNBase and step through the middlegame from move 22 onward. The point at which AlphaZero’s advantage becomes structural rather than tactical — the moment when Stockfish is defending correctly but the position is still getting worse — is not a single move but a sequence. The PGNBase analysis tools let you mark candidate positions and compare engine lines at any depth.
How AlphaZero Changed Opening Theory
Chess theory has always flowed from the top down: elite players develop ideas, those ideas spread through tournament play and books, and the broader community absorbs them over years. AlphaZero introduced a different direction — theory developed from scratch through self-play, with no human input, that converged on human conclusions and in some areas went beyond them.
The King’s Indian Attack as a universal system. AlphaZero showed a strong preference throughout the match for the setup with g3 and Bg2, playing it against a wide range of Black defenses. Its handling of this structure — patient piece improvement, willingness to sit on microscopically better positions for thirty or forty moves, and precise timing of pawn breaks — gave human analysts a new framework for the system. The game above is a direct example: AlphaZero’s Bg2 fianchetto becomes a long-range weapon only in the endgame, after forty moves of preparation.
Wing pawn advances. After the ten sample games were published, analysts noted that AlphaZero systematically used a4 and h4 to gain space and create weaknesses, treating wing pawns as attacking pieces rather than passive structural elements. This observation influenced how some players at the elite level approached the middlegame in subsequent years.
The Sicilian Defense. Rather than entering the sharpest theoretical lines of the Open Sicilian, AlphaZero frequently chose slower setups that allowed it to build long-term pressure. Grandmaster Matthew Sadler and Natasha Regan, in their 2019 book Game Changer — the most thorough published analysis of AlphaZero’s games in English — concluded that AlphaZero played the Sicilian the way a positional grandmaster would: avoiding sharp forced lines in favor of positions where the engine’s ability to assess long-term compensation gave it a structural edge. If you want to see how the Sicilian looks from a player who prefers structure over tactics, the Sicilian Defense guide covers the main variations and what each one demands from both sides.
The Berlin Defense. As noted above, AlphaZero’s independent rediscovery of the Berlin Defense — and its strong preference for it against 1.e4 e5 — was one of the more striking findings from the 2022 PNAS interpretability study. The network’s internal representations developed concepts corresponding to material value, piece mobility, king safety, and pawn structure without being told those categories exist. It invented the vocabulary of chess strategy on its own.
Leela Chess Zero and AlphaZero’s Lasting Influence
DeepMind never released AlphaZero’s weights, code, or training infrastructure publicly. The engine exists only in two papers and ten games. What it inspired, however, was substantial.
Shortly after the AlphaZero paper circulated, a group of volunteers launched Leela Chess Zero (Lc0), an open-source project applying the same neural-network self-play approach using donated GPU compute from the chess community. Within a year, Lc0 had reached superhuman strength. It now competes regularly in the Top Chess Engine Championship (TCEC) alongside Stockfish.
The two programs play completely differently. Stockfish’s brute-force alpha-beta search versus Lc0’s neural-network Monte Carlo Tree Search represents a genuine philosophical divide about what chess-playing intelligence means. Their head-to-head results are roughly matched, with Stockfish maintaining a narrow edge in overall performance while Lc0 regularly wins games through the kind of deep positional sacrifice that AlphaZero pioneered. Neither approach has decisively beaten the other in controlled conditions.
The influence on human chess is harder to quantify. But in the years since 2017, several trends at elite level — wider use of the Berlin Defense, more aggressive wing pawn advances, and a greater willingness to accept long-term positional compensation over immediate material equality — have been at least partly attributed to the framework AlphaZero demonstrated. Whether grandmasters consciously absorbed the engine’s ideas or arrived at them independently is impossible to say with certainty. The correlation exists.
For anyone interested in the full arc of how computers have reshaped competitive chess — from Shannon’s 1950 proposal through the Deep Blue era and into neural networks — the how engines changed chess overview covers the whole story, including why the evaluation functions that traditional engines rely on are fundamentally different from what AlphaZero learned.
The Immortal Game of 1851, which you can read about in the Immortal Game post, involved Adolf Anderssen sacrificing two rooks, a bishop, and his queen to deliver checkmate with three minor pieces. AlphaZero’s wins look nothing like that. There are no sudden queen sacrifices or mating attacks. There is just an accumulation of small advantages — a better bishop here, a more active rook there, a slightly more advanced pawn — that grows over forty moves into something insurmountable. It is chess as a long argument, not a short fight.
Why AlphaZero Still Matters
AlphaZero did not win every game. Stockfish drew 72 times from the 100-game match, and under fairer hardware conditions the margin narrows further. The question of which program is “better” was never definitively resolved, and in some sense it cannot be: the two programs embody different philosophies about what playing strength means.
What cannot be disputed is the effect on how we think about chess and artificial intelligence simultaneously.
Before AlphaZero, computer chess meant faster search, better evaluation of known patterns, and deeper calculation. After AlphaZero, it also meant self-play training, neural network inference, and positional intuition that surprised experts. A program that had never read a single game of chess had independently converged on concepts that took human players centuries to articulate. That finding — documented in the 2022 PNAS study — is not primarily about chess. It is about how learning works.
For chess players specifically, the ten published games remain worth studying, not for the moves alone but for what they imply about what the game is. The engine that dismantled Stockfish in December 2017 did so by playing positions that looked, to human eyes, like someone who had understood something about chess that the rest of us were still learning.
Load the Queen’s Indian game in PGNBase, step through the endgame, and try to find the move where Stockfish’s position became objectively hopeless. It is harder to identify than the queen sacrifice in the Immortal Game. That difficulty is precisely the point.