Ataraxos AI Beats Stratego Legend Pim Niemeijer 15-1

Ataraxos AI Beats Stratego Legend Pim Niemeijer 15-1

Stratego was one of the few major competitive board games where humans still had the edge over machines. That may now be over. A research team from Carnegie Mellon University, New York University, Stanford University and MIT has built Ataraxos, a Stratego AI that clearly beat the most successful player in the game's history. According to the team, training cost less than $8,000.

The work is published in Nature. As far as the authors know, Ataraxos is the first AI to reach superhuman level in Stratego.

A one-sided match against the best

In an official 20-game series, Ataraxos played Dutch player Pim Niemeijer. The final score was 15 wins, four draws and one loss for the AI.

Niemeijer's record is unmatched. He holds four world championships, 15 Dutch national titles and two online world championships. He also spent more than 600 weeks at number one in the world rankings. George Franka, the only player who has competed in every world championship since 1997, calls him "the best Stratego player of all time."

The series ran over three weeks to avoid fatigue and to give Niemeijer time to prepare. He was paid $1,000 to take part, plus $100 per win and $50 per draw, so he had a reason to play his best. He also knew the AI would not adapt to his style.

Counting draws as half a win, Ataraxos reached an effective win rate of 85 percent. The authors call that unprecedented at the top level. Max Roelofs, a three-time world championship runner-up, notes that margins among elite players are usually razor-thin because the game forces risk-taking. The AI also played with a structural handicap: Niemeijer could adjust to it across games, but it could not adjust to him. Three-time world champion Vincent de Boer considers that a major disadvantage, according to the paper.

At an exhibition during the 2025 Stratego World Championship, Ataraxos won 38 of 40 games against tournament players. Opponents describe its style as hard to read. It bluffs in ways humans see as too risky, stalls stubbornly when behind, and seems almost eerily lucky with piece placement.

Why Stratego was hard

Each side places 40 pieces face down. The paper counts more than 10^33 possible setups. First author Samuel Sokota explained on X that in games with hidden information, the value of a move depends not just on what comes next but also on what happened before and during the decision.

Methods from poker AI worked around this, but only when hidden information was limited. Texas Hold'em has just 1,326 possible starting hands, and the compute those methods need grows with the amount of hidden information.

A fraction of DeepMind's budget

DeepMind tried before with DeepNash, which did not get past top human level. DeepNash trained on 1,024 TPU nodes for two to three months. The Ataraxos authors estimate that would cost $3 million to $4.5 million at 2025 prices. At the 2023 world championship, DeepNash won 19 of 28 games but lost to most top players, Niemeijer included.

Ataraxos needed one week on 16 Nvidia H100 GPUs, plus four days on four GPUs for its belief network. The researchers put that at roughly 1/500th of the compute cost, 1/30th of the self-play games and 1/100th of the training examples. They credit a custom GPU simulator and much higher sample efficiency.

A direct match never happened. The team offered to build the infrastructure, but DeepMind said the DeepNash code no longer works.

How it learns

Ataraxos uses no human data. It learns only through self-play. The key is regularization: an extra condition during training pushes the AI to vary its setups and moves instead of settling on a fixed strategy too early. Predictable Stratego players get exploited.

The pressure eases over time, and step sizes shrink with it. Early on, the AI varies heavily and learns in big jumps. Later, it makes small corrections. The authors compare regularization to an energy reserve: burn it too fast and learning stalls.

A belief network predicts the opponent's hidden pieces. Before each move, Ataraxos samples possible game states, tests candidate moves and runs an extra learning step for that decision.

Beyond one game

The same method won four 50-game Barrage Stratego series against three of the four top-ranked players. It set records in every variant of the card game Hanabi and beat PerfectDou and DouZero in the Chinese card game Dou dizhu.

The authors argue that large amounts of hidden information no longer block reinforcement learning and search, provided fast, accurate simulators exist. They name financial markets, military conflicts and negotiations as examples. One limit remains: because the search mimics only a single learning step, adding more compute time does not keep improving results. The code is public.

The Bigger Picture

The headline is the win, but the cost figure may matter more. A university team beat a result that a major lab spent millions on, which suggests that clever training design can still outweigh raw compute in some domains.

The leap to markets or negotiations is a bigger one. The authors themselves tie it to a condition: good simulators. That links this work to the wider push for simulated environments, such as the world model startups attracting large funding rounds. Real-world settings rarely come with clean rules.

It is worth watching whether others reproduce the results with the open code, and whether a stronger search method removes the compute ceiling the authors flag.