Ataraxos Beats Stratego's Top Player 15-1-4 After Training on 16 H100s for a Week
A.I. / news
Ataraxos Beats Stratego's Top Player 15-1-4 After Training on 16 H100s for a Week
A Nature paper from MIT, Carnegie Mellon, NYU and Stanford puts the compute bill under $8,000, against an estimated $3 million to $4.5 million for DeepMind's DeepNash.

Ataraxos, a Stratego-playing system built by researchers at MIT, Carnegie Mellon University, New York University and Stanford, beat Pim Niemeijer 15 wins, 1 loss and 4 draws in a 20-game series, according to The Decoder. Niemeijer is a four-time world champion who held the top ranking for more than 600 weeks. The paper appeared in Nature.
The sources disagree on the publication date. The Decoder gives Oct. 1, 2026. Remio says Sept. 30. The Nature page itself was not fetched for this report.

What the training run cost
Training took one week on 16 Nvidia H100 GPUs, plus four days on four H100s for the belief network, The Decoder reports. The cost came to under $8,000 at 2025 prices. Remio adds that the paper's figure covers reported training compute only, not researcher salaries or infrastructure.
The authors estimate DeepMind's 2022 DeepNash system would cost $3 million to $4.5 million at 2025 hardware prices, based on 1,024 TPU v3 nodes for two to three months. Ataraxos used about one-hundredth of the training examples and one-thirtieth of the self-play games, according to Remio.
- Ataraxos (upper bound)8000 USD
- DeepNash (low estimate)3M USD
Source: Remio and The Decoder, citing the Nature paper's estimates, accessed 2026-10-02
| Measure | Ataraxos | DeepNash |
|---|---|---|
| Hardware | 16 H100s, 1 week, plus 4 H100s, 4 days | 1,024 TPU v3 nodes, 2 to 3 months |
| Training cost | Under $8,000 | $3 million to $4.5 million (estimate) |
| Match result | 15-1-4 vs Niemeijer | Not compared in sources |
How it plays
The system combines a blueprint strategy learned through self-play with decision-time planning, Remio reports. A belief network predicts where the opponent's hidden pieces sit before each move. The Decoder puts the number of possible hidden setups at more than 10^33, which is why the game stumped earlier systems.
The Decoder also reports that regularisation forces the system to vary its play, so opponents cannot predict it. Remio says Ataraxos won 38 of 40 games at a 2025 World Championship demonstration. Both sources also report results beyond Stratego: record scores in Barrage Stratego and in the card game Hanabi, where The Decoder says it needed "two orders of magnitude less compute" than earlier benchmarks.
The limits the authors state
The search step mimics a single learning step, The Decoder reports, so more compute at play time cannot keep improving it. Senior author Gabriele Farina is named by Remio; his title was not given in the material fetched.
The authors say humans must keep authority over any real-world use such as negotiation or cybersecurity, and that this requires "a way to inspect the model's decisions" before deployment.
The cost claim is the authors' own, and the DeepNash number is an estimate rather than a reported bill. The same pattern of small, cheap models rivalling large ones appears in our report on Qwen3.8-27B fitting in 17GB, and the money tied up in Nvidia hardware is the subject of our piece on Amazon's Nvidia chip leaseback.
What to watch
No code or weights release was mentioned in the sources fetched. Independent replication of the sub-$8,000 figure would need the paper's training configuration.
Sources
More in A.I.
- 01GPT-Synopsys: OpenAI Gets Paid Only When the Chips It Helps Design Beat the Customer's BaselineThe Sept. 30 deal pairs an OpenAI model with Synopsys' design software, with no price, no release date and no named customer.
- 02Nine Mathematicians Advising OpenAI Ask AI Labs to Stop Testing Hard Problems on Models Nobody Else Can UseThe Advisory Group on Mathematics and AI published its rules on Sept. 29: release fast, fund human understanding, disclose prompts and costs.
- 03Qwen3.8-27B Ships Under Apache 2.0 and Fits in 17GB, but Spends 160 Million Tokens Where the Median Spends 43 MillionAlibaba's open-weight model scores 52 on Artificial Analysis's Intelligence Index. Its own benchmark figures are vendor-supplied, and users report slow runs.
- 04OpenAI Ties Moonshot AI to a July Campaign That Replayed Encrypted Reasoning, Offers No Evidence PubliclyOpenAI says 16,000 requests from more than 4,000 accounts tried to recover hidden model reasoning. Its attribution to Moonshot rests on its own assertion.