MIT’s Ataraxos Beats Stratego Champions Without Requisitioning a Small Power Plant
MIT-linked Ataraxos beats elite Stratego players with far less compute. A new Nature paper shows real progress, with real limits beyond the game board.
Cambridge has found another way to make game night require a methods section. This time, the impressive part is that the computer can win without submitting a capital request resembling a municipal infrastructure plan.
Ataraxos, an AI system developed by researchers at MIT, Carnegie Mellon, NYU, and Stanford, is the subject of a Nature paper published September 30. It defeated Stratego champion Pim Niemeijer with 15 wins, one loss, and four draws. The researchers report roughly one five-hundredth of the reinforcement-learning compute cost of the earlier DeepNash effort.
The calendar needs its own chair at this table. An earlier preprint was submitted November 10, 2025. Today’s news is the journal publication, not a match played this morning or a newly launched commercial service. Science takes time to put on its formal shoes.
My verdict: a meaningful research win, especially for people who would prefer progress in AI to include better algorithms. The achievement deserves attention without being promoted to chief operating officer of everything.
Everybody brought a secret to game night
Stratego puts 40 pieces on each side, with identities concealed from the opponent. You are trying to capture a flag while learning what the other pieces are. A move therefore changes both the board and what someone can infer about it. It is an unusually tidy laboratory for reasoning when information is unevenly distributed.
For historical perspective, DeepMind’s December 2022 DeepNash account reported an all-time top-three ranking on the Gravon platform and expert-level human play. AI did not discover Stratego yesterday. Ataraxos advances an already serious line of work.
The distinction is useful outside gaming. Knowing how to choose when the facts are visible and knowing how to choose when someone else possesses crucial facts are different challenges. Imagine buying used equipment from a seller who knows its maintenance history. You can ask questions, inspect components, walk away, or make an offer. Each action costs something and may reveal something. That is an analogy, not an Ataraxos deployment.
In Massachusetts terms, it resembles a Kendall Square networking conversation in which everyone knows their own runway and nobody quite knows whether “we’re being selective” means profitable or fundraising. The information asymmetry is familiar. The board game at least supplies rules.
Do the homework, then think before moving
MIT’s announcement identifies senior author Gabriele Farina, an EECS assistant professor and principal investigator at the Laboratory for Information and Decision Systems. This is a substantive Cambridge research connection within a multi-university collaboration, not an honorary Boston company assembled from somebody’s alumni biography.
The basic approach combines self-play with planning during the game. Playing against itself builds an initial strategy. A generative model estimates plausible identities for hidden opposing pieces; the system then examines possible continuations before choosing a move. It learns a starting point and does additional work at the moment a decision matters.
That combination appeals to me because it gives uncertainty an actual job in the computation. In an illustrative situation, an approaching unknown piece could represent strength or a bluff. Committing everything to one interpretation would be brittle. Considering several plausible explanations creates room to take a calculated risk. A useful strategist needs a way to handle being unsure, not merely a more confident delivery voice.
Readers of our Subconscious coverage will recognize the broader editorial interest: how much useful work can a machine get from its computation? Context compression and game-playing algorithms solve different problems. Both make a better conversation than treating electricity consumption as a personality.
The trophy comes with a denominator
The Nature evaluation against Niemeijer comprised 20 games over three weeks. Counting draws as half-wins gives an 85% effective win rate. The paper also records 38 wins and two losses in a demonstration at the August 2025 world championship. That was a demo against attendees, not a claim that the system entered and won the human tournament.
The paper reports a Stratego training cost of a few thousand dollars. Its techniques also produced superhuman Barrage Stratego play and state-of-the-art results for Hanabi and dou dizhu. Those distinctions matter: different games, different comparisons, no universal trophy.
A lower training bill is consequential because it can change who gets to experiment. If a research direction requires industrial spending before anyone can test a sensible idea, the set of people able to contribute shrinks. More economical methods offer a path to broader participation. They do not automatically make every implementation cheap, or include the salaries, failed experiments, and institutional support behind a published result.
This is the regional strength I find more interesting than declaring Boston the capital of all future intelligence. Our Julia column explored another local habit: improving the machinery that lets other people do difficult work. There is civic pride available in making the calculation better. It does not require a ribbon-cutting for the calculation.
Please leave the procurement department on human difficulty
MIT describes negotiations and cybersecurity as possible future applications. Farina also says interpretability needs work so humans can audit recommendations. These are prospective uses and an acknowledged research need, not evidence of successful operational deployment.
A game supplies a bounded world, permitted moves, and an agreed outcome. A contract negotiation might contain ambiguous language, changing incentives, damaged relationships, or an email attachment that invalidates everybody’s assumptions. Before applying a strategy engine there, I would want evidence that its model captures the relevant choices and consequences. Winning the simulated version of the wrong problem would be a particularly expensive form of academic excellence.
That is also why our argument for Massachusetts AI ambition should leave room for several kinds of success. A foundation-model business, a scientific-software community, and a strategic-learning paper contribute differently. They need not be forced into matching fleece vests for the ecosystem photograph.
Ataraxos earns an enthusiastic, appropriately bounded cheer. Better play with substantially less training computation is a result worth understanding. The next challenge is making those decision methods useful where the rules are less accommodating. For today, Cambridge can enjoy a machine that did its homework, kept its composure, and left enough room in the budget to buy the board game.