Deep Blue beat Garry Kasparov in chess in 1997, AlphaGo beat Lee Sedol in Go in 2016, and poker bots have been beating pros for years. But a classic game called A strategist endured. Even DeepMind, with its extraordinary budget, was unable to build a machine that would reliably beat the best human players.
Now a team of researchers from Carnegie Mellon, MIT, New York University and Stanford University have done it. Their AI named Ataraxos beat Pim Niemeijer, arguably the best A strategist Players all-time, 15 to 1, with four ties. And it only took 16 GPUs and a few thousand dollars to train it.
Hidden armies
In A strategistEach player receives 40 pieces representing military ranks, from marshal to spy, as well as bombs and a flag. You win by capturing the opponent’s flag. Your opponent knows it Where Your pieces are, but not What they are. Identities are only revealed when two parts collide in battle – the weaker one is removed and the victor’s identity is revealed. That makes A strategist a game with incomplete information, just like poker, which computers cracked years ago. “There is something very special there A strategist“That means it’s a huge amount of hidden information unfolding over a very long period of time,” said Eugene Vinitsky, a researcher at NYU and co-author of the study.
In some forms of poker, the hidden information is tiny. In Texas Hold’em, “you only have two hidden cards,” said Gabriele Farina, a computer scientist at MIT and another co-author. That leaves only 1,326 possible hands, few enough for a machine to weigh them all. “In A strategist“There are 40 pieces on the board that can be in any order,” Farina said. That’s more than ten million possible lineups. Then there is the length of the game.
“In chess the game usually lasts 40 moves, but in A strategist“A game can easily last 2,000 moves,” Farina said. Over and beyond A strategist is a bluff game. Sometimes you move a weak piece like a marshal just to scare off your opponent. If players bluff too often, their threats have no meaning; If they never bluff, they become predictable. This balancing act, the team explains, is what caused previous AIs like DeepMind’s DeepNash, which launched in 2022, to fail.