Invited talk: New Results on Real-Time Reasoning in Imperfect-Information Games -- Tuomas Sandholm
Abstract
There has been tremendous progress in 2-player 0-sum games over the last 23 years. The greatest scalability improvements have come from real-time reasoning (aka. subgame-solving) techniques. Subgame solving is drastically more complicated in imperfect-information games because strategies outside a subgame affect what the strategies in the subgame should be. In the first part of this talk, I will present Obscuro, the first superhuman AI for Fog-of-War chess [Zhang & Sandholm, ICLR-26], a recognized challenge problem after superhuman level was reached in no-limit Texas hold’em. Most prior subgame-solving techniques require the construction of the “common knowledge set”, making them unusable with this much imperfect information. Our new techniques do not require that. Experiments against the prior state-of-the-art AI and human players - including the world’s best - show that Obscuro is significantly stronger. In the second part of the talk, I will present two very recent subgame-solving discoveries: 1) a new equilibrium refinement for subgame solving that maintains the safety guarantee of prior techniques while performing significantly better in practice [Kubíček, Lisý & Sandholm, IJCAI-26], and 2) how to fix modern policy-gradient algorithms, which converge to equilibrium when applied to a full game but can produce highly exploitable strategies when further trained locally at run time, by extending safe subgame solving based on gadget games to RL [Kubíček, Lisý & Sandholm, MALGAI-26].