{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T01:56:35Z","timestamp":1790819795904,"version":"4.1.0"},"reference-count":55,"publisher":"Springer Science and Business Media LLC","issue":"8134","license":[{"start":{"date-parts":[[2026,9,30]],"date-time":"2026-09-30T00:00:00Z","timestamp":1790726400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,9,30]],"date-time":"2026-09-30T00:00:00Z","timestamp":1790726400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Nature"],"published-print":{"date-parts":[[2026,10,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another. Unfortunately, the presence of large amounts of hidden information renders established reinforcement learning and search approaches ineffective. Even with multimillion-dollar industrial research efforts\n                    <jats:sup>1<\/jats:sup>\n                    , top-human-level play at Stratego\u2014a board wargame with hidden information on a massive scale\u2014has remained beyond the reach of artificial intelligence (AI). Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information. Ataraxos defeated the most decorated human Stratego player of all time by a large margin\u2014achieving, to our knowledge, the first superhuman result in the game\u2019s history\u2014while consuming orders of magnitude less compute and data than previous efforts. Using the same techniques, we built a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and\n                    <jats:italic>dou dizhu<\/jats:italic>\n                    , all with low cost and high sample efficiency. The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.\n                  <\/jats:p>","DOI":"10.1038\/s41586-026-11036-y","type":"journal-article","created":{"date-parts":[[2026,9,30]],"date-time":"2026-09-30T15:04:02Z","timestamp":1790780642000},"page":"55-59","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Scalable decision-making for games of imperfect information"],"prefix":"10.1038","volume":"658","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-8437-5368","authenticated-orcid":false,"given":"Samuel","family":"Sokota","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eugene","family":"Vinitsky","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hengyuan","family":"Hu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhiyuan","family":"Fan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"J. Zico","family":"Kolter","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3976-0061","authenticated-orcid":false,"given":"Gabriele","family":"Farina","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,9,30]]},"reference":[{"key":"11036_CR1","doi-asserted-by":"publisher","first-page":"990","DOI":"10.1126\/science.add4679","volume":"378","author":"J Perolat","year":"2022","unstructured":"Perolat, J. et al. Mastering the game of Stratego with model-free multiagent reinforcement learning. Science 378, 990\u2013996 (2022).","journal-title":"Science"},{"key":"11036_CR2","unstructured":"von Neumann, J. & Morgenstern, O. Theory of Games and Economic Behavior (Princeton Univ. Press, 1944)."},{"key":"11036_CR3","doi-asserted-by":"publisher","unstructured":"Kuhn, H. W. Extensive games and the problem of information. In Contributions to the Theory of Games Vol. 2 (eds Kuhn, H. W. & Tucker, A. W.) 193\u2013216 https:\/\/doi.org\/10.1515\/9781400881970-012 (Princeton Univ. Press, 1953).","DOI":"10.1515\/9781400881970-012"},{"key":"11036_CR4","doi-asserted-by":"publisher","first-page":"508","DOI":"10.1126\/science.aam6960","volume":"356","author":"M Morav\u010d\u00edk","year":"2017","unstructured":"Morav\u010d\u00edk, M. et al. DeepStack: expert-level artificial intelligence in heads-up no-limit poker. Science 356, 508\u2013513 (2017).","journal-title":"Science"},{"key":"11036_CR5","unstructured":"Brown, N. & Sandholm, T. Safe and nested subgame solving for imperfect-information games. Adv. Neural Inf. Process. Syst. 30, 689\u2013699 (2017)."},{"key":"11036_CR6","doi-asserted-by":"publisher","first-page":"418","DOI":"10.1126\/science.aao1733","volume":"359","author":"N Brown","year":"2018","unstructured":"Brown, N. & Sandholm, T. Superhuman AI for heads-up no-limit poker: Libratus beats top professionals. Science 359, 418\u2013424 (2018).","journal-title":"Science"},{"key":"11036_CR7","unstructured":"Brown, N., Sandholm, T. & Amos, B. Depth-limited solving for imperfect-information games. Adv. Neural Inf. Process. Syst. 31, 7663\u20137674 (2018)."},{"key":"11036_CR8","doi-asserted-by":"publisher","first-page":"885","DOI":"10.1126\/science.aay2400","volume":"365","author":"N Brown","year":"2019","unstructured":"Brown, N. & Sandholm, T. Superhuman AI for multiplayer poker. Science 365, 885\u2013890 (2019).","journal-title":"Science"},{"key":"11036_CR9","doi-asserted-by":"publisher","unstructured":"Zarick, R., Pellegrino, B., Brown, N. & Banister, C. Unlocking the potential of deep counterfactual value networks. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2007.10442 (2020).","DOI":"10.48550\/arXiv.2007.10442"},{"key":"11036_CR10","unstructured":"Brown, N., Bakhtin, A., Lerer, A. & Gong, Q. Combining deep reinforcement learning and search for imperfect-information games. Adv. Neural Inf. Process. Syst. 33, 17057\u201317069 (2020)."},{"key":"11036_CR11","unstructured":"Sokota, S. et al. Abstracting imperfect information away from two-player zero-sum games. In International Conference on Machine Learning (eds Krause, A. et al.) 32169\u201332193 (PMLR, 2023)."},{"key":"11036_CR12","unstructured":"Sokota, S. et al. A unified approach to reinforcement learning, quantal response equilibria, and two-player zero-sum games. In International Conference on Learning Representations (eds Liu, Y. et al.) 18182\u201318194 (ICLR, 2023)."},{"key":"11036_CR13","unstructured":"Sokota, S. et al. The update-equivalence framework for decision-time planning. In International Conference on Learning Representations (eds Kim, B. et al.) 25831\u201325847 (ICLR, 2024)."},{"key":"11036_CR14","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1145\/1365490.1365500","volume":"6","author":"J Nickolls","year":"2008","unstructured":"Nickolls, J., Buck, I., Garland, M. & Skadron, K. Scalable parallel programming with CUDA. Queue 6, 40\u201353 https:\/\/doi.org\/10.1145\/1365490.1365500 (2008).","journal-title":"Queue"},{"key":"11036_CR15","unstructured":"International Stratego Rating. History of Pim Niemeijer. Kleier.net https:\/\/web.archive.org\/web\/20260910042337\/https:\/\/www.kleier.net\/cgi\/player.php?pid=2047 (2026)."},{"key":"11036_CR16","unstructured":"Stratego. Wikipedia https:\/\/en.wikipedia.org\/w\/index.php?title=Stratego&oldid=1362636676 (2026)."},{"key":"11036_CR17","unstructured":"International Stratego Rating. History of George Franka. Kleier.net https:\/\/web.archive.org\/web\/20260910042436\/https:\/\/www.kleier.net\/cgi\/player.php?pid=776 (2026)."},{"key":"11036_CR18","unstructured":"International Stratego Rating. History of Max Roelofs. Kleier.net https:\/\/web.archive.org\/web\/20260910042604\/https:\/\/www.kleier.net\/cgi\/player.php?pid=2406 (2026)."},{"key":"11036_CR19","unstructured":"International Stratego Rating. History of Vincent de Boer Kleier.net https:\/\/web.archive.org\/web\/20260910042516\/https:\/\/www.kleier.net\/cgi\/player.php?pid=296 (2026)."},{"key":"11036_CR20","doi-asserted-by":"publisher","unstructured":"Bard, N. et al. The Hanabi challenge: a new frontier for AI research. Artif. Intell. https:\/\/doi.org\/10.1016\/j.artint.2019.103216 (2020).","DOI":"10.1016\/j.artint.2019.103216"},{"key":"11036_CR21","unstructured":"Dou dizhu. Wikipedia https:\/\/en.wikipedia.org\/w\/index.php?title=Dou_dizhu&oldid=1339210829 (2026)."},{"key":"11036_CR22","doi-asserted-by":"publisher","first-page":"141","DOI":"10.3390\/info11030141","volume":"11","author":"G Yuexian","year":"2020","unstructured":"Yuexian, G., Li, W., Xiao, Y., Khalid, M. N. A. & Iida, H. Nature of attractive multiplayer games: case study on China\u2019s most popular card game\u2014doudizhu. Information 11, 141 https:\/\/doi.org\/10.3390\/info11030141 (2020).","journal-title":"Information"},{"key":"11036_CR23","doi-asserted-by":"publisher","first-page":"34954","DOI":"10.52202\/068431-2533","volume":"35","author":"G Yang","year":"2022","unstructured":"Yang, G. et al. PerfectDou: dominating doudizhu with perfect information distillation. Adv. Neural Inf. Process. Syst. 35, 34954\u201334965 (2022).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"11036_CR24","unstructured":"Zha, D. et al. DouZero: mastering doudizhu with self-play deep reinforcement learning. In International Conference on Machine Learning (ICML) 12333\u201312344 (PMLR, 2021)."},{"key":"11036_CR25","doi-asserted-by":"publisher","first-page":"58","DOI":"10.1145\/203330.203343","volume":"38","author":"G Tesauro","year":"1995","unstructured":"Tesauro, G. Temporal difference learning and TD-Gammon. Commun. ACM 38, 58\u201368 https:\/\/doi.org\/10.1145\/203330.203343 (1995).","journal-title":"Commun. ACM"},{"key":"11036_CR26","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver, D. et al. Mastering the game of Go with deep neural networks and tree search. Nature 529, 484\u2013489 https:\/\/doi.org\/10.1038\/nature16961 (2016).","journal-title":"Nature"},{"key":"11036_CR27","doi-asserted-by":"publisher","first-page":"354","DOI":"10.1038\/nature24270","volume":"550","author":"D Silver","year":"2017","unstructured":"Silver, D. et al. Mastering the game of Go without human knowledge. Nature 550, 354\u2013359 https:\/\/doi.org\/10.1038\/nature24270 (2017).","journal-title":"Nature"},{"key":"11036_CR28","doi-asserted-by":"publisher","first-page":"604","DOI":"10.1038\/s41586-020-03051-4","volume":"588","author":"J Schrittwieser","year":"2020","unstructured":"Schrittwieser, J. et al. Mastering Atari, Go, chess and shogi by planning with a learned model. Nature 588, 604\u2013609 https:\/\/doi.org\/10.1038\/s41586-020-03051-4 (2020).","journal-title":"Nature"},{"key":"11036_CR29","doi-asserted-by":"publisher","unstructured":"Wu, D. J. Accelerating self-play learning in Go. Preprint at https:\/\/doi.org\/10.48550\/arXiv.1902.10565 (2019).","DOI":"10.48550\/arXiv.1902.10565"},{"key":"11036_CR30","unstructured":"Sutton, R. S. The bitter lesson. Incomplete Ideas http:\/\/www.incompleteideas.net\/IncIdeas\/BitterLesson.html (2019)."},{"key":"11036_CR31","unstructured":"Bakhtin, A. et al. Mastering the game of No-Press Diplomacy via human-regularized reinforcement learning and planning. In International Conference on Learning Representations (eds Liu, Y. et al.) 22428\u201322456 (ICLR, 2023)."},{"key":"11036_CR32","unstructured":"Vaswani, A. et al. Attention is all you need. Adv. Neural Inf. Process. Syst. 30, 5998\u20136008 (2017)."},{"key":"11036_CR33","doi-asserted-by":"publisher","unstructured":"Monroe, D. & Chalmers, P. A. Mastering chess with a transformer model. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2409.12272 (2024).","DOI":"10.48550\/arXiv.2409.12272"},{"key":"11036_CR34","unstructured":"Gehring, J., Auli, M., Grangier, D., Yarats, D. & Dauphin, Y. N. Convolutional sequence to sequence learning. In International Conference on Machine Learning (eds Precup, D. & Teh, Y. W.) 1243\u20131252 (PMLR, 2017)."},{"key":"11036_CR35","doi-asserted-by":"crossref","unstructured":"Sutton, R. S. et al. Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction. In International Conference on Autonomous Agents and Multiagent Systems (eds Tumer, K. et al.) 761\u2013768 (IFAAMAS, 2011).","DOI":"10.65109\/QDHN7183"},{"key":"11036_CR36","doi-asserted-by":"publisher","first-page":"9","DOI":"10.1023\/A:1022633531479","volume":"3","author":"RS Sutton","year":"1988","unstructured":"Sutton, R. S. Learning to predict by the methods of temporal differences. Mach. Learn. 3, 9\u201344 (1988).","journal-title":"Mach. Learn."},{"key":"11036_CR37","unstructured":"Schulman, J., Moritz, P., Levine, S., Jordan, M. & Abbeel, P. High-dimensional continuous control using generalized advantage estimation. In International Conference on Learning Representations (eds Bengio, Y. & LeCun, Y.) (ICLR, 2016)."},{"key":"11036_CR38","unstructured":"Cusumano-Towner, M. F. et al. Robust autonomy emerges from self-play. In International Conference on Machine Learning (eds Singh, A. et al.) 11710\u201311737 (PMLR, 2025)."},{"key":"11036_CR39","doi-asserted-by":"publisher","first-page":"633","DOI":"10.1038\/s41586-025-09422-z","volume":"645","author":"D Guo","year":"2025","unstructured":"Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633\u2013638 (2025).","journal-title":"Nature"},{"key":"11036_CR40","unstructured":"Ziebart, B. D., Maas, A., Bagnell, J. A. & Dey, A. K. Maximum entropy inverse reinforcement learning. In AAAI Conference on Artificial Intelligence (eds Fox, D. & Gomes, C. P.) 1433\u20131438 (AAAI Press, 2008)."},{"key":"11036_CR41","doi-asserted-by":"publisher","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A. & Klimov, O. Proximal policy optimization algorithms. Preprint at https:\/\/doi.org\/10.48550\/arXiv.1707.06347 (2017).","DOI":"10.48550\/arXiv.1707.06347"},{"key":"11036_CR42","unstructured":"Pascanu, R., Mikolov, T. & Bengio, Y. On the difficulty of training recurrent neural networks. In International Conference on Machine Learning (eds Dasgupta, S. & McAllester, D.) 1310\u20131318 (PMLR, 2013)."},{"key":"11036_CR43","unstructured":"Kingma, D. P. & Ba, J. Adam: a method for stochastic optimization. In International Conference on Learning Representations (eds Bengio, Y. & LeCun, Y.) (ICLR, 2015)."},{"key":"11036_CR44","first-page":"1929","volume":"15","author":"N Srivastava","year":"2014","unstructured":"Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I. & Salakhutdinov, R. Dropout: a simple way to prevent neural networks from overfitting. J. Mach. Learn. Res. 15, 1929\u20131958 (2014).","journal-title":"J. Mach. Learn. Res."},{"key":"11036_CR45","unstructured":"Moore, S. Stratego evaluator. GitHub https:\/\/github.com\/braathwaate\/strategoevaluator (2012)."},{"key":"11036_CR46","unstructured":"de Boer, V. Invincible: A Stratego Bot. Master\u2019s thesis, Delft Univ. Technology (2007)."},{"key":"11036_CR47","unstructured":"Stratego Classic Challenge rating\/ranking 2022. Gravon https:\/\/www.gravon.de\/gravon\/stratego\/rating2022.jsp (2026)."},{"key":"11036_CR48","unstructured":"Statistics for starship. Gravon https:\/\/www.gravon.de\/gravon\/stratego\/player0.jsp?nick=starship (2026)."},{"key":"11036_CR49","unstructured":"DeepNash surprises top Stratego players. Stratego News https:\/\/web.archive.org\/web\/20250119231008\/http:\/\/www.strategonews.com\/wc2023\/deepnash-surprises-top-stratego-players\/ (2023)."},{"key":"11036_CR50","unstructured":"Cloud TPU pricing. Google Cloud https:\/\/cloud.google.com\/tpu\/pricing?hl=en (2025)."},{"key":"11036_CR51","unstructured":"Voltage Park pricing. Voltage Park https:\/\/www.voltagepark.com\/pricing (2025)."},{"key":"11036_CR52","doi-asserted-by":"publisher","unstructured":"Xu, M. et al. Spatial-temporal transformer networks for traffic flow forecasting. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2001.02908 (2021).","DOI":"10.48550\/arXiv.2001.02908"},{"key":"11036_CR53","unstructured":"Zhang, B. H. & Sandholm, T. General search techniques without common knowledge for imperfect-information games, and application to superhuman Fog of War chess. In International Conference on Learning Representations (eds Vondrick, C. et al.) 93536\u201393563 (ICLR, 2026)."},{"key":"11036_CR54","unstructured":"Fickinger, A., Hu, H., Amos, B., Russell, S. & Brown, N. Scalable online planning via reinforcement learning fine-tuning. In Advances in Neural Information Processing Systems Vol. 34 (eds Ranzato, M. et al.) 16951\u201316963 (Curran Associates, 2021)."},{"key":"11036_CR55","doi-asserted-by":"publisher","unstructured":"Forkel, J., Ruhdorfer, C., Beukman, M., Bulling, A. & Foerster, J. High entropy leads to symmetry equivariant policies in Dec-POMDPs. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2511.22581 (2026).","DOI":"10.48550\/arXiv.2511.22581"}],"container-title":["Nature"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s41586-026-11036-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41586-026-11036-y","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41586-026-11036-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T01:04:30Z","timestamp":1790816670000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s41586-026-11036-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,9,30]]},"references-count":55,"journal-issue":{"issue":"8134","published-print":{"date-parts":[[2026,10,1]]}},"alternative-id":["11036"],"URL":"https:\/\/doi.org\/10.1038\/s41586-026-11036-y","relation":{},"ISSN":["0028-0836","1476-4687"],"issn-type":[{"value":"0028-0836","type":"print"},{"value":"1476-4687","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,9,30]]},"assertion":[{"value":"20 November 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 August 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 September 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","label":"Competing interests","group":{"name":"EthicsHeading","label":"Ethics"}}]}}