{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:32:34Z","timestamp":1740123154313,"version":"3.37.3"},"reference-count":28,"publisher":"Springer Science and Business Media LLC","issue":"8","license":[{"start":{"date-parts":[[2022,6,9]],"date-time":"2022-06-09T00:00:00Z","timestamp":1654732800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,6,9]],"date-time":"2022-06-09T00:00:00Z","timestamp":1654732800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100006919","name":"Massachusetts Institute of Technology","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100006919","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2022,8]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We address the problem of interpretability in iterative game solving for imperfect-information games such as poker. This lack of interpretability has two main sources: first, the use of an uninterpretable feature representation, and second, the use of black box methods such as neural networks, for the fitting procedure. In this paper, we present advances on both fronts. Namely, first we propose a novel, compact, and easy-to-understand game-state feature representation for Heads-up No-limit (HUNL) Poker. Second, we make use of globally optimal decision trees, paired with a counterfactual regret minimization (CFR) self-play algorithm, to train our poker bot which produces an entirely interpretable agent. Through experiments against Slumbot, the winner of the most recent Annual Computer Poker Competition, we demonstrate that our approach yields a HUNL Poker agent that is capable of beating the Slumbot. Most exciting of all, the resulting poker bot is highly interpretable, allowing humans to learn from the novel strategies it discovers.<\/jats:p>","DOI":"10.1007\/s10994-022-06179-8","type":"journal-article","created":{"date-parts":[[2022,6,9]],"date-time":"2022-06-09T23:03:36Z","timestamp":1654815816000},"page":"3063-3083","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["World-class interpretable poker"],"prefix":"10.1007","volume":"111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1985-1003","authenticated-orcid":false,"given":"Dimitris","family":"Bertsimas","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alex","family":"Paskov","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,6,9]]},"reference":[{"key":"6179_CR1","volume-title":"Machine learning under a modern optimization lens","author":"D Bertsimas","year":"2019","unstructured":"Bertsimas, D., & Dunn, J. (2019). Machine learning under a modern optimization lens. Dynamic Ideas LLC."},{"key":"6179_CR2","volume-title":"Classification and regression trees","author":"L Breiman","year":"1984","unstructured":"Breiman, L., Friedman, J. H., Olshen, R. A., & Stone, C. J. (1984). Classification and regression trees. Wadsworth and Brooks."},{"key":"6179_CR3","unstructured":"Brown, N., Ganzfried, S., & Sandholm, T. (2015). Hierarchical abstraction, distributed equilibrium computation, and post-processing, with application to a champion no-limit Texas Hold\u2019em agent (Vol. 1, pp. 7\u201315)."},{"key":"6179_CR4","unstructured":"Brown, N., & Sandholm, T. (2015). Simultaneous abstraction and equilibrium finding in games. In IJCAI."},{"key":"6179_CR5","unstructured":"Brown, N., Sandholm, T., & Amos, B. (2018). Depth-limited solving for imperfect-information games. In Advances in neural information processing systems (Vol.\u00a031, pp. 7663\u20137674). Curran Associates, Inc."},{"key":"6179_CR6","first-page":"793","volume":"97","author":"N Brown","year":"2019","unstructured":"Brown, N., Lerer, A., Gross, S., & Sandholm, T. (2019). Deep counterfactual regret minimization. Conference on Machine Learning, 97, 793\u2013802.","journal-title":"Conference on Machine Learning"},{"issue":"6374","key":"6179_CR7","doi-asserted-by":"publisher","first-page":"418","DOI":"10.1126\/science.aao1733","volume":"359","author":"N Brown","year":"2018","unstructured":"Brown, N., & Sandholm, T. (2018). Superhuman AI for heads-up no-limit poker: Libratus beats top professionals. Science, 359(6374), 418\u2013424.","journal-title":"Science"},{"key":"6179_CR8","doi-asserted-by":"publisher","first-page":"1829","DOI":"10.1609\/aaai.v33i01.33011829","volume":"33","author":"N Brown","year":"2019","unstructured":"Brown, N., & Sandholm, T. (2019). Solving imperfect-information games via discounted regret minimization. Proceedings of the AAAI Conference on Artificial Intelligence, 33, 1829\u20131836.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"issue":"7","key":"6179_CR9","doi-asserted-by":"publisher","first-page":"1039","DOI":"10.1007\/s10994-017-5633-9","volume":"106","author":"DBJ Dunn","year":"2017","unstructured":"Dunn, D. B. J. (2017). Optimal classification trees. Machine Learning, 106(7), 1039\u20131082.","journal-title":"Machine Learning"},{"key":"6179_CR10","unstructured":"Ganzfried, S., & Chiswick, M. (2019). Most important fundamental rule of poker strategy. CoRR, arxiv:1906.09895."},{"key":"6179_CR11","unstructured":"Ganzfried, S., & Sandholm, T. (2010). Computing equilibria by incorporating qualitative models. AAMAS."},{"key":"6179_CR12","unstructured":"Ganzfried, S., & Sandholm, T. (2013). Action translation in extensive-form games with large action spaces: axioms, paradoxes, and the pseudo-harmonic mapping. In IJCAI."},{"key":"6179_CR13","doi-asserted-by":"crossref","unstructured":"Ganzfried, S., & Sandholm, T. (2014). Potential-aware imperfect-recall abstraction with earth mover\u2019s distance in imperfect-information games. In Proceedings of the twenty-eighth AAAI conference on artificial intelligence (pp. 682\u2013690). AAAI Press.","DOI":"10.1609\/aaai.v28i1.8816"},{"issue":"4","key":"6179_CR14","doi-asserted-by":"publisher","first-page":"2017","DOI":"10.3390\/g8040049","volume":"8","author":"S Ganzfried","year":"2017","unstructured":"Ganzfried, S., & Yusuf, F. (2017). Computing human-understandable strategies: deducing fundamental rules of poker strategy. Games, 8(4), 2017. https:\/\/doi.org\/10.3390\/g8040049.","journal-title":"Games"},{"key":"6179_CR15","doi-asserted-by":"crossref","unstructured":"Gilpin, A., & Sandholm, T. (2006). A competitive Texas Hold\u2019em poker player via automated abstraction and real-time equilibrium computation. In Proceedings of the 21st national conference on artificial intelligence (Vol.\u00a02).","DOI":"10.1145\/1160633.1160911"},{"key":"6179_CR16","doi-asserted-by":"crossref","unstructured":"Gilpin, A., & Sandholm, T. (2007a). Better automated abstraction techniques for imperfect information games, with application to Texas Hold\u2019em poker. In AAMAS \u201907.","DOI":"10.1145\/1329125.1329358"},{"key":"6179_CR17","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1145\/1284320.1284324","volume":"54","author":"A Gilpin","year":"2007","unstructured":"Gilpin, A., & Sandholm, T. (2007b). Lossless abstraction of imperfect information games. J. ACM, 54, 25.","journal-title":"J. ACM"},{"key":"6179_CR18","unstructured":"Gilpin, A., & Sandholm, T. (2008). An experimental comparison using poker: Expectation-based versus potential-aware automated abstraction in imperfect information games. In AAAI."},{"key":"6179_CR19","unstructured":"Gilpin, A., Sandholm, T., & S\u00f8rensen, T. (2007). Potential-aware automated abstraction of sequential games, and holistic equilibrium analysis of Texas Hold\u2019em poker. In Proceedings of the 22nd national conference on artificial intelligence - Volume 1 (Vol.\u00a01, pp. 50\u201357)."},{"key":"6179_CR20","doi-asserted-by":"publisher","unstructured":"Gilpin, A., Sandholm, T., & S\u00f8rensen, T. (2008). A heads-up no-limit Texas Hold\u2019em poker player: Discretized betting models and automatically generated equilibrium-finding programs (Vol.\u00a02, pp. 911\u2013918). https:\/\/doi.org\/10.1145\/1402298.1402350.","DOI":"10.1145\/1402298.1402350"},{"key":"6179_CR21","unstructured":"Jackson, E.\u00a0G. (2017). Targeted cfr. In AAAI workshops."},{"key":"6179_CR22","unstructured":"Johanson, M. (2013). Measuring the size of large no-limit poker games. In CoRR."},{"key":"6179_CR23","unstructured":"Johanson, M., Waugh, K., Bowling, M., & Zinkevich, M. (2011). Accelerating best response calculation in large extensive games (pp. 258\u2013265)."},{"key":"6179_CR24","unstructured":"Li, H., Hu, K., Ge, Z., Jiang, T., Qi, Y., & Song, L. (2020). Double neural counterfactual regret minimization. In International conference on learning representations."},{"issue":"6337","key":"6179_CR25","doi-asserted-by":"publisher","first-page":"508","DOI":"10.1126\/science.aam6960","volume":"356","author":"M Morav\u010d\u0131k","year":"2017","unstructured":"Morav\u010d\u0131k, M., Schmid, M., Burch, N., Lis\u00fd, V., Morrill, D., Bard, N., et al. (2017). Deepstack: Expert-level artificial intelligence in heads-up no-limit poker. Science, 356(6337), 508\u2013513.","journal-title":"Science"},{"key":"6179_CR26","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1126\/science.aar6404","volume":"362","author":"D Silver","year":"2018","unstructured":"Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., et al. (2018). A general reinforcement learning algorithm that masters chess, shogi, and go through self-play. Science, 362, 1140\u20131144.","journal-title":"Science"},{"key":"6179_CR27","unstructured":"Zarick, R., Pellegrino, B., Brown, N., & Banister, C. (2020). Unlocking the potential of deep counterfactual value networks. In CoRR, arXiv:abs\/2007.10442."},{"key":"6179_CR28","unstructured":"Zinkevich, M., Johanson, M., Bowling, M., & Piccione, C. (2008). Regret minimization in games with incomplete information. In Advances in neural information processing systems (Vol.\u00a020, pp. 1729\u20131736). Curran Associates, Inc."}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06179-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-022-06179-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06179-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,8,1]],"date-time":"2022-08-01T20:11:42Z","timestamp":1659384702000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-022-06179-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,9]]},"references-count":28,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2022,8]]}},"alternative-id":["6179"],"URL":"https:\/\/doi.org\/10.1007\/s10994-022-06179-8","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"type":"print","value":"0885-6125"},{"type":"electronic","value":"1573-0565"}],"subject":[],"published":{"date-parts":[[2022,6,9]]},"assertion":[{"value":"2 February 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 February 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 April 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 June 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Both authors consent to participate in this research","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"Both authors consent to publications","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"Not applicable","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval"}}]}}