{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,3]],"date-time":"2025-12-03T17:30:36Z","timestamp":1764783036469,"version":"3.46.0"},"reference-count":21,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Meas. Anal. Comput. Syst."],"published-print":{"date-parts":[[2025,12]]},"abstract":"<jats:p>We introduce a new hypothesis testing-based learning dynamics in which players update their strategies by combining hypothesis testing with utility-driven exploration. In this dynamics, each player forms beliefs about opponents' strategies and episodically tests these beliefs using empirical observations. Beliefs are resampled either when the hypothesis test is rejected or through exploration, where the probability of exploration decreases with the player's (transformed) utility. In general finite normal-form games, we show that the learning process converges to a set of approximate Nash equilibria and, more importantly, to a refinement that selects equilibria maximizing the minimum (transformed) utility across all players. Our result establishes convergence to equilibrium in general finite games and reveals a novel mechanism for equilibrium selection induced by the structure of the learning dynamics.<\/jats:p>","DOI":"10.1145\/3771571","type":"journal-article","created":{"date-parts":[[2025,12,2]],"date-time":"2025-12-02T20:07:03Z","timestamp":1764706023000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Learning with Episodic Hypothesis Testing in General Games: A Framework for Equilibrium Selection"],"prefix":"10.1145","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-8397-5929","authenticated-orcid":false,"given":"Ruifan","family":"Yang","sequence":"first","affiliation":[{"name":"Cornell University, Ithaca, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5334-4163","authenticated-orcid":false,"given":"Manxi","family":"Wu","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,12,2]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13235-018-0244-z"},{"volume-title":"Elements of information theory","author":"Cover Thomas M","key":"e_1_2_1_2_1","unstructured":"Thomas M Cover. 1999. Elements of information theory. John Wiley & Sons."},{"key":"e_1_2_1_3_1","first-page":"341","volume-title":"Theoretical Economics","volume":"1","author":"Foster Dean","year":"2006","unstructured":"Dean Foster and H Young. 2006. Regret Testing: Learning to Play Nash Equilibrium Without Knowing You Have an Opponent. Theoretical Economics, Vol. 1 (10 2006), 341-367."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1006\/game.1997.0595"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0899-8256(03)00025-3"},{"volume-title":"The theory of learning in games","author":"Fudenberg Drew","key":"e_1_2_1_6_1","unstructured":"Drew Fudenberg and David K Levine. 1998. The theory of learning in games. Vol. 2. MIT press."},{"key":"e_1_2_1_7_1","volume-title":"On the properties of the softmax function with application in game theory and reinforcement learning. arXiv preprint arXiv:1704.00805","author":"Gao Bolin","year":"2017","unstructured":"Bolin Gao and Lacra Pavel. 2017. On the properties of the softmax function with application in game theory and reinforcement learning. arXiv preprint arXiv:1704.00805 (2017)."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.GEB.2006.06.001"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1111\/1468-0262.00153"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.JET.2022.105551"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-27819-1_3"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.2307\/2951492"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/S10994-006-0219-Y\/METRICS"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCB.2009.2017273"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.GEB.2012.03.006"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.2307\/2171894"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/PL00007181\/METRICS"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1111\/J.1468-0262.2005.00585.X"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.GEB.2012.02.017"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.2307\/2951778"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.GEB.2008.02.011"}],"container-title":["Proceedings of the ACM on Measurement and Analysis of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3771571","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,3]],"date-time":"2025-12-03T17:25:49Z","timestamp":1764782749000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3771571"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12]]},"references-count":21,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,12]]}},"alternative-id":["10.1145\/3771571"],"URL":"https:\/\/doi.org\/10.1145\/3771571","relation":{},"ISSN":["2476-1249"],"issn-type":[{"type":"electronic","value":"2476-1249"}],"subject":[],"published":{"date-parts":[[2025,12]]},"assertion":[{"value":"2025-12-02","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}