{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T15:53:42Z","timestamp":1783698822076,"version":"3.55.0"},"reference-count":84,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,3,7]],"date-time":"2024-03-07T00:00:00Z","timestamp":1709769600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Science Foundation","award":["IIS 17-51278"],"award-info":[{"award-number":["IIS 17-51278"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Recomm. Syst."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>\n            Current practice for evaluating recommender systems typically focuses on point estimates of user-oriented effectiveness metrics or business metrics, sometimes combined with additional metrics for considerations such as diversity and novelty. In this article, we argue for the need for researchers and practitioners to attend more closely to various\n            <jats:italic>distributions<\/jats:italic>\n            that arise from a recommender system (or other information access system) and the sources of uncertainty that lead to these distributions. One immediate implication of our argument is that both researchers and practitioners must report and examine more thoroughly the distribution of utility between and within different stakeholder groups. However, distributions of various forms arise in many more aspects of the recommender systems experimental process, and distributional thinking has substantial ramifications for how we design, evaluate, and present recommender systems evaluation and research results. Leveraging and emphasizing distributions in the evaluation of recommender systems is a necessary step to ensure that the systems provide appropriate and equitably distributed benefit to the people they affect.\n          <\/jats:p>","DOI":"10.1145\/3613455","type":"journal-article","created":{"date-parts":[[2023,8,5]],"date-time":"2023-08-05T08:53:45Z","timestamp":1691225625000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["Distributionally-Informed Recommender System Evaluation"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2467-0108","authenticated-orcid":false,"given":"Michael D.","family":"Ekstrand","sequence":"first","affiliation":[{"name":"People &amp; Information Research Team, Boise State University, Boise, USA and Department of Information Science, Drexel University, Philadelphia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9538-047X","authenticated-orcid":false,"given":"Ben","family":"Carterette","sequence":"additional","affiliation":[{"name":"Spotify, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2345-1288","authenticated-orcid":false,"given":"Fernando","family":"Diaz","sequence":"additional","affiliation":[{"name":"Language Technologies Institute, Carnegie Mellon University, Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,3,7]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079628.3079657"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/2382438.2382442"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","unstructured":"J. J. Allaire Charles Teague Carlos Scheidegger Yihui Xie and Christophe Dervieux. 2022. Quarto. DOI:10.5281\/zenodo.5960048","DOI":"10.5281\/zenodo.5960048"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00008"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3527546.3527559"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401033"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210063"},{"key":"e_1_3_3_9_2","first-page":"43","volume-title":"Proceedings of the 14th Conference on Uncertainty in Artificial Intelligence","author":"Breese John S.","year":"1998","unstructured":"John S. Breese, David Heckerman, and Carl Kadie. 1998. Empirical analysis of predictive algorithms for collaborative filtering. In Proceedings of the 14th Conference on Uncertainty in Artificial Intelligence. 43\u201352."},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/792550.792552"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080836"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210014"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.18637\/jss.v076.i01"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/1571941.1572017"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/2009916.2010037"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/2808194.2809469"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3341981.3358959"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-78646-7_5"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/2063576.2063668"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/2396761.2396782"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240323.3240370"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-011-9167-7"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/1645953.1646033"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526711"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324922000043"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390334.1390446"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1108\/eb050097"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/963770.963776"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3411962"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3412778"},{"key":"e_1_3_3_31_2","unstructured":"Michael D. Ekstrand. 2021. Multiversal Simulacra: Understanding Hypotheticals and Possible Worlds Through Simulation. arXiv:2110.00811 (Oct.2021). https:\/\/arxiv.org\/abs\/2110.00811"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460231.3470938"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1561\/1500000079"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11257-020-09284-2"},{"key":"e_1_3_3_35_2","volume-title":"Proceedings of the 30th Florida Artificial Intelligence Research Society Conference","author":"Ekstrand Michael D.","year":"2017","unstructured":"Michael D. Ekstrand and Vaibhav Mahant. 2017. Sturgeon and the cool kids: Problems with Top-N recommender evaluation. In Proceedings of the 30th Florida Artificial Intelligence Research Society Conference. AAAI Press."},{"key":"e_1_3_3_36_2","first-page":"172","volume-title":"Proceedings of the 1st Conference on Fairness, Accountability and Transparency","volume":"81","author":"Ekstrand Michael D.","year":"2018","unstructured":"Michael D. Ekstrand, Mucun Tian, Ion Madrazo Azpiazu, Jennifer D. Ekstrand, Oghenemaro Anuyah, David McNeill, and Maria Soledad Pera. 2018. All the cool kids, how do they fit in?: Popularity and demographic biases in Recommender Evaluation and Effectiveness. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency. Sorelle A. Friedler and Christo Wilson (Eds.), Vol. 81. PMLR, 172\u2013186."},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3406522.3446033"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3190580.3190586"},{"key":"e_1_3_3_39_2","unstructured":"Google. 2022. Search Quality Evaluator Guidelines. https:\/\/static.googleusercontent.com\/media\/guidelines.raterhub.com\/en\/\/searchqualityevaluatorguidelines.pdf"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-0716-2197-4_15"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/963770.963772"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2008.22"},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-021-05946-3"},{"key":"e_1_3_3_44_2","volume-title":"Proceedings of the Perspectives on the Evaluation of Recommender Systems Workshop 2021","volume":"2955","author":"Ihemelandu Ngozi","year":"2021","unstructured":"Ngozi Ihemelandu and Michael D. Ekstrand. 2021. Statistical inference: The missing piece of recsys experiment reliability discourse. In Proceedings of the Perspectives on the Evaluation of Recommender Systems Workshop 2021, Vol. 2955. CEUR-WS."},{"key":"e_1_3_3_45_2","volume-title":"Proceedings of the ImpactRS Workshop at RecSys 2019","author":"Jannach Dietmar","year":"2019","unstructured":"Dietmar Jannach, Omer Sar Shalem, and Joseph A. Konstan. 2019. Towards more impactful recommender systems research. In Proceedings of the ImpactRS Workshop at RecSys 2019."},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/582415.582418"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/1458082.1458176"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383313.3412235"},{"key":"e_1_3_3_49_2","volume-title":"Proceedings of the ML Evaluation Standards Workshop at ICLR 2022","author":"Larson Stefan","year":"2022","unstructured":"Stefan Larson. 2022. Towards yet another checklist for new datasets. In Proceedings of the ML Evaluation Standards Workshop at ICLR 2022."},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/1835449.1835486"},{"key":"e_1_3_3_51_2","first-page":"50","volume-title":"Proceedings of the 23rd Conference on Uncertainty in Artificial Intelligence","author":"Marlin Benjamin M.","year":"2007","unstructured":"Benjamin M. Marlin, Richard S. Zemel, Sam Roweis, and Malcolm Slaney. 2007. Collaborative filtering and the missing at random assumption. In Proceedings of the 23rd Conference on Uncertainty in Artificial Intelligence. AUAI, 50\u201354."},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460231.3474259"},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3041021.3054197"},{"key":"e_1_3_3_54_2","doi-asserted-by":"crossref","unstructured":"Martin Mladenov Chih-Wei Hsu Vihan Jain Eugene Ie Christopher Colby Nicolas Mayoraz Hubert Pham Dustin Tran Ivan Vendrov and Craig Boutilier. 2021. RecSim NG: Toward Principled Uncertainty Modeling for Recommender Ecosystems. arXiv:2103.08057 (March2021). http:\/\/arxiv.org\/abs\/2103.08057","DOI":"10.1145\/3383313.3411527"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/1416950.1416952"},{"key":"e_1_3_3_56_2","doi-asserted-by":"crossref","DOI":"10.3998\/mpub.9739","volume-title":"Rules, Games, and Common-Pool Resources","author":"Ostrom Elinor","year":"1994","unstructured":"Elinor Ostrom, Roy Gardner, James Walker, James M. Walker, and Jimmy Walker. 1994. Rules, Games, and Common-Pool Resources. University of Michigan Press."},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1002\/asi.24203"},{"key":"e_1_3_3_58_2","doi-asserted-by":"publisher","DOI":"10.1002\/meet.1450440226"},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390255"},{"key":"e_1_3_3_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532018"},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/2365952.2365962"},{"key":"e_1_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.346"},{"key":"e_1_3_3_63_2","unstructured":"David Rohde Stephen Bonner Travis Dunlop Flavian Vasile and Alexandros Karatzoglou. 2018. RecoGym: A Reinforcement Learning Environment for the Problem of Product Recommendation in Online Advertising. arXiv:1808.00720 (Aug.2018). https:\/\/arxiv.org\/abs\/1808.00720"},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3451964.3451976"},{"key":"e_1_3_3_65_2","volume-title":"Proceedings of the 2nd International Workshop on Evaluating Information Access","author":"Sakai Tetsuya","year":"2008","unstructured":"Tetsuya Sakai and Stephen Robertson. 2008. Modelling A user population for designing information retrieval metrics. In Proceedings of the 2nd International Workshop on Evaluating Information Access."},{"key":"e_1_3_3_66_2","doi-asserted-by":"publisher","DOI":"10.1145\/122860.122897"},{"key":"e_1_3_3_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308560.3317595"},{"key":"e_1_3_3_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/1321440.1321528"},{"key":"e_1_3_3_69_2","unstructured":"Ian Soboroff. 2021. The Datasets Were Not Built to Be Solved. They Were Built as Tools to Understand the Problem and the Systems We Build to \u201cSolve\u201d Them.https:\/\/twitter.com\/ian_soboroff\/status\/1426901262369439751"},{"key":"e_1_3_3_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/1835804.1835895"},{"key":"e_1_3_3_71_2","unstructured":"Jacopo Tagliabue Federico Bianchi Tobias Schnabel Giuseppe Attanasio Ciro Greco Gabriel de Souza P. Moreira and Patrick John Chia. 2022. EvalRS: A Rounded Evaluation of Recommender Systems. arXiv:2207.05772 (July2022). http:\/\/arxiv.org\/abs\/2207.05772"},{"key":"e_1_3_3_72_2","doi-asserted-by":"publisher","DOI":"10.5555\/636669.636684"},{"key":"e_1_3_3_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/2043932.2043987"},{"key":"e_1_3_3_74_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343413.3378004"},{"key":"e_1_3_3_75_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-015-9274-y"},{"key":"e_1_3_3_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331259"},{"key":"e_1_3_3_77_2","doi-asserted-by":"publisher","DOI":"10.1145\/2911451.2914708"},{"key":"e_1_3_3_78_2","doi-asserted-by":"publisher","DOI":"10.1145\/3483382.3483384"},{"key":"e_1_3_3_79_2","doi-asserted-by":"publisher","DOI":"10.1145\/3086701"},{"key":"e_1_3_3_80_2","doi-asserted-by":"publisher","DOI":"10.1145\/2348283.2348385"},{"key":"e_1_3_3_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/3471158.3472260"},{"key":"e_1_3_3_82_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532007"},{"key":"e_1_3_3_83_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390334.1390363"},{"key":"e_1_3_3_84_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240323.3240355"},{"key":"e_1_3_3_85_2","doi-asserted-by":"publisher","DOI":"10.1145\/1871437.1871672"}],"container-title":["ACM Transactions on Recommender Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3613455","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3613455","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:36:30Z","timestamp":1750178190000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3613455"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,7]]},"references-count":84,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3613455"],"URL":"https:\/\/doi.org\/10.1145\/3613455","relation":{},"ISSN":["2770-6699"],"issn-type":[{"value":"2770-6699","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,7]]},"assertion":[{"value":"2022-12-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-06","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}