{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,1]],"date-time":"2026-08-01T00:22:40Z","timestamp":1785543760040,"version":"3.56.0"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,2,7]],"date-time":"2023-02-07T00:00:00Z","timestamp":1675728000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2023,7,31]]},"abstract":"<jats:p>\n            Recommender models are hard to evaluate, particularly under offline setting. In this article, we provide a comprehensive and critical analysis of the data leakage issue in recommender system offline evaluation. Data leakage is caused by not observing global timeline in evaluating recommenders e.g., train\/test data split does not follow global timeline. As a result, a model learns from the user-item interactions that are not expected to be available at the prediction time. We first show the temporal dynamics of user-item interactions along global timeline, then explain why data leakage exists for collaborative filtering models. Through carefully designed experiments, we show that all models indeed recommend future items that are not available at the time point of a test instance, as the result of data leakage. The experiments are conducted with four widely used baseline models\u2014BPR, NeuMF, SASRec, and LightGCN, on four popular offline datasets\u2014MovieLens-25M, Yelp, Amazon-music, and Amazon-electronic, adopting leave-last-one-out data split.\n            <jats:xref ref-type=\"fn\">\n              <jats:sup>1<\/jats:sup>\n            <\/jats:xref>\n            We further show that data leakage does impact models\u2019 recommendation accuracy. Their relative performance orders thus become unpredictable with different amount of leaked future data in training. To evaluate recommendation systems in a realistic manner in offline setting, we propose a timeline scheme, which calls for a revisit of the recommendation model design.\n          <\/jats:p>","DOI":"10.1145\/3569930","type":"journal-article","created":{"date-parts":[[2022,10,28]],"date-time":"2022-10-28T11:47:07Z","timestamp":1666957627000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":73,"title":["A Critical Study on Data Leakage in Recommender System Offline Evaluation"],"prefix":"10.1145","volume":"41","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3210-3595","authenticated-orcid":false,"given":"Yitong","family":"Ji","sequence":"first","affiliation":[{"name":"Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0764-4258","authenticated-orcid":false,"given":"Aixin","family":"Sun","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8996-7581","authenticated-orcid":false,"given":"Jie","family":"Zhang","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3144-6374","authenticated-orcid":false,"given":"Chenliang","family":"Li","sequence":"additional","affiliation":[{"name":"Wuhan University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,2,7]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240323.3240363"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463245"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331199"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/2532508.2532512"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/2043932.2043996"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/1864708.1864753"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11257-012-9136-x"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2043932.2043990"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-020-09371-3"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3137597.3137601"},{"key":"e_1_3_3_12_2","unstructured":"Jiawei Chen Hande Dong Xiang Wang Fuli Feng Meng Wang and Xiangnan He. 2020. Bias and debias in recommender system: A survey and future directions. ACM Transactions on Information Systems (TOIS) ."},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488560.3498519"},{"key":"e_1_3_3_14_2","first-page":"101","volume-title":"Proceedings of the RecSys","author":"Dacrema Maurizio Ferrari","year":"2019","unstructured":"Maurizio Ferrari Dacrema, Paolo Cremonesi, and Dietmar Jannach. 2019. Are we really making much progress? A worrying analysis of recent neural recommendation approaches. In Proceedings of the RecSys. ACM, 101\u2013109."},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-19274-7_47"},{"key":"e_1_3_3_16_2","unstructured":"Milena Filipovic Blagoj Mitrevski Diego Antognini Emma Lejal Glaude Boi Faltings and Claudiu Musat. 2021. Modeling online behavior in recommender systems: The importance of temporal context. In Proceedings of the Perspectives on the Evaluation of Recommender Systems Workshop 2021 co-located with RecSys 2021 . CEUR-WS.org."},{"key":"e_1_3_3_17_2","unstructured":"Yang Gao Yi-Fan Li Yu Lin Hang Gao and Latifur Khan. 2020. Deep learning on knowledge graph for recommender system: A survey. arXiv:2004.00387. Retrieved from https:\/\/arxiv.org\/abs\/2004.00387."},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.5555\/1577069.1755883"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4899-7637-6_8"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401063"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3038912.3052569"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210017"},{"key":"e_1_3_3_23_2","doi-asserted-by":"crossref","unstructured":"Amir Hossein Jadidinejad Craig Macdonald and Iadh Ounis. 2020. The simpson\u2019s paradox in the offline evaluation of recommendation systems. ACM Transactions on Information Systems (TOIS) 40 1 (2022) 1\u201322.","DOI":"10.1145\/3458509"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3298689.3347069"},{"key":"e_1_3_3_25_2","volume-title":"Proceedings of the REVEAL18 Workshop on offline Evaluation for Recommender Systems","author":"Jeunen Olivier","year":"2018","unstructured":"Olivier Jeunen, Koen Verstrepen, and B. Goethals. 2018. Fair offline evaluation methodologies for implicit-feedback recommender systems with MNAR data. In Proceedings of the REVEAL18 Workshop on offline Evaluation for Recommender Systems."},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401233"},{"key":"e_1_3_3_27_2","volume-title":"Proceedings of the ICTIR","author":"Ji Yitong","year":"2022","unstructured":"Yitong Ji, Aixin Sun, Jie Zhang, and Chenliang Li. 2022. Do loyal users enjoy better recommendations? Understanding recommender accuracy from a time perspective. In Proceedings of the ICTIR."},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00035"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/2382577.2382579"},{"key":"e_1_3_3_30_2","unstructured":"Sergey Kolesnikov and Mikhail Andronov. 2022. CVTT: Cross-validation through time. arXiv:2205.05393. Retrieved from https:\/\/arxiv.org\/abs\/2205.05393."},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403226"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132847.3132926"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401096"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383313.3418479"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1018"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/2645710.2645723"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240323.3240356"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460231.3474239"},{"key":"e_1_3_3_39_2","first-page":"452","volume-title":"Proceedings of the UAI","author":"Rendle Steffen","year":"2009","unstructured":"Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian personalized ranking from implicit feedback. In Proceedings of the UAI. AUAI Press, 452\u2013461."},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4899-7637-6"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/2645710.2645746"},{"key":"e_1_3_3_42_2","volume-title":"Proceedings of the RecSys.","author":"Scheidt Teresa","year":"2021","unstructured":"Teresa Scheidt and Joeran Beel. 2021. Time-dependent evaluation of recommender systems. In Proceedings of the RecSys.CEUR-WS.org."},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/2556270"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290989"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383313.3412489"},{"key":"e_1_3_3_46_2","doi-asserted-by":"crossref","unstructured":"J. Vinagre A. M. Jorge C. Rocha and J. Gama. 2019. Statistically robust evaluation of stream-based recommender systems. IEEE Transactions on Knowledge and Data Engineering 33 7 (2021) 2971\u20132982.","DOI":"10.1109\/TKDE.2019.2960216"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401133"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33015329"},{"key":"e_1_3_3_49_2","volume-title":"Proceedings of the Perspectives on the Evaluation of Recommender Systems Workshop 2021 co-located with RecSys 2021.","author":"Woolridge Daniel","year":"2021","unstructured":"Daniel Woolridge, Sean Wilner, and Madeleine Glick. 2021. Sequence or pseudo-sequence? An analysis of sequential recommendation datasets. In Proceedings of the Perspectives on the Evaluation of Recommender Systems Workshop 2021 co-located with RecSys 2021.CEUR-WS.org."},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3412754"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401154"},{"key":"e_1_3_3_52_2","article-title":"Evaluating recommender systems: Survey and framework","author":"Zangerle Eva","year":"2022","unstructured":"Eva Zangerle and Christine Bauer. 2022. Evaluating recommender systems: Survey and framework. ACM Comput. Surv. (2022).","journal-title":"ACM Comput. Surv."},{"issue":"1","key":"e_1_3_3_53_2","first-page":"5:1\u20135:38","article-title":"Deep learning based recommender system: A survey and new perspectives","volume":"52","author":"Zhang Shuai","year":"2019","unstructured":"Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay. 2019. Deep learning based recommender system: A survey and new perspectives. ACM Comput. Surv. 52, 1 (2019), 5:1\u20135:38.","journal-title":"ACM Comput. Surv."},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462875"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401167"},{"key":"e_1_3_3_56_2","article-title":"A revisiting study of appropriate offline evaluation for top-N recommendation algorithms","author":"Zhao Wayne Xin","year":"2022","unstructured":"Wayne Xin Zhao, Zihan Lin, Zhichao Feng, Pengfei Wang, and Ji-Rong Wen. 2022. A revisiting study of appropriate offline evaluation for top-N recommendation algorithms. ACM Trans. Inf. Syst.(2022).","journal-title":"ACM Trans. Inf. Syst."},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482016"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3569930","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3569930","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:07:50Z","timestamp":1750183670000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3569930"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,7]]},"references-count":56,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,7,31]]}},"alternative-id":["10.1145\/3569930"],"URL":"https:\/\/doi.org\/10.1145\/3569930","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,7]]},"assertion":[{"value":"2022-01-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-10-13","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}