{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,7]],"date-time":"2026-03-07T18:00:00Z","timestamp":1772906400948,"version":"3.50.1"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2025,4,10]],"date-time":"2025-04-10T00:00:00Z","timestamp":1744243200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Recomm. Syst."],"published-print":{"date-parts":[[2025,12,31]]},"abstract":"<jats:p>\n            A large catalogue size is one of the central challenges in training recommendation models: A large number of items makes them memory- and computationally inefficient to compute scores for all items during training, forcing these models to deploy negative sampling. However, negative sampling increases the proportion of positive interactions in the training data, and therefore, models trained with negative sampling tend to overestimate the probabilities of positive interactions\u2014a phenomenon we call\n            <jats:italic>overconfidence<\/jats:italic>\n            . While the absolute values of the predicted scores\/probabilities are not important for the ranking of retrieved recommendations, overconfident models may fail to estimate nuanced differences in the top-ranked items, resulting in degraded performance. In this article, we show that overconfidence explains why the popular SASRec model underperforms when compared to BERT4Rec. This is contrary to the BERT4Rec authors\u2019 explanation that the difference in performance is due to the bi-directional attention mechanism. To mitigate overconfidence, we propose a novel Generalised Binary Cross-Entropy Loss function (gBCE) and theoretically prove that it can mitigate overconfidence. We further propose the gSASRec model, an improvement over SASRec that deploys an increased number of negatives and the gBCE loss. Through detailed experiments on three datasets, we show that gSASRec does not exhibit the overconfidence problem. As a result, gSASRec can outperform BERT4Rec (e.g., +9.47% NDCG on the MovieLens-1M dataset) while requiring less training time (e.g., -73% training time on MovieLens-1M). Moreover, in contrast to BERT4Rec, gSASRec is suitable for large datasets that contain more than 1 million items. Finally, we show how addressing overconfidence can improve model calibration\u2014the ability of a model to predict actual interaction probabilities accurately. By applying gBCE to the SASRec model on MovieLens-1M dataset, we reduce the models\u2019 expected calibration error by 98.9% (from 0.966 to 0.01).\n          <\/jats:p>","DOI":"10.1145\/3699521","type":"journal-article","created":{"date-parts":[[2024,10,7]],"date-time":"2024-10-07T10:16:08Z","timestamp":1728296168000},"page":"1-34","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Improving Effectiveness by Reducing Overconfidence in Large Catalogue Sequential Recommendation with gBCE loss"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0911-3605","authenticated-orcid":false,"given":"Aleksandr Vladimirovich","family":"Petrov","sequence":"first","affiliation":[{"name":"School Of Computing Science, University of Glasgow, Glasgow, United Kingdom of Great Britain and Northern Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3143-279X","authenticated-orcid":false,"given":"Craig","family":"Macdonald","sequence":"additional","affiliation":[{"name":"School Of Computing Science, University of Glasgow, Glasgow, United Kingdom of Great Britain and Northern Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,4,10]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Jason Wise. 2023. How Many Videos Are on YouTube in 2024? Earthweb. Retrieved from https:\/\/earthweb.com\/how-many-videos-are-on-youtube\/"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3573128.3604896"},{"key":"e_1_3_3_4_2","article-title":"From RankNet to LambdaRank to LambdaMART: An overview","volume":"11","author":"Burges Christopher","year":"2010","unstructured":"Christopher Burges. 2010. From RankNet to LambdaRank to LambdaMART: An overview. Learning 11 (Jan.2010).","journal-title":"Learning"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383313.3412259"},{"key":"e_1_3_3_6_2","doi-asserted-by":"crossref","unstructured":"Yongjun Chen Jia Li Zhiwei Liu Nitish Shirish Keskar Huan Wang Julian McAuley and Caiming Xiong. 2022. Generating Negative Samples for Sequential Recommendation. arxiv:2208.03645 [cs]","DOI":"10.1145\/3485447.3512090"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/2020408.2020579"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/312624.312692"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2959100.2959190"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00949"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460231.3475943"},{"key":"e_1_3_3_12_2","first-page":"4171","volume-title":"NAACL-HLT","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT. 4171\u20134186."},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.98"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557266"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10710-017-9314-z"},{"key":"e_1_3_3_16_2","first-page":"1321","volume-title":"ICML","author":"Guo Chuan","year":"2017","unstructured":"Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In ICML. 1321\u20131330."},{"issue":"4","key":"e_1_3_3_17_2","first-page":"19:1\u201319:19","article-title":"The MovieLens datasets: History and context","volume":"5","author":"Harper F. Maxwell","year":"2015","unstructured":"F. Maxwell Harper and Joseph A. Konstan. 2015. The MovieLens datasets: History and context. ACM Trans. Interact. Intell. Syst. 5, 4 (Dec.2015), 19:1\u201319:19.","journal-title":"ACM Trans. Interact. Intell. Syst."},{"key":"e_1_3_3_18_2","first-page":"173","volume-title":"WWW","author":"He Xiangnan","year":"2017","unstructured":"Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017. Neural collaborative filtering. In WWW. 173\u2013182."},{"key":"e_1_3_3_19_2","volume-title":"Information Retrieval: Computational and Theoretical Aspects","author":"Heaps H. S.","year":"1978","unstructured":"H. S. Heaps. 1978. Information Retrieval: Computational and Theoretical Aspects. Academic Press, Inc., USA."},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3604915.3608839"},{"key":"e_1_3_3_21_2","volume-title":"ICLR","author":"Hidasi Bal\u00e1zs","year":"2016","unstructured":"Bal\u00e1zs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2016. Session-based recommendations with recurrent neural networks. In ICLR."},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1001"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401233"},{"key":"e_1_3_3_24_2","unstructured":"Zhaohong Jia Hu Zhang Wei Gao Fulan Qian Hai Chen Chunchun Li and Kai Li. 2022. Dual Attention Recommendation Algorithm Based on Item Attributes. Retrieved from https:\/\/www.researchsquare.com\/article\/rs-2206282\/v1"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00035"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3604915.3610644"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2009.263"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3535335"},{"key":"e_1_3_3_29_2","volume-title":"ICLR","author":"Lan Zhenzhong","year":"2020","unstructured":"Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In ICLR."},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-02181-7"},{"key":"e_1_3_3_31_2","unstructured":"Tsung-Yi Lin Priya Goyal Ross Girshick Kaiming He and Piotr Doll\u00e1r. 2018. Focal Loss for Dense Object Detection. arxiv:1708.02002 [cs]"},{"key":"e_1_3_3_32_2","volume-title":"NeurIPS","author":"Morozov Stanislav","year":"2018","unstructured":"Stanislav Morozov and Artem Babenko. 2018. Non-metric similarity graphs for maximum inner product search. In NeurIPS, Vol. 31."},{"key":"e_1_3_3_33_2","volume-title":"Probabilistic Machine Learning: An Introduction","author":"Murphy Kevin P.","year":"2022","unstructured":"Kevin P. Murphy. 2022. Probabilistic Machine Learning: An Introduction. The MIT Press, Cambridge, Massachusetts. Q325.5 .M872 2022"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.2981890"},{"key":"e_1_3_3_35_2","volume-title":"AAAI","author":"Naeini Mahdi Pakdaman","year":"2015","unstructured":"Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015. Obtaining well calibrated probabilities using Bayesian binning. In AAAI, Vol. 29."},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080724"},{"key":"e_1_3_3_37_2","first-page":"188","volume-title":"RecSys","author":"Pellegrini Roberto","year":"2022","unstructured":"Roberto Pellegrini, Wenjie Zhao, and Iain Murray. 2022. Don\u2019t recommend the obvious: Estimate probability ratios. In RecSys. 188\u2013197."},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3523227.3546785"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3523227.3548487"},{"key":"e_1_3_3_40_2","article-title":"RSS: Effective and efficient training for sequential recommendation using recency sampling","author":"Petrov Aleksandr","year":"2023","unstructured":"Aleksandr Petrov and Craig Macdonald. 2023. RSS: Effective and efficient training for sequential recommendation using recency sampling. ACM Trans. Recomm. Syst. 3, 1 (2023). Association for Computing Machinery.","journal-title":"ACM Trans. Recomm. Syst."},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3616855.3635821"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-56063-7_10"},{"key":"e_1_3_3_43_2","volume-title":"GenIR@SIGIR","author":"Petrov Aleksandr V.","year":"2023","unstructured":"Aleksandr V. Petrov and Craig Macdonald. 2023. Generative sequential recommendation with GPTRec. In GenIR@SIGIR."},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3604915.3608783"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.2478\/congeo-2019-0022"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488560.3498433"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-0716-2197-4_4"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/2556195.2556248"},{"key":"e_1_3_3_49_2","volume-title":"UAI","author":"Rendle Steffen","year":"2009","unstructured":"Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian personalized ranking from implicit feedback. In UAI."},{"key":"e_1_3_3_50_2","first-page":"811","volume-title":"WWW","author":"Rendle Steffen","year":"2010","unstructured":"Steffen Rendle, Christoph Freudenthaler, and Lars Schmidt-Thieme. 2010. Factorizing personalized Markov chains for next-basket recommendation. In WWW. 811."},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3523227.3548486"},{"issue":"1","key":"e_1_3_3_52_2","first-page":"1","article-title":"Shopper intent prediction from clickstream e-commerce data with minimal browsing information","volume":"10","author":"Requena Borja","year":"2020","unstructured":"Borja Requena, Giovanni Cassani, Jacopo Tagliabue, Ciro Greco, and Lucas Lacasa. 2020. Shopper intent prediction from clickstream e-commerce data with minimal browsing information. Scient. Rep. 10, 1 (2020), 1\u201323.","journal-title":"Scient. Rep."},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.1093\/jxb\/10.2.290"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-018-1254-2"},{"key":"e_1_3_3_55_2","first-page":"270","volume-title":"Broadcast News Transcription and Understanding Workshop","author":"Stolcke A.","year":"1998","unstructured":"A. Stolcke. 1998. Entropy-based pruning of backoff language models. In Broadcast News Transcription and Understanding Workshop. 270\u2013274."},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357895"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3159652.3159656"},{"key":"e_1_3_3_58_2","volume-title":"NeurIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS."},{"key":"e_1_3_3_59_2","volume-title":"ICML","author":"Wei Hongxin","year":"2022","unstructured":"Hongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng, Bo An, and Yixuan Li. 2022. Mitigating neural network overconfidence with logit normalization. In ICML."},{"key":"e_1_3_3_60_2","volume-title":"IJCAI","author":"Weston Jason","year":"2011","unstructured":"Jason Weston, Samy Bengio, and Nicolas Usunier. 2011. WSABIE: Scaling up to large vocabulary image annotation. In IJCAI."},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3637061"},{"key":"e_1_3_3_62_2","unstructured":"Yonghui Wu Mike Schuster Zhifeng Chen Quoc V. Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey Jeff Klingner Apurva Shah Melvin Johnson Xiaobing Liu \u0141ukasz Kaiser Stephan Gouws Yoshikiyo Kato Taku Kudo Hideto Kazawa Keith Stevens George Kurian Nishant Patil Wei Wang Cliff Young Jason Smith Jason Riesa Alex Rudnick Oriol Vinyals Greg Corrado Macduff Hughes and Jeffrey Dean. 2016. Google\u2019s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. arxiv:1609.08144 [cs]"},{"key":"e_1_3_3_63_2","first-page":"3410","volume-title":"WWW","author":"Xie Jiayi","year":"2024","unstructured":"Jiayi Xie, Shang Liu, Gao Cong, and Zhenzhong Chen. 2024. UnifiedSSR: A unified framework of sequential search and recommendation. In WWW. 3410\u20133419."},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00099"},{"key":"e_1_3_3_65_2","doi-asserted-by":"publisher","DOI":"10.5555\/3172077.3172336"},{"key":"e_1_3_3_66_2","doi-asserted-by":"publisher","DOI":"10.1145\/3616855.3635848"},{"key":"e_1_3_3_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/2983323.2983758"},{"key":"e_1_3_3_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290975"},{"key":"e_1_3_3_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657761"},{"key":"e_1_3_3_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3411954"},{"key":"e_1_3_3_71_2","first-page":"3854","volume-title":"WWW","author":"Zhou Peilin","year":"2024","unstructured":"Peilin Zhou, You-Liang Huang, Yueqi Xie, Jingqi Gao, Shoujin Wang, Jae Boum Kim, and Sunghun Kim. 2024. Is contrastive learning necessary? A study of data augmentation vs contrastive learning in sequential recommendation. In WWW. 3854\u20133863."}],"container-title":["ACM Transactions on Recommender Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3699521","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3699521","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:09:51Z","timestamp":1750295391000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3699521"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,10]]},"references-count":70,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,12,31]]}},"alternative-id":["10.1145\/3699521"],"URL":"https:\/\/doi.org\/10.1145\/3699521","relation":{},"ISSN":["2770-6699"],"issn-type":[{"value":"2770-6699","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,10]]},"assertion":[{"value":"2024-02-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-20","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}