{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T05:17:35Z","timestamp":1784179055843,"version":"3.55.0"},"reference-count":64,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,8,2]],"date-time":"2024-08-02T00:00:00Z","timestamp":1722556800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Recomm. Syst."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>Many modern sequential recommender systems use deep neural networks, which can effectively estimate the relevance of items but require a lot of time to train. Slow training increases the costs of training, hinders product development timescales, and prevents the model from being regularly updated to adapt to changing user preferences. The training of such sequential models involves appropriately sampling past user interactions to create a realistic training objective. The existing training objectives have limitations. For instance, next item prediction never uses the beginning of the sequence as a learning target, thereby potentially discarding valuable data. However, the item masking used by the state-of-the-art BERT4Rec recommender model is only weakly related to the goal of the sequential recommendation; therefore, it requires much more time to obtain an effective model. Hence, we propose a novel Recency-based Sampling of Sequences (RSS) training objective (which is parameterized by a choice of recency importance function) that addresses both limitations. We apply our method to various recent and state-of-the-art model architectures\u2014such as GRU4Rec, Caser, and SASRec. We show that the models enhanced with our method can achieve performances exceeding or very close to the effective BERT4Rec, but with much less training time. For example, on the MovieLens-20M dataset, RSS applied to the SASRec model can result in a 60% improvement in NDCG over a vanilla SASRec, and a 16% improvement over a fully trained BERT4Rec model, despite taking 93% less training time than BERT4Rec. We also experiment with two families of recency importance functions and show that they perform similarly. We further empirically demonstrate that RSS-enhanced SASRec successfully learns to distinguish differences between recent and older interactions\u2014a property that the original SASRec model does not exhibit. Overall, we show that RSS is a viable (and frequently better) alternative to the existing training objectives, which is both effective and efficient for training sequential recommender model when the computational resources for training are limited.<\/jats:p>","DOI":"10.1145\/3604436","type":"journal-article","created":{"date-parts":[[2023,6,23]],"date-time":"2023-06-23T09:20:58Z","timestamp":1687512058000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["RSS: Effective and Efficient Training for Sequential Recommendation Using Recency Sampling"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0911-3605","authenticated-orcid":false,"given":"Aleksandr","family":"Petrov","sequence":"first","affiliation":[{"name":"University of Glasgow, Glasgow, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3143-279X","authenticated-orcid":false,"given":"Craig","family":"Macdonald","sequence":"additional","affiliation":[{"name":"University of Glasgow, Glasgow, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,8,2]]},"reference":[{"key":"e_1_3_3_2_2","first-page":"265","volume-title":"Proceedings of the USENIX","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et\u00a0al. 2016. TensorFlow: A system for large-scale machine learning. In Proceedings of the USENIX. 265\u2013283."},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-88942-5_24"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3481905"},{"issue":"23","key":"e_1_3_3_5_2","first-page":"81","article-title":"From RankNet to LambdaRank to LambdaMART: An overview","volume":"11","author":"Burges Christopher J. C.","year":"2010","unstructured":"Christopher J. C. Burges. 2010. From RankNet to LambdaRank to LambdaMART: An overview. Learning 11, 23-581 (2010), 81.","journal-title":"Learning"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383313.3412259"},{"key":"e_1_3_3_7_2","unstructured":"Olivier Chapelle and Yi Chang. 2011. Yahoo! learning to rank challenge overview. In Proceedings of the Learning to Rank Challenge. PMLR 1\u201324."},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/2020408.2020579"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383313.3412216"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460231.3475943"},{"key":"e_1_3_3_11_2","first-page":"4171","volume-title":"Proceedings of the NAACL-HLT","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the NAACL-HLT. 4171\u20134186."},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58285-2_23"},{"key":"e_1_3_3_13_2","first-page":"1","volume-title":"Proceedings of the ACM SIGIR Forum","volume":"54","author":"Fuhr Norbert","year":"2021","unstructured":"Norbert Fuhr. 2021. Proof by experimentation? Towards better IR research. In Proceedings of the ACM SIGIR Forum, Vol. 54. 1\u20134."},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463240"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-0716-2197-4_15"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/2827872"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3271761"},{"key":"e_1_3_3_18_2","volume-title":"Proceedings of the ICLR","author":"Hidasi Bal\u00e1zs","year":"2016","unstructured":"Bal\u00e1zs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2016. Session-based recommendations with recurrent neural networks. In Proceedings of the ICLR."},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313447"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210017"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00035"},{"key":"e_1_3_3_22_2","first-page":"3146","volume-title":"Proceedings of the NeurIPS","author":"Ke Guolin","year":"2017","unstructured":"Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017. LightGBM: A highly efficient gradient boosting decision tree. In Proceedings of the NeurIPS. 3146\u20133154."},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDMW53433.2021.00013"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-0716-2197-4_3"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403226"},{"key":"e_1_3_3_26_2","first-page":"14","volume-title":"Proceedings of the Workshop on New Trends on Content-Based Recommender @ RecSys (CEUR Workshop)","volume":"1448","author":"Kula Maciej","year":"2015","unstructured":"Maciej Kula. 2015. Metadata embeddings for user and item cold-start recommendations. In Proceedings of the Workshop on New Trends on Content-Based Recommender @ RecSys (CEUR Workshop), Vol. 1448. 14\u201321."},{"key":"e_1_3_3_27_2","article-title":"MOI-Mixer: Improving MLP-mixer with multi order interactions in sequential recommendation","author":"Lee Hojoon","year":"2021","unstructured":"Hojoon Lee, Dongyoon Hwang, Sunghwan Hong, Changyeon Kim, Seungryong Kim, and Jaegul Choo. 2021. MOI-Mixer: Improving MLP-mixer with multi order interactions in sequential recommendation. Retrieved from https:\/\/arXiv:2108.07505.","journal-title":"Retrieved from https:\/\/arXiv:2108.07505"},{"key":"e_1_3_3_28_2","doi-asserted-by":"crossref","unstructured":"H. Li X. Wang Z. Zhang J. Ma P. Cui and W. Zhu. 2021. Intention-aware sequential recommendation with structured intent transition. IEEE Trans. Knowl. and Data Eng. 34 11 (2021) 5403\u20135414.","DOI":"10.1109\/TKDE.2021.3050571"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462973"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1561\/1500000016"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463036"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11257-018-9209-6"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330984"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403091"},{"key":"e_1_3_3_35_2","first-page":"1","article-title":"Variational Bayesian representation learning for grocery recommendation","volume":"24","author":"Meng Zaiqiao","year":"2021","unstructured":"Zaiqiao Meng, Richard McCreadie, Craig Macdonald, and Iadh Ounis. 2021. Variational Bayesian representation learning for grocery recommendation. Info. Retrieval J. 24 (102021), 1\u201323.","journal-title":"Info. Retrieval J."},{"key":"e_1_3_3_36_2","doi-asserted-by":"crossref","unstructured":"Umaporn Padungkiatwattana Thitiya Sae-Diae Saranya Maneeroj and Atsuhiro Takasu. 2022. ARERec: Attentive local interaction model for sequential recommendation. IEEE Access 10 (2022) 31340\u201331358.","DOI":"10.1109\/ACCESS.2022.3160466"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3523227.3546785"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3523227.3548487"},{"key":"e_1_3_3_39_2","first-page":"41","volume-title":"Proceedings of the WSDM WebTour","author":"Petrov Aleksandr","year":"2021","unstructured":"Aleksandr Petrov and Yuriy Makarov. 2021. Attention-based neural re-ranking approach for next city in trip recommendations. In Proceedings of the WSDM WebTour. 41\u201345."},{"key":"e_1_3_3_40_2","volume-title":"Proceedings of the ICLR","author":"Qin Zhen","year":"2021","unstructured":"Zhen Qin, Le Yan, Honglei Zhuang, Yi Tay, Rama Kumar Pasumarthi, Xuanhui Wang, Michael Bendersky, and Marc Najork. 2021. Are neural rankers still outperformed by gradient boosted decision trees?. In Proceedings of the ICLR."},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3473339"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3382764"},{"key":"e_1_3_3_43_2","unstructured":"Ruihong Qiu Zi Huang and Hongzhi Yin. 2021. Memory augmented multi-instance contrastive predictive coding for sequential recommendation. Retrieved from https:\/\/arxiv.org\/abs\/2109.00368."},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488560.3498433"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401109"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3190616"},{"key":"e_1_3_3_47_2","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language Models Are Unsupervised Multitask Learners. https:\/\/cdn.openai.com\/better-language-models\/language_models_are_unsupervised_multitask_learners.pdf."},{"issue":"140","key":"e_1_3_3_48_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 140 (2020), 1\u201367.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_3_49_2","first-page":"452","volume-title":"Proceedings of the CUAI","author":"Rendle Steffen","year":"2009","unstructured":"Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian personalized ranking from implicit feedback. In Proceedings of the CUAI. 452\u2013461."},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772773"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357895"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3159652.3159656"},{"key":"e_1_3_3_53_2","volume-title":"Natural Language Processing with Transformers, Revised Edition","author":"Tunstall Lewis","year":"2022","unstructured":"Lewis Tunstall, Leandro von Werra, and Thomas Wolf. 2022. Natural Language Processing with Transformers, Revised Edition. O\u2019Reilly Media, Inc."},{"key":"e_1_3_3_54_2","first-page":"5998","volume-title":"Proceedings of the NeurIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the NeurIPS. 5998\u20136008."},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3271786"},{"key":"e_1_3_3_56_2","doi-asserted-by":"crossref","unstructured":"Chenyang Wang Weizhi Ma Chong Chen Min Zhang Yiqun Liu and Shaoping Ma. 2023. Sequential recommendation with multiple contrast signals. ACM Transactions on Information Systems 41 1 (2023) 1\u201327.","DOI":"10.1145\/3522673"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482296"},{"key":"e_1_3_3_58_2","unstructured":"Xu Xie Fei Sun Zhaoyang Liu Shiwen Wu Jinyang Gao Bolin Ding and Bin Cui. 2020. Contrastive learning for sequential recommendation. Retrieved from https:\/\/arXiv:2010.14395."},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/2983323.2983758"},{"key":"e_1_3_3_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290975"},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-022-1184-8"},{"key":"e_1_3_3_62_2","article-title":"Dynamic graph neural networks for sequential recommendation","author":"Zhang Mengqi","year":"2022","unstructured":"Mengqi Zhang, Shu Wu, Xueli Yu, Qiang Liu, and Liang Wang. 2022. Dynamic graph neural networks for sequential recommendation. IEEE Trans. Knowl. Data Eng. (TKDE) (2022).","journal-title":"IEEE Trans. Knowl. Data Eng. (TKDE)"},{"key":"e_1_3_3_63_2","volume-title":"Proceedings of the ICJAI","author":"Zhao Pengyu","year":"2021","unstructured":"Pengyu Zhao, Tianxiao Shui, Yuanxing Zhang, Kecheng Xiao, and Kaigui Bian. 2021. Adversarial oracular seq2seq learning for sequential recommendation. In Proceedings of the ICJAI."},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3411954"},{"key":"e_1_3_3_65_2","first-page":"580","volume-title":"Proceedings of the UAI","author":"Zimdars Andrew","year":"2001","unstructured":"Andrew Zimdars, David Maxwell Chickering, and Christopher Meek. 2001. Using temporal data for making recommendations. In Proceedings of the UAI. 580\u2013588."}],"container-title":["ACM Transactions on Recommender Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3604436","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3604436","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:18Z","timestamp":1750178838000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3604436"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,2]]},"references-count":64,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3604436"],"URL":"https:\/\/doi.org\/10.1145\/3604436","relation":{},"ISSN":["2770-6699"],"issn-type":[{"value":"2770-6699","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,2]]},"assertion":[{"value":"2023-01-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-02","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}