{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,7]],"date-time":"2026-02-07T09:07:42Z","timestamp":1770455262793,"version":"3.49.0"},"reference-count":39,"publisher":"Association for Computing Machinery (ACM)","issue":"2","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62472277"],"award-info":[{"award-number":["62472277"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62072304"],"award-info":[{"award-number":["62072304"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172275"],"award-info":[{"award-number":["62172275"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62472176"],"award-info":[{"award-number":["62472176"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62507015"],"award-info":[{"award-number":["62507015"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62537001"],"award-info":[{"award-number":["62537001"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62572309"],"award-info":[{"award-number":["62572309"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Shanghai East Talents Program","award":["2023-177"],"award-info":[{"award-number":["2023-177"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>\n                    Sequential recommenders aim to enhance prediction accuracy by leveraging user interaction sequences, with transformer-based models showing particularly strong performance. Among them, cold-start sequential recommenders are particularly challenging because these models typically require extensive historical data to perform optimally. Some works attempt to address this issue by enhancing the adaptive ability of the sequence recommenders with meta-learning approaches. However, they are unsuitable for enhancing the popular Transformer-based sequence recommenders: MAML-based models cannot adapt the large number of parameters of Transformers, while transition-based and metric-based meta-learning models rely on unique architectures that are incompatible with Transformer-based frameworks. Also, they usually lack mechanisms to recognize and cater to multiple interests within short interaction sequences. To address these limitations, we propose MESA, a meta-modulation plugin module specifically designed to enhance the cold-start recommendation of Transformer-based sequential recommender systems. (1) We design a meta-modulation method to directly modulate the parameters in Transformer-based sequence encoders, thus enabling the model to adapt more effectively to new users in cold-start scenarios. (2) Additionally, MESA integrates the Mixture of Experts (MoE) mechanism, which refines sequence representations by utilizing multiple experts, each focusing on different aspects of user interests. This structure enhances the personalization of the recommendation by effectively handling diverse user interests within the sequences. Experiments demonstrate the effectiveness of MESA in cold-start scenarios.\n                    <jats:italic toggle=\"yes\">Our codes are available here<\/jats:italic>\n                    :\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/Mushroom-cat\/MESA\">https:\/\/github.com\/Mushroom-cat\/MESA<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3788282","type":"journal-article","created":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T11:15:12Z","timestamp":1768821312000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["MESA: Plugin Meta-Modulation for Transformer-Based Cold-Start Sequential Recommendation"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-6300-8864","authenticated-orcid":false,"given":"Xinyang","family":"Li","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6406-4992","authenticated-orcid":false,"given":"Yanmin","family":"Zhu","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1752-5423","authenticated-orcid":false,"given":"Chunyang","family":"Wang","sequence":"additional","affiliation":[{"name":"East China Normal University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0207-9643","authenticated-orcid":false,"given":"Jiadi","family":"Yu","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1384-198X","authenticated-orcid":false,"given":"Feilong","family":"Tang","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,2,4]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Lei Jimmy Ba Jamie Ryan Kiros and Geoffrey E. Hinton. 2016. Layer normalization. arXiv:1607.06450. Retrieved from https:\/\/arxiv.org\/abs\/1607.06450"},{"key":"e_1_3_2_3_2","first-page":"57","volume-title":"RecTemp@RecSys","volume":"1922","author":"Bogina Veronika","year":"2017","unstructured":"Veronika Bogina and Tsvi Kuflik. 2017. Incorporating dwell time in session-based recommendations with recurrent neural networks. In RecTemp@RecSys, Vol. 1922, CEUR-WS.org, 57\u201359."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2021.3126456"},{"key":"e_1_3_2_5_2","first-page":"2605","volume-title":"IJCAI","author":"Cheng Chen","year":"2013","unstructured":"Chen Cheng, Haiqin Yang, Michael R. Lyu, and Irwin King. 2013. Where you like to go next: Successive point-of-interest recommendation. In IJCAI. IJCAI\/AAAI, 2605\u20132611."},{"key":"e_1_3_2_6_2","first-page":"4171","volume-title":"NAACL-HLT","volume":"1","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, Vol. 1, Association for Computational Linguistics, 4171\u20134186."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1207\/s15516709cog1402_1"},{"key":"e_1_3_2_8_2","first-page":"2036","volume-title":"WWW","author":"Fan Ziwei","year":"2022","unstructured":"Ziwei Fan, Zhiwei Liu, Yu Wang, Alice Wang, Zahra Nazari, Lei Zheng, Hao Peng, and Philip S. Yu. 2022. Sequential recommendation via stochastic self-attention. In WWW. ACM, 2036\u20132047."},{"key":"e_1_3_2_9_2","first-page":"1126","volume-title":"ICML","volume":"70","author":"Finn Chelsea","year":"2017","unstructured":"Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In ICML, Vol. 70, PMLR, 1126\u20131135."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401063"},{"key":"e_1_3_2_12_2","first-page":"173","volume-title":"WWW","author":"He Xiangnan","year":"2017","unstructured":"Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017. Neural collaborative filtering. In WWW. ACM, 173\u2013182."},{"key":"e_1_3_2_13_2","first-page":"3088","volume-title":"CIKM","author":"He Zhankui","year":"2021","unstructured":"Zhankui He, Handong Zhao, Zhe Lin, Zhaowen Wang, Ajinkya Kale, and Julian J. McAuley. 2021. Locker: Locally constrained self-attentive sequential recommendation. In CIKM. ACM, 3088\u20133092."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3271761"},{"key":"e_1_3_2_15_2","volume-title":"ICLR","author":"Hidasi Bal\u00e1zs","year":"2016","unstructured":"Bal\u00e1zs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2016. Session-based recommendations with recurrent neural networks. In ICLR."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466753"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1991.3.1.79"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00035"},{"key":"e_1_3_2_19_2","volume-title":"ICLR","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In ICLR."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330859"},{"key":"e_1_3_2_21_2","first-page":"1306","volume-title":"WWW","author":"Lin Xixun","year":"2021","unstructured":"Xixun Lin, Jia Wu, Chuan Zhou, Shirui Pan, Yanan Cao, and Bin Wang. 2021. Task-adaptive neural process for user Cold-Start recommendation. In WWW. ACM, 1306\u20131316."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3616855.3635784"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3718734"},{"key":"e_1_3_2_24_2","unstructured":"Steffen Rendle Christoph Freudenthaler Zeno Gantner and Lars Schmidt-Thieme. 2012. BPR: Bayesian personalized ranking from implicit feedback. arXiv:1205.2618. Retrieved from https:\/\/arxiv.org\/abs\/1205.2618"},{"key":"e_1_3_2_25_2","first-page":"811","volume-title":"WWW","author":"Rendle Steffen","year":"2010","unstructured":"Steffen Rendle, Christoph Freudenthaler, and Lars Schmidt-Thieme. 2010. Factorizing personalized Markov chains for next-basket recommendation. In WWW. ACM, 811\u2013820."},{"key":"e_1_3_2_26_2","volume-title":"ICLR","author":"Shazeer Noam","year":"2017","unstructured":"Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In ICLR. OpenReview.net."},{"key":"e_1_3_2_27_2","first-page":"1713","volume-title":"CIKM","author":"Song Jiayu","year":"2021","unstructured":"Jiayu Song, Jiajie Xu, Rui Zhou, Lu Chen, Jianxin Li, and Chengfei Liu. 2021. CBML: A cluster-based meta-learning model for session-based recommendation. In CIKM. ACM, 1713\u20131722."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357895"},{"key":"e_1_3_2_29_2","first-page":"17","volume-title":"DLRS@RecSys","author":"Tan Yong Kiam","year":"2016","unstructured":"Yong Kiam Tan, Xinxing Xu, and Yong Liu. 2016. Improved recurrent neural networks for session-based recommendations. In DLRS@RecSys. ACM, 17\u201322."},{"key":"e_1_3_2_30_2","first-page":"5998","volume-title":"NIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NIPS, 5998\u20136008."},{"key":"e_1_3_2_31_2","unstructured":"Chunyang Wang Yanmin Zhu Haobing Liu Tianzi Zang Jiadi Yu and Feilong Tang. 2022. 2022. Deep meta-learning in recommendation systems: A survey. arXiv:2206.04415. Retrieved from https:\/\/arxiv.org\/abs\/2206.04415"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463089"},{"key":"e_1_3_2_33_2","first-page":"155","volume-title":"SIGIR","author":"Wang Pengfei","year":"2019","unstructured":"Pengfei Wang, Hanxiong Chen, Yadong Zhu, Huawei Shen, and Yongfeng Zhang. 2019. Unified collaborative filtering over graph embeddings. In SIGIR. ACM, 155\u2013164."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331267"},{"key":"e_1_3_2_35_2","first-page":"661","volume-title":"ICDM","author":"Wei Tianxin","year":"2020","unstructured":"Tianxin Wei, Ziwei Wu, Ruirui Li, Ziniu Hu, Fuli Feng, Xiangnan He, Yizhou Sun, and Wei Wang. 2020. Fast adaptation for cold-start collaborative filtering with meta-learning. In ICDM. IEEE, 661\u2013670."},{"key":"e_1_3_2_36_2","first-page":"1021","volume-title":"WWW","author":"Wu Shiguang","year":"2023","unstructured":"Shiguang Wu, Yaqing Wang, Qinghe Jing, Daxiang Dong, Dejing Dou, and Quanming Yao. 2023. ColdNAS: Search to modulate for user cold-start recommendation. In WWW. ACM, 1021\u20131031."},{"key":"e_1_3_2_37_2","first-page":"359","volume-title":"KDD","author":"Yin Jianwen","year":"2020","unstructured":"Jianwen Yin, Chenghao Liu, Weiqing Wang, Jianling Sun, and Steven C. H. Hoi. 2020. Learning transferrable parameters for long-tailed sequential user behavior modeling. In KDD. ACM, 359\u2013367."},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","first-page":"114077","DOI":"10.1109\/ACCESS.2019.2936461","article-title":"GACOforRec: Session-based graph convolutional neural networks recommendation model","volume":"7","author":"Zhang Mingge","year":"2019","unstructured":"Mingge Zhang and Zhenyu Yang. 2019. GACOforRec: Session-based graph convolutional neural networks recommendation model. IEEE Access 7 (2019), 114077\u2013114085.","journal-title":"IEEE Access"},{"key":"e_1_3_2_39_2","first-page":"2220","volume-title":"WWW","author":"Zhang Yin","year":"2021","unstructured":"Yin Zhang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi, Lichan Hong, and Ed H. Chi. 2021. A model of two tales: Dual transfer learning framework for improved long-tail item recommendation. In WWW. ACM, 2220\u20132231."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i5.16601"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3788282","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T12:45:20Z","timestamp":1770209120000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3788282"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,4]]},"references-count":39,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3788282"],"URL":"https:\/\/doi.org\/10.1145\/3788282","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,4]]},"assertion":[{"value":"2024-09-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-04","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}