{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T02:27:33Z","timestamp":1783736853994,"version":"3.55.0"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2024,1,22]],"date-time":"2024-01-22T00:00:00Z","timestamp":1705881600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2024,5,31]]},"abstract":"<jats:p>Self-attention models have achieved the state-of-the-art performance in sequential recommender systems by capturing the sequential dependencies among user\u2013item interactions. However, they rely on adding positional embeddings to the item sequence to retain the sequential information, which may break the semantics of item embeddings due to the heterogeneity between these two types of embeddings. In addition, most existing works assume that such dependencies exist solely in the item embeddings, but neglect their existence among the item features. In our previous study, we proposed a novel sequential recommendation model, i.e., MLP4Rec, based on the recent advances of MLP-Mixer architectures, which is naturally sensitive to the order of items in a sequence because matrix elements related to different positions of a sequence will be given different weights in training. We developed a tri-directional fusion scheme to coherently capture sequential, cross-channel, and cross-feature correlations with linear computational complexity as well as much fewer model parameters than existing self-attention methods. However, the cascading mixer structure, the large number of normalization layers between different mixer layers, and the noise generated by these operations limit the efficiency of information extraction and the effectiveness of MLP4Rec. In this extended version, we propose a novel framework \u2013 SMLP4Rec for sequential recommendation to address the aforementioned issues. The new framework changes the flawed cascading structure to a parallel mode, and integrates normalization layers to minimize their impact on the model\u2019s efficiency while maximizing their effectiveness. As a result, the training speed and prediction accuracy of SMLP4Rec are vastly improved in comparison to MLP4Rec. Extensive experimental results demonstrate that the proposed method is significantly superior to the state-of-the-art approaches. The implementation code is available online to ease reproducibility.<\/jats:p>","DOI":"10.1145\/3637871","type":"journal-article","created":{"date-parts":[[2023,12,18]],"date-time":"2023-12-18T11:55:39Z","timestamp":1702900539000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":45,"title":["SMLP4Rec: An Efficient All-MLP Architecture for Sequential Recommendations"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4470-5972","authenticated-orcid":false,"given":"Jingtong","family":"Gao","sequence":"first","affiliation":[{"name":"City University of Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2926-4416","authenticated-orcid":false,"given":"Xiangyu","family":"Zhao","sequence":"additional","affiliation":[{"name":"City University of Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1009-2343","authenticated-orcid":false,"given":"Muyang","family":"Li","sequence":"additional","affiliation":[{"name":"University of Sydney"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2871-0023","authenticated-orcid":false,"given":"Minghao","family":"Zhao","sequence":"additional","affiliation":[{"name":"NetEase FUXI AI Lab"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6986-5825","authenticated-orcid":false,"given":"Runze","family":"Wu","sequence":"additional","affiliation":[{"name":"NetEase FUXI AI Lab"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8522-6142","authenticated-orcid":false,"given":"Ruocheng","family":"Guo","sequence":"additional","affiliation":[{"name":"Bytedance AI Lab"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6857-261X","authenticated-orcid":false,"given":"Yiding","family":"Liu","sequence":"additional","affiliation":[{"name":"Baidu Inc"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0684-6205","authenticated-orcid":false,"given":"Dawei","family":"Yin","sequence":"additional","affiliation":[{"name":"Baidu Inc"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,1,22]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","unstructured":"Heng-Tze Cheng Levent Koc Jeremiah Harmsen Tal Shaked Tushar Chandra Hrishi Aradhye Glen Anderson Greg Corrado Wei Chai Mustafa Ispir Rohan Anil Zakaria Haque Lichan Hong Vihan Jain Xiaobing Liu and Hemal Shah. 2016. Wide & deep learning for recommender systems. In Proceedings of the DLRS.","DOI":"10.1145\/2988450.2988454"},{"key":"e_1_3_2_3_2","volume-title":"Proceedings of the CIKM","author":"Du Hanwen","year":"2022","unstructured":"Hanwen Du, Hui Shi, Pengpeng Zhao, Deqing Wang, Victor S. Sheng, Yanchi Liu, Guanfeng Liu, and Lei Zhao. 2022. Contrastive learning with bidirectional transformers for sequential recommendation. In Proceedings of the CIKM."},{"key":"e_1_3_2_4_2","unstructured":"Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (GELUs). arXiv:1606.08415v5. Retrieved from https:\/\/arxiv.org\/abs\/1606.08415v5"},{"key":"e_1_3_2_5_2","unstructured":"Bal\u00e1zs Hidasi Alexandros Karatzoglou Linas Baltrunas and Domonkos Tikk. 2015. Session-based recommendations with recurrent neural networks. arXiv:1511.06939v4. Retrieved from https:\/\/arxiv.org\/abs\/1511.06939v4"},{"key":"e_1_3_2_6_2","volume-title":"Proceedings of the RecSys","author":"Hidasi Bal\u00e1zs","year":"2016","unstructured":"Bal\u00e1zs Hidasi, Massimo Quadrana, Alexandros Karatzoglou, and Domonkos Tikk. 2016. Parallel recurrent neural network architectures for feature-rich session-based recommendations. In Proceedings of the RecSys."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00035"},{"key":"e_1_3_2_8_2","volume-title":"Proceedings of the ICML","author":"Katharopoulos Angelos","year":"2020","unstructured":"Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Fran\u00e7ois Fleuret. 2020. Transformers are RNNs: Fast autoregressive transformers with linear attention. In Proceedings of the ICML."},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the AACL","author":"Kenton Jacob Devlin, Ming-Wei Chang,","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, KentonLee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the AACL."},{"key":"e_1_3_2_10_2","volume-title":"Proceedings of the ICLR (Poster)","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the ICLR (Poster)."},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Shun Kiyono Sosuke Kobayashi Jun Suzuki and Kentaro Inui. 2021. SHAPE: Shifted absolute position embedding for transformers. In Proceedings of the EMNLP.","DOI":"10.18653\/v1\/2021.emnlp-main.266"},{"key":"e_1_3_2_12_2","unstructured":"Hojoon Lee Dongyoon Hwang Sunghwan Hong Changyeon Kim Seungryong Kim and Jaegul Choo. 2021. MOI-mixer: Improving MLP-mixer with multi order interactions in sequential recommendation. arXiv:2108.07505v1. Retrieved from https:\/\/arxiv.org\/abs\/2108.07505v1"},{"key":"e_1_3_2_13_2","volume-title":"Proceedings of the RecSys","author":"Li Chengxi","year":"2023","unstructured":"Chengxi Li, Yejing Wang, Qidong Liu, Xiangyu Zhao, Wanyu Wang, Yiqi Wang, Lixin Zou, Wenqi Fan, and Qing Li. 2023. STRec: Sparse transformer for sequential recommendations. In Proceedings of the RecSys."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132847.3132926"},{"key":"e_1_3_2_15_2","volume-title":"Proceedings of the WWW","author":"Li Muyang","year":"2023","unstructured":"Muyang Li, Zijian Zhang, Xiangyu Zhao, Wanyu Wang, Minghao Zhao, Runze Wu, and Ruocheng Guo. 2023. AutoMLP: Automated MLP for sequential recommendations. In Proceedings of the WWW."},{"key":"e_1_3_2_16_2","volume-title":"Proceedings of the IJCAI","author":"Li Muyang","year":"2022","unstructured":"Muyang Li, Xiangyu Zhao, Chuan Lyu, Minghao Zhao, Runze Wu, and Ruocheng Guo. 2022. MLP4Rec: A pure MLP architecture for sequential recommendations. In Proceedings of the IJCAI."},{"key":"e_1_3_2_17_2","volume-title":"Proceedings of the CIKM","author":"Li Yang","year":"2021","unstructured":"Yang Li, Tong Chen, Peng-Fei Zhang, and Hongzhi Yin. 2021. Lightweight self-attentive sequential recommendation. In Proceedings of the CIKM."},{"key":"e_1_3_2_18_2","volume-title":"Proceedings of the KDD","author":"Li Zhi","year":"2018","unstructured":"Zhi Li, Hongke Zhao, Qi Liu, Zhenya Huang, Tao Mei, and Enhong Chen. 2018. Learning from history and present: Next-item recommendation via discriminatively exploiting user behaviors. In Proceedings of the KDD."},{"key":"e_1_3_2_19_2","volume-title":"Proceedings of the NeurIPS","author":"Liu Hanxiao","year":"2021","unstructured":"Hanxiao Liu, Zihang Dai, David So, and Quoc V. Le. 2021. Pay attention to MLPs. In Proceedings of the NeurIPS."},{"key":"e_1_3_2_20_2","unstructured":"Langming Liu Liu Cai Chi Zhang Xiangyu Zhao Jingtong Gao Wanyu Wang Yifu Lv Wenqi Fan Yiqi Wang Ming He Zitao Liu and Qing Li. 2023. LinRec: Linear attention mechanism for long-term sequential recommender systems. In Proceedings of the SIGIR."},{"key":"e_1_3_2_21_2","volume-title":"Proceedings of the CIKM","author":"Liu Qidong","year":"2023","unstructured":"Qidong Liu, Fan Yan, Xiangyu Zhao, Zhaocheng Du, Huifeng Guo, Ruiming Tang, and Feng Tian. 2023. Diffusion augmentation for sequential recommendation. In Proceedings of the CIKM."},{"key":"e_1_3_2_22_2","unstructured":"Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013. Efficient estimation ofword representations in vector space. arXiv:1301.3781v3. Retrieved from https:\/\/arxiv.org\/abs\/1301.3781v3"},{"key":"e_1_3_2_23_2","unstructured":"Volodymyr Mnih Nicolas Heess Alex Graves and Koray Kavukcuoglu. 2014. Recurrent models of visual attention. In Proceedings of the NeurIPS."},{"key":"e_1_3_2_24_2","volume-title":"Proceedings of the RecSys","author":"Quadrana Massimo","year":"2017","unstructured":"Massimo Quadrana, Alexandros Karatzoglou, Bal\u00e1zs Hidasi, and Paolo Cremonesi. 2017. Personalizing session-based recommendations with hierarchical recurrent neural networks. In Proceedings of the RecSys."},{"key":"e_1_3_2_25_2","volume-title":"Proceedings of the UAI","author":"Rendle Steffen","year":"2009","unstructured":"Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian personalized ranking from implicit feedback. In Proceedings of the UAI."},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772773"},{"key":"e_1_3_2_27_2","volume-title":"Proceedings of the WACV","author":"Shen Zhuoran","year":"2021","unstructured":"Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li. 2021. Efficient attention: Attention with linear complexities. In Proceedings of the WACV."},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","unstructured":"Jianlin Su Yu Lu Shengfeng Pan Bo Wen and Yunfeng Liu. 2021. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing 568 (2021) 127063.","DOI":"10.1016\/j.neucom.2023.127063"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357895"},{"key":"e_1_3_2_30_2","volume-title":"Proceedings of the AAAI","author":"Tang Chuanxin","year":"2022","unstructured":"Chuanxin Tang, Yucheng Zhao, Guangting Wang, Chong Luo, Wenxuan Xie, and Wenjun Zeng. 2022. Sparse MLP for image recognition: Is self-attention really necessary?. In Proceedings of the AAAI."},{"key":"e_1_3_2_31_2","volume-title":"Proceedings of the SIGIR","author":"Tian Yu","year":"2022","unstructured":"Yu Tian, Jianxin Chang, Yanan Niu, Yang Song, and Chenliang Li. 2022. When multi-level meets multi-interest: A multi-grained neural model for sequential recommendation. In Proceedings of the SIGIR."},{"key":"e_1_3_2_32_2","unstructured":"Ilya O. Tolstikhin Neil Houlsby Alexander Kolesnikov Lucas Beyer Xiaohua Zhai Thomas Unterthiner Jessica Yung Andreas Steiner Daniel Keysers Jakob Uszkoreit Mario Lucic and Alexey Dosovitskiy. 2021. MLP-mixer: An all-MLP architecture for vision. In Proceedings of the NeurIPS."},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","unstructured":"Hugo Touvron Piotr Bojanowski Mathilde Caron Matthieu Cord Alaaeldin El-Nouby Edouard Grave Gautier Izacard Armand Joulin Gabriel Synnaeve Jakob Verbeek and Herv\u00e9 J\u00e9gou. 2022. ResMLP: Feedforward networks for image classification with data-efficient training. IEEE Transactions on Pattern Analysis and Machine Intelligence.","DOI":"10.1109\/TPAMI.2022.3206148"},{"key":"e_1_3_2_34_2","volume-title":"Proceedings of the NeurIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the NeurIPS."},{"key":"e_1_3_2_35_2","volume-title":"Proceedings of the ICLR","author":"Wang Benyou","year":"2020","unstructured":"Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen. 2020. On position embeddings in BERT. In Proceedings of the ICLR."},{"key":"e_1_3_2_36_2","unstructured":"Sinong Wang Belinda Z. Li Madian Khabsa Han Fang and Hao Ma. 2020. Linformer: Self-attention with linear complexity. arXiv:2006.04768v3. Retrieved from https:\/\/arxiv.org\/abs\/2006.04768v3"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","unstructured":"Xin Xia Junliang Yu Qinyong Wang Chaoqun Yang Nguyen Quoc Viet Hung and Hongzhi Yin. 2023. Efficient on-device session-based recommendation. ACM Transactions on Information Systems 41 4 (2023) 1\u201324.","DOI":"10.1145\/3580364"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3543507.3583361"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/2911451.2914683"},{"key":"e_1_3_2_40_2","volume-title":"Proceedings of the SIGIR","author":"Yuan Enming","year":"2022","unstructured":"Enming Yuan, Wei Guo, Zhicheng He, Huifeng Guo, Chengkai Liu, and Ruiming Tang. 2022. Multi-behavior sequential transformer recommender. In Proceedings of the SIGIR."},{"key":"e_1_3_2_41_2","volume-title":"Proceedings of the CIKM","author":"Zhang Chi","year":"2022","unstructured":"Chi Zhang, Yantong Du, Xiangyu Zhao, Qilong Han, Rui Chen, and Li Li. 2022. Hierarchical item inconsistency signal learning for sequence denoising in sequential recommendation. In Proceedings of the CIKM."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/600"},{"key":"e_1_3_2_43_2","volume-title":"Proceedings of the CIKM","author":"Zhao Kesen","year":"2022","unstructured":"Kesen Zhao, Xiangyu Zhao, Zijian Zhang, and Muyang Li. 2022. MAE4Rec: Storage-saving transformer for sequential recommendations. In Proceedings of the CIKM."},{"key":"e_1_3_2_44_2","unstructured":"Wayne Xin Zhao Shanlei Mu Yupeng Hou Zihan Lin Yushuo Chen Xingyu Pan Kaiyuan Li Yujie Lu Hui Wang Changxin Tian Yingqian Min Zhichao Feng Xinyan Fan Xu Chen Pengfei Wang Wendi Ji Yaliang Li Xiaoling Wang and Ji-Rong Wen. 2021. RecBole: Towards a unified comprehensive and efficient framework for recommendation algorithms. In Proceedings of the CIKM."},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","unstructured":"Xiangyu Zhao Long Xia Jiliang Tang and Dawei Yin. 2019. Deep reinforcement learning for search recommendation and online advertising: A survey. ACM SIGWEB Newsletter 2019 (2019) 1\u201315.","DOI":"10.1145\/3320496.3320500"},{"key":"e_1_3_2_46_2","volume-title":"Proceedings of the RecSys","author":"Zhao Xiangyu","year":"2018","unstructured":"Xiangyu Zhao, Long Xia, Liang Zhang, Zhuoye Ding, Dawei Yin, and Jiliang Tang. 2018. Deep reinforcement learning for page-wise recommendations. In Proceedings of the RecSys."},{"key":"e_1_3_2_47_2","volume-title":"Proceedings of the KDD","author":"Zhao Xiangyu","year":"2018","unstructured":"Xiangyu Zhao, Liang Zhang, Zhuoye Ding, Long Xia, Jiliang Tang, and Dawei Yin. 2018. Recommendations with negative feedback via pairwise deep reinforcement learning. In Proceedings of the KDD."},{"key":"e_1_3_2_48_2","unstructured":"Jianqiao Zheng Sameera Ramasinghe and Simon Lucey. 2021. Rethinking positional encoding. arXiv:2107.02561v3. Retrieved from https:\/\/arxiv.org\/abs\/2107.02561v3"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3411954"},{"key":"e_1_3_2_50_2","volume-title":"Proceedings of the WWW","author":"Zhou Kun","year":"2022","unstructured":"Kun Zhou, Hui Yu, Wayne Xin Zhao, and Ji-Rong Wen. 2022. Filter-enhanced MLP is all you need for sequential recommendation. In Proceedings of the WWW."}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3637871","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3637871","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:49:18Z","timestamp":1750286958000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3637871"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,22]]},"references-count":49,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,5,31]]}},"alternative-id":["10.1145\/3637871"],"URL":"https:\/\/doi.org\/10.1145\/3637871","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,22]]},"assertion":[{"value":"2023-01-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-28","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}