{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,24]],"date-time":"2026-02-24T12:54:51Z","timestamp":1771937691842,"version":"3.50.1"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2024,4,29]],"date-time":"2024-04-29T00:00:00Z","timestamp":1714348800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61972337, 61502414"],"award-info":[{"award-number":["61972337, 61502414"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2024,9,30]]},"abstract":"<jats:p>Casting sequential recommendation (SR) as a reinforcement learning (RL) problem is promising and some RL-based methods have been proposed for SR. However, these models are sub-optimal due to the following limitations: (a) they fail to leverage the supervision signals in the RL training to capture users\u2019 explicit preferences, leading to slow convergence; and (b) they do not utilize auxiliary information (e.g., knowledge graph) to avoid blindness when exploring users\u2019 potential interests. To address the above-mentioned limitations, we propose a multiplex information-guided RL model (MELOD), which employs a novel RL training framework with Teach and Explore components for SR. We adopt a Teach component to accurately capture users\u2019 explicit preferences and speed up RL convergence. Meanwhile, we design a dynamic intent induction network (DIIN) as a policy function to generate diverse predictions. We utilize the DIIN for the Explore component to mine users\u2019 potential interests by conducting a sequential and knowledge information joint-guided exploration. Moreover, a sequential and knowledge-aware reward function is designed to achieve stable RL training. These components significantly improve MELOD\u2019s performance and convergence against existing RL algorithms to achieve effectiveness and efficiency. Experimental results on seven real-world datasets show that our model significantly outperforms state-of-the-art methods.<\/jats:p>","DOI":"10.1145\/3630003","type":"journal-article","created":{"date-parts":[[2023,10,23]],"date-time":"2023-10-23T21:23:58Z","timestamp":1698096238000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Teach and Explore: A Multiplex Information-guided Effective and Efficient Reinforcement Learning for Sequential Recommendation"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0976-1159","authenticated-orcid":false,"given":"Surong","family":"Yan","sequence":"first","affiliation":[{"name":"Zhejiang University of Finance and Economics, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3332-0687","authenticated-orcid":false,"given":"Chenglong","family":"Shi","sequence":"additional","affiliation":[{"name":"Zhejiang University of Finance and Economics, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5445-3676","authenticated-orcid":false,"given":"Haosen","family":"Wang","sequence":"additional","affiliation":[{"name":"Southeast University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-7510-5583","authenticated-orcid":false,"given":"Lei","family":"Chen","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-5671-1199","authenticated-orcid":false,"given":"Ling","family":"Jiang","sequence":"additional","affiliation":[{"name":"Zhejiang University of Finance and Economics, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7094-5017","authenticated-orcid":false,"given":"Ruilin","family":"Guo","sequence":"additional","affiliation":[{"name":"Zhejiang University of Finance and Economics, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3124-5487","authenticated-orcid":false,"given":"Kwei-Jay","family":"Lin","sequence":"additional","affiliation":[{"name":"University of California Irvine, USA and Chang Gung University, Taoyuan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,4,29]]},"reference":[{"key":"e_1_3_4_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dss.2015.03.008"},{"issue":"1","key":"e_1_3_4_3_2","first-page":"5:1\u20135:38","article-title":"Deep learning-based recommender system: A survey and new perspectives","volume":"52","author":"Zhang Shuai","year":"2019","unstructured":"Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay. 2019. Deep learning-based recommender system: A survey and new perspectives. ACM Comput. Surv. 52, 1 (2019), 5:1\u20135:38.","journal-title":"ACM Comput. Surv."},{"key":"e_1_3_4_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482016"},{"key":"e_1_3_4_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219846"},{"key":"e_1_3_4_6_2","first-page":"582","volume-title":"Proceedings of the 12th ACM International Conference on Web Search and Data Mining","author":"Yuan Fajie","unstructured":"Fajie Yuan, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. Jose, and Xiangnan He. A simple convolutional generative network for next item recommendation. In Proceedings of the 12th ACM International Conference on Web Search and Data Mining. 582\u2013590."},{"key":"e_1_3_4_7_2","volume-title":"Proceedings of the 4th International Conference on Learning Representations","year":"2016","unstructured":"Bal \\({&#x000E1;}\\) zs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2016. Session-based recommendations with recurrent neural networks. In Proceedings of the 4th International Conference on Learning Representations."},{"key":"e_1_3_4_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3159652.3159656"},{"key":"e_1_3_4_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00035"},{"key":"e_1_3_4_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462968"},{"key":"e_1_3_4_11_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.3301346"},{"key":"e_1_3_4_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3159652.3159668"},{"key":"e_1_3_4_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3109859.3109896"},{"key":"e_1_3_4_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357895"},{"key":"e_1_3_4_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462978"},{"key":"e_1_3_4_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210007"},{"key":"e_1_3_4_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.1998.712192"},{"key":"e_1_3_4_18_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913495721"},{"key":"e_1_3_4_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-26706-7_2"},{"key":"e_1_3_4_20_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_4_21_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature16961"},{"key":"e_1_3_4_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330668"},{"key":"e_1_3_4_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401134"},{"key":"e_1_3_4_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401237"},{"key":"e_1_3_4_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3426723"},{"key":"e_1_3_4_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772773"},{"key":"e_1_3_4_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2016.0030"},{"key":"e_1_3_4_28_2","first-page":"5998","volume-title":"Proceedings of the 30th Annual Conference on Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 30th Annual Conference on Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_4_29_2","first-page":"4171","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171\u20134186."},{"key":"e_1_3_4_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00099"},{"key":"e_1_3_4_31_2","article-title":"Contrastive self-supervised sequential recommendation with robust augmentation","author":"Liu Zhiwei","year":"2021","unstructured":"Zhiwei Liu, Yongjun Chen, Jia Li, Philip S. Yu, Julian McAuley, and Caiming Xiong. 2021. Contrastive self-supervised sequential recommendation with robust augmentation. Retrieved from https:\/\/arXiv:2108.06479","journal-title":"Retrieved from https:\/\/arXiv:2108.06479"},{"key":"e_1_3_4_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488560.3498433"},{"key":"e_1_3_4_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401142"},{"key":"e_1_3_4_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488560.3498524"},{"key":"e_1_3_4_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539342"},{"key":"e_1_3_4_36_2","article-title":"A survey on knowledge graph-based recommender systems","author":"Guo Qingyu","year":"2020","unstructured":"Qingyu Guo, Fuzhen Zhuang, Chuan Qin, Hengshu Zhu, Xing Xie, Hui Xiong, and Qing He. 2020. A survey on knowledge graph-based recommender systems. IEEE Trans. Knowl. Data Eng. 34, 8 (2020), 3549\u20133568.","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"e_1_3_4_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532009"},{"issue":"2","key":"e_1_3_4_38_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3432244","article-title":"Toward dynamic user intention: Temporal evolutionary effects of item relations in sequential recommendation","volume":"39","author":"Wang Chenyang","year":"2020","unstructured":"Chenyang Wang, Weizhi Ma, Min Zhang, Chong Chen, Yiqun Liu, and Shaoping Ma. 2020. Toward dynamic user intention: Temporal evolutionary effects of item relations in sequential recommendation. ACM Trans. Info. Syst. 39, 2 (2020), 1\u201333.","journal-title":"ACM Trans. Info. Syst."},{"key":"e_1_3_4_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532025"},{"key":"e_1_3_4_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3486673"},{"key":"e_1_3_4_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3271739"},{"key":"e_1_3_4_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330989"},{"key":"e_1_3_4_43_2","first-page":"2787","volume-title":"Proceedings of the 27th Annual Conference on Advances in Neural Information Processing Systems","author":"Bordes Antoine","year":"2013","unstructured":"Antoine Bordes, Nicolas Usunier, Alberto Garc\u00eda-Dur\u00e1n, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. In Proceedings of the 27th Annual Conference on Advances in Neural Information Processing Systems. 2787\u20132795."},{"key":"e_1_3_4_44_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v29i1.9491"},{"key":"e_1_3_4_45_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"Sun Zhiqing","year":"2019","unstructured":"Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. 2019. RotatE: Knowledge graph embedding by relational rotation in complex space. In Proceedings of the 7th International Conference on Learning Representations."},{"key":"e_1_3_4_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210017"},{"key":"e_1_3_4_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/2959100.2959167"},{"key":"e_1_3_4_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488560.3498494"},{"key":"e_1_3_4_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462913"},{"key":"e_1_3_4_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401147"},{"key":"e_1_3_4_51_2","first-page":"249","volume-title":"Proceedings of the 13th International Conference on Artificial Intelligence and Statistics","author":"Glorot Xavier","year":"2010","unstructured":"Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics. 249\u2013256."},{"key":"e_1_3_4_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485447.3512111"},{"key":"e_1_3_4_53_2","doi-asserted-by":"publisher","DOI":"10.1162\/dint_a_00008"},{"key":"e_1_3_4_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403226"},{"key":"e_1_3_4_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557464"},{"key":"e_1_3_4_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591656"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3630003","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3630003","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:57:00Z","timestamp":1750291020000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3630003"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,29]]},"references-count":55,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,9,30]]}},"alternative-id":["10.1145\/3630003"],"URL":"https:\/\/doi.org\/10.1145\/3630003","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,29]]},"assertion":[{"value":"2023-01-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-16","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}