{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T10:45:06Z","timestamp":1781779506345,"version":"3.54.5"},"reference-count":53,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2023,7,15]],"date-time":"2023-07-15T00:00:00Z","timestamp":1689379200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,7,15]],"date-time":"2023-07-15T00:00:00Z","timestamp":1689379200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001773","name":"University of New South Wales","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001773","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["World Wide Web"],"published-print":{"date-parts":[[2023,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Deep reinforcement learning (DRL) has shown promising results in modeling dynamic user preferences in RS in recent literature. However, training a DRL agent in the sparse RS environment poses a significant challenge. This is because the agent must balance between exploring informative user-item interaction trajectories and using existing trajectories for policy learning, a known exploration and exploitation trade-off. This trade-off greatly affects the recommendation performance when the environment is sparse. In DRL-based RS, balancing exploration and exploitation is even more challenging as the agent needs to deeply explore informative trajectories and efficiently exploit them in the context of RS. To address this issue, we propose a novel intrinsically motivated reinforcement learning (IMRL) method that enhances the agent\u2019s capability to explore informative interaction trajectories in the sparse environment. We further enrich these trajectories via an adaptive counterfactual augmentation strategy with a customised threshold to improve their efficiency in exploitation. Our approach is evaluated on six offline datasets and three online simulation platforms, demonstrating its superiority over existing state-of-the-art methods. The extensive experiments show that our IMRL method outperforms other methods in terms of recommendation performance in the sparse RS environment.<\/jats:p>","DOI":"10.1007\/s11280-023-01187-7","type":"journal-article","created":{"date-parts":[[2023,7,15]],"date-time":"2023-07-15T14:01:56Z","timestamp":1689429716000},"page":"3253-3274","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Intrinsically motivated reinforcement learning based recommendation with counterfactual data augmentation"],"prefix":"10.1007","volume":"26","author":[{"given":"Xiaocong","family":"Chen","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Siyu","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lianyong","family":"Qi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yong","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lina","family":"Yao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,7,15]]},"reference":[{"key":"1187_CR1","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2023.110335","volume":"264","author":"X Chen","year":"2023","unstructured":"Chen, X., Yao, L., McAuley, J., Zhou, G., Wang, X.: Deep reinforcement learning in recommender systems: A survey and new perspectives. Knowl. Based Syst. 264, 110335 (2023)","journal-title":"Knowl. Based Syst."},{"key":"1187_CR2","doi-asserted-by":"crossref","unstructured":"Zheng, G., Zhang, F., Zheng, Z., Xiang, Y., Yuan, N.J., Xie, X., Li, Z.: Drn: A deep reinforcement learning framework for news recommendation. In: Proceedings of the 2018 World Wide Web Conference, 167\u2013176 (2018)","DOI":"10.1145\/3178876.3185994"},{"key":"1187_CR3","unstructured":"Dulac-Arnold, G., Evans, R., van Hasselt, H., Sunehag, P., Lillicrap, T., Hunt, J., Mann, T., Weber, T., Degris, T., Coppin, B.: Deep reinforcement learning in large discrete action spaces. arXiv:1512.07679 (2015)"},{"key":"1187_CR4","doi-asserted-by":"crossref","unstructured":"Xu, J., Wei, Z., Xia, L., Lan, Y., Yin, D., Cheng, X., Wen, J.-R.: Reinforcement learning to rank with pairwise policy gradient. In: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, 509\u2013518 (2020)","DOI":"10.1145\/3397271.3401148"},{"key":"1187_CR5","unstructured":"Degris, T., White, M., Sutton, R.S.: Off-policy actor-critic. arXiv:1205.4839 (2012)"},{"key":"1187_CR6","doi-asserted-by":"crossref","unstructured":"Chen, X., Huang, C., Yao, L., Wang, X., Zhang, W., etal: Knowledge-guided deep reinforcement learning for interactive recommendation. In: 2020 International Joint Conference on Neural Networks (IJCNN), 1\u20138 (2020). IEEE","DOI":"10.1109\/IJCNN48605.2020.9207010"},{"key":"1187_CR7","doi-asserted-by":"crossref","unstructured":"Chen, H., Dai, X., Cai, H., Zhang, W., Wang, X., Tang, R., Zhang, Y., Yu, Y.: Large-scale interactive recommendation with tree-structured policy gradient. In: Proceedings of the AAAI Conference on Artificial Intelligence, 33, 3312\u20133320 (2019)","DOI":"10.1609\/aaai.v33i01.33013312"},{"key":"1187_CR8","doi-asserted-by":"crossref","unstructured":"Cai, Q., Filos-Ratsikas, A., Tang, P., Zhang, Y.: Reinforcement mechanism design for e-commerce. In: Proceedings of the 2018 World Wide Web Conference, 1339\u20131348 (2018)","DOI":"10.1145\/3178876.3186039"},{"key":"1187_CR9","doi-asserted-by":"crossref","unstructured":"Chen, X., Yao, L., Sun, A., Wang, X., Xu, X., Zhu, L.: Generative inverse deep reinforcement learning for online recommendation. In: Proceedings of the 30th ACM International Conference on Information & Knowledge Management, 201\u2013210 (2021)","DOI":"10.1145\/3459637.3482347"},{"key":"1187_CR10","doi-asserted-by":"crossref","unstructured":"Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., Abbeel, P.: Overcoming exploration in reinforcement learning with demonstrations. In: 2018 IEEE International Conference on Robotics and Automation (ICRA), 6292\u20136299 (2018). IEEE","DOI":"10.1109\/ICRA.2018.8463162"},{"key":"1187_CR11","doi-asserted-by":"crossref","unstructured":"Chen, M., Wang, Y., Xu, C., Le, Y., Sharma, M., Richardson, L., Wu S.-L., Chi, E.: Values of user exploration in recommender systems. In: Fifteenth ACM Conference on Recommender Systems, 85\u201395 (2021)","DOI":"10.1145\/3460231.3474236"},{"key":"1187_CR12","doi-asserted-by":"crossref","unstructured":"Chen, M.: Exploration in recommender systems. In: Fifteenth ACM Conference on Recommender Systems, pp. 551-553 (2021)","DOI":"10.1145\/3460231.3474601"},{"key":"1187_CR13","unstructured":"Schaul, T., Quan, J., Antonoglou, I., Silver, D. (2015) Prioritized experience replay. arXiv:1511.05952"},{"key":"1187_CR14","doi-asserted-by":"crossref","unstructured":"Wang, Z., Zhang, J., Xu, H., Chen, X., Zhang, Y., Zhao, W.X.,Wen, J.-R.: Counterfactual data-augmented sequential recommendation. In: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 347\u2013356 (2021)","DOI":"10.1145\/3404835.3462855"},{"key":"1187_CR15","doi-asserted-by":"crossref","unstructured":"Zhang, S., Yao, D., Zhao, Z., Chua, T.-S., Wu, F.: Causerec: Counterfactual user sequence synthesis for sequential recommendation. In: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 367\u2013377 (2021)","DOI":"10.1145\/3404835.3462908"},{"key":"1187_CR16","doi-asserted-by":"crossref","unstructured":"Chen, X., Yao, L., McAuley, J., Guan, W., Chang, X.,Wang, X.: Localitysensitive state-guided experience replay optimization for sparse rewards in online recommendation. In: Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, 1316\u20131325 (2022)","DOI":"10.1145\/3477495.3532015"},{"key":"1187_CR17","doi-asserted-by":"crossref","unstructured":"Chen, X., Yao, L., Chang, X., Wang, S.: Empowerment-driven policy gradient learning with counterfactual augmentation in recommender systems. In: 2022 IEEE International Conference on Data Mining (ICDM), 885\u2013890 (2022). IEEE","DOI":"10.1109\/ICDM54844.2022.00102"},{"key":"1187_CR18","unstructured":"Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R., Smola, A.: Deep sets. arXiv:1703.06114 (2017)"},{"key":"1187_CR19","doi-asserted-by":"publisher","first-page":"96","DOI":"10.1214\/09-SS057","volume":"3","author":"J Pearl","year":"2009","unstructured":"Pearl, J.: Causal inference in statistics: An overview. Stat. Surv. 3, 96\u2013146 (2009)","journal-title":"Stat. Surv."},{"key":"1187_CR20","unstructured":"Pitis, S., Creager, E., Garg, A.: Counterfactual data augmentation using locally factored dynamics. arXiv:2007.02863 (2020)"},{"key":"1187_CR21","volume-title":"Elements of Causal Inference - Foundations and Learning Algorithms","author":"J Peters","year":"2017","unstructured":"Peters, J., Janzing, D., Sch\u00f6lkopf, B.: Elements of Causal Inference - Foundations and Learning Algorithms. Adaptive Computation and Machine Learning Series. The MIT Press, Cambridge, MA, USA (2017)"},{"key":"1187_CR22","unstructured":"Lu, C., Huang, B., Wang, K., Hern\u00e1ndez-Lobato, J.M., Zhang, K., Sch\u00f6lkopf, B.: Sample-efficient reinforcement learning via counterfactual-based data augmentation. arXiv:2012.09092 (2020)"},{"issue":"9","key":"1187_CR23","doi-asserted-by":"publisher","first-page":"1591","DOI":"10.1109\/TCYB.2013.2290775","volume":"44","author":"Y Xiang","year":"2013","unstructured":"Xiang, Y., Truong, M.: Acquisition of causal models for local distributions in bayesian networks. IEEE Trans. Cybern. 44(9), 1591\u20131604 (2013)","journal-title":"IEEE Trans. Cybern."},{"issue":"1","key":"1187_CR24","doi-asserted-by":"publisher","first-page":"14","DOI":"10.1109\/TIT.1972.1054753","volume":"18","author":"S Arimoto","year":"1972","unstructured":"Arimoto, S.: An algorithm for computing the capacity of arbitrary discrete memoryless channels. IEEE Trans. Inf. Theory 18(1), 14\u201320 (1972)","journal-title":"IEEE Trans. Inf. Theory"},{"issue":"4","key":"1187_CR25","doi-asserted-by":"publisher","first-page":"460","DOI":"10.1109\/TIT.1972.1054855","volume":"18","author":"R Blahut","year":"1972","unstructured":"Blahut, R.: Computation of channel capacity and rate-distortion functions. IEEE Trans. Inf. Theory 18(4), 460\u2013473 (1972)","journal-title":"IEEE Trans. Inf. Theory"},{"key":"1187_CR26","first-page":"2125","volume":"28","author":"S Mohamed","year":"2015","unstructured":"Mohamed, S., Jimenez Rezende, D.: Variational information maximisation for intrinsically motivated reinforcement learning. Adv. Neural Inf. Process. Syst. 28, 2125\u20132133 (2015)","journal-title":"Adv. Neural Inf. Process. Syst."},{"issue":"1","key":"1187_CR27","doi-asserted-by":"publisher","first-page":"16","DOI":"10.1177\/1059712310392389","volume":"19","author":"T Jung","year":"2011","unstructured":"Jung, T., Polani, D., Stone, P.: Empowerment for continuous agent-environment systems. Adapt. Behav. 19(1), 16\u201339 (2011)","journal-title":"Adapt. Behav."},{"key":"1187_CR28","first-page":"7869","volume":"32","author":"F Leibfried","year":"2019","unstructured":"Leibfried, F., Pascual-D\u00edaz, S., Grau-Moya, J.: A unified bellman optimality principle combining reward maximization and empowerment. Adv. Neural Inf. Process. Syst. 32, 7869\u20137880 (2019)","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"1187_CR29","unstructured":"Kumar, N.M.: Empowerment-driven exploration using mutual information estimation. arXiv:1810.05533 (2018)"},{"key":"1187_CR30","doi-asserted-by":"publisher","unstructured":"Elements of Information Theory. John Wiley & Sons, Ltd. https:\/\/doi.org\/10.1002\/0471200611","DOI":"10.1002\/0471200611"},{"key":"1187_CR31","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., Levine, S.: Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In: International Conference on Machine Learning, pp. 1861\u20131870 (2018). PMLR"},{"key":"1187_CR32","unstructured":"Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al.: Soft actor-critic algorithms and applications. arXiv:1812.05905 (2018)"},{"issue":"7540","key":"1187_CR33","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., Graves, A., Riedmiller, M., Fidjeland, A.K., Ostrovski, G., et al.: Human-level control through deep reinforcement learning. Nature 518(7540), 529\u2013533 (2015)","journal-title":"Nature"},{"key":"1187_CR34","unstructured":"Chen, X., Li, S., Li, H., Jiang, S., Qi, Y., Song, L.: Generative adversarial user model for reinforcement learning based recommendation system. In: International Conference on Machine Learning, pp. 1052\u20131061 (2019). PMLR"},{"key":"1187_CR35","doi-asserted-by":"crossref","unstructured":"Kang, W.-C., McAuley, J.: Self-attentive sequential recommendation. In: 2018 IEEE International Conference on Data Mining (ICDM), pp. 197\u2013206 (2018). IEEE","DOI":"10.1109\/ICDM.2018.00035"},{"key":"1187_CR36","doi-asserted-by":"crossref","unstructured":"Zhao, X., Zhang, L., Ding, Z., Xia, L., Tang, J., Yin, D.: Recommendations with negative feedback via pairwise deep reinforcement learning. In: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 1040\u20131048 (2018)","DOI":"10.1145\/3219819.3219886"},{"key":"1187_CR37","doi-asserted-by":"crossref","unstructured":"Cai, R., Wu, J., San, A., Wang, C., Wang, H.: Category-aware collaborative sequential recommendation. In: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 388\u2013397 (2021)","DOI":"10.1145\/3404835.3462832"},{"key":"1187_CR38","doi-asserted-by":"crossref","unstructured":"Mu, S., Li, Y., Zhao, W.X., Wang, J., Ding, B., Wen, J.-R.: Alleviating spurious correlations in knowledge-aware recommendations through counterfactual generator. In: Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, 1401\u20131411 (2022)","DOI":"10.1145\/3477495.3531934"},{"key":"1187_CR39","doi-asserted-by":"crossref","unstructured":"Xian, Y., Fu, Z., Muthukrishnan, S., de Melo, G., Zhang, Y.: Reinforcement knowledge graph reasoning for explainable recommendation. In: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 285\u2013294 (2019). ACM","DOI":"10.1145\/3331184.3331203"},{"key":"1187_CR40","doi-asserted-by":"crossref","unstructured":"Shi, J.-C., Yu, Y., Da, Q., Chen, S.-Y., Zeng, A.-X.: Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, pp. 4902-4909 (2019)","DOI":"10.1609\/aaai.v33i01.33014902"},{"key":"1187_CR41","unstructured":"Ie, E., Hsu, C.-w., Mladenov, M., Jain, V., Narvekar, S., Wang, J., Wu, R., Boutilier, C.: Recsim: A configurable simulation platform for recommender systems. arXiv:1909.04847 (2019) [cs.LG]IMRL"},{"key":"1187_CR42","unstructured":"Rohde, D., Bonner, S., Dunlop, T., Vasile, F., Karatzoglou, A.: Recogym: A reinforcement learning environment for the problem of product recommendation in online advertising. arXiv:1808.00720 (2018)"},{"key":"1187_CR43","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al.: Pytorch: An imperative style, high-performance deep learning library. In: Advances in Neural Information Processing Systems, pp. 8026\u20138037 (2019)"},{"key":"1187_CR44","doi-asserted-by":"crossref","unstructured":"Chen, M., Beutel, A., Covington, P., Jain, S., Belletti, F., Chi, E.H.: Top-k off-policy correction for a reinforce recommender system. In: Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining, 456\u2013464 (2019)","DOI":"10.1145\/3289600.3290999"},{"key":"1187_CR45","doi-asserted-by":"crossref","unstructured":"Wei, T., Feng, F., Chen, J.,Wu, Z., Yi, J., He, X.: Model-agnostic counterfactual reasoning for eliminating popularity bias in recommender system. In: Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, 1791\u20131800 (2021)","DOI":"10.1145\/3447548.3467289"},{"key":"1187_CR46","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Feng, F., He, X., Wei, T., Song, C., Ling, G., Zhang, Y.: Causal intervention for leveraging popularity bias in recommendation. arXiv:2105.06067 (2021)","DOI":"10.1145\/3404835.3462875"},{"key":"1187_CR47","doi-asserted-by":"crossref","unstructured":"Gershman, S.J.: Reinforcement learning and causal models. The Oxford handbook of causal reasoning, 295 (2017)","DOI":"10.1093\/oxfordhb\/9780199399550.013.20"},{"key":"1187_CR48","unstructured":"Zhu, S., Ng, I., Chen, Z.: Causal discovery with reinforcement learning. arXiv:1906.04477 (2019)"},{"key":"1187_CR49","unstructured":"Dasgupta, I., Wang, J., Chiappa, S., Mitrovic, J., Ortega, P., Raposo, D., Hughes, E., Battaglia, P., Botvinick, M., Kurth-Nelson, Z.: Causal reasoning from meta-reinforcement learning. arXiv:1901.08162 (2019)"},{"key":"1187_CR50","unstructured":"Zhang, J., Kumor, D., Bareinboim, E.: Causal imitation learning with unobserved confounders. Advances in neural information processing systems 33 (2020)"},{"key":"1187_CR51","doi-asserted-by":"crossref","unstructured":"Madumal, P., Miller, T., Sonenberg, L., Vetere, F.: Explainable reinforcement learning through a causal lens. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, pp. 2493\u20132500 (2020)","DOI":"10.1609\/aaai.v34i03.5631"},{"key":"1187_CR52","doi-asserted-by":"crossref","unstructured":"Ji, J., Li, Z., Xu, S., Xiong, M., Tan, J., Ge, Y., Wang, H., Zhang, Y.: Counterfactual collaborative reasoning. In: Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining, pp. 249\u2013257 (2023)","DOI":"10.1145\/3539597.3570464"},{"issue":"3","key":"1187_CR53","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3559757","volume":"41","author":"X He","year":"2023","unstructured":"He, X., Zhang, Y., Feng, F., Song, C., Yi, L., Ling, G., Zhang, Y.: Addressing confounding feature issue for causal recommendation. ACM Trans. Info. Syst. 41(3), 1\u201323 (2023)","journal-title":"ACM Trans. Info. Syst."}],"container-title":["World Wide Web"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11280-023-01187-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11280-023-01187-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11280-023-01187-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,11]],"date-time":"2023-10-11T04:24:46Z","timestamp":1696998286000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11280-023-01187-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,15]]},"references-count":53,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2023,9]]}},"alternative-id":["1187"],"URL":"https:\/\/doi.org\/10.1007\/s11280-023-01187-7","relation":{},"ISSN":["1386-145X","1573-1413"],"issn-type":[{"value":"1386-145X","type":"print"},{"value":"1573-1413","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,15]]},"assertion":[{"value":"16 January 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 April 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 June 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 July 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"None","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}},{"value":"Not Applicable","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Approval"}}]}}