{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T05:13:35Z","timestamp":1784178815517,"version":"3.55.0"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62376275, 62472426, 62502091"],"award-info":[{"award-number":["62376275, 62472426, 62502091"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Key Laboratory of Research on Large Models and Intelligent Governance"},{"name":"Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE"},{"name":"Central Universities in UIBE","award":["24QN06, 24PYTS22"],"award-info":[{"award-number":["24QN06, 24PYTS22"]}]},{"DOI":"10.13039\/501100004260","name":"Renmin University of China","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004260","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>In real-world streaming recommender systems, user preferences evolve dynamically over time. Existing bandit-based methods treat time merely as a timestamp, neglecting its explicit relationship with user preferences and leading to suboptimal performance. Moreover, the online learning methods often suffer from inefficient exploration\u2013exploitation during the early online phase. To address these issues, we propose HyperBandit+, a novel contextual bandit policy which integrates a time-aware hypernetwork to adapt to time-varying user preferences and employs a large language model-assisted warm-start mechanism (LLM Start) to enhance exploration\u2013exploitation efficiency at the early online phase. Specifically, HyperBandit+ leverages a neural network that takes time features as input and generates parameters for estimating time-varying rewards by capturing the correlation between time and user preferences. Additionally, the LLM Start mechanism employs multi-step data augmentation to simulate realistic interaction data for effective offline learning, providing warm-start parameters for the bandit policy at the early online phase. To meet real-time streaming recommendation demands, we adopt low-rank factorization to reduce hypernetwork training complexity. Theoretically, we rigorously establish a sublinear regret upper bound that accounts for both the hypernetwork and the LLM warm-start mechanism. Extensive experiments on real-world datasets demonstrate that HyperBandit+ consistently outperforms state-of-the-art baselines in terms of accumulated rewards.<\/jats:p>","DOI":"10.1145\/3793543","type":"journal-article","created":{"date-parts":[[2026,2,14]],"date-time":"2026-02-14T14:27:43Z","timestamp":1771079263000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Enhancing Bandit Algorithms with LLMs for Time-varying User Preferences in Streaming Recommendations"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-3567-8071","authenticated-orcid":false,"given":"Chenglei","family":"Shen","sequence":"first","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7454-4384","authenticated-orcid":false,"given":"Yi","family":"Zhan","sequence":"additional","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5676-4339","authenticated-orcid":false,"given":"Weijie","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence and Data Science, University of International Business and Economics, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7397-5632","authenticated-orcid":false,"given":"Xiao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7170-111X","authenticated-orcid":false,"given":"Jun","family":"Xu","sequence":"additional","affiliation":[{"name":"Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,3,25]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Ali Baheri and Cecilia O. Alm. 2023. LLMs-augmented contextual bandit. arXiv:2311.02268. Retrieved from https:\/\/arxiv.org\/abs\/2311.02268"},{"key":"e_1_3_2_3_2","first-page":"9679","article-title":"Contextual bandits with cross-learning","volume":"32","author":"Balseiro Santiago","year":"2019","unstructured":"Santiago Balseiro, Negin Golrezaei, Mohammad Mahdian, Vahab Mirrokni, and Jon Schneider. 2019. Contextual bandits with cross-learning. In Advances in Neural Information Processing Systems 32, 9679\u20139688.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2010.07.024"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1287\/stsy.2019.0033"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1561\/2200000024","article-title":"Regret analysis of stochastic and nonstochastic multi-armed bandit problems","volume":"5","author":"Bubeck S\u00e9bastien","year":"2012","unstructured":"S\u00e9bastien Bubeck and Nicol\u00f2 Cesa-Bianchi. 2012. Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Journal of Foundations and Trends in Machine Learning 5 (2012), 1\u2013122.","journal-title":"Journal of Foundations and Trends in Machine Learning"},{"key":"e_1_3_2_7_2","first-page":"129","volume-title":"Proceedings of 24th International Conference on Machine Learning","author":"Cao Zhe","year":"2007","unstructured":"Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. 2007. Learning to rank: From pairwise approach to listwise approach. In Proceedings of 24th International Conference on Machine Learning, 129\u2013136."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511546921"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","first-page":"1243","DOI":"10.1145\/1989323.1989465","volume-title":"Proceedings of the 2011 ACM SIGMOD International Conference on Management of Data","author":"Chandramouli Badrish","year":"2011","unstructured":"Badrish Chandramouli, Justin J. Levandoski, Ahmed Eldawy, and Mohamed F. Mokbel. 2011. StreamRec: A real-time recommender system. In Proceedings of the 2011 ACM SIGMOD International Conference on Management of Data, 1243\u20131246."},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1145\/3038912.3052627","volume-title":"Proceedings of 26th International Conference on World Wide Web","author":"Chang Shiyu","year":"2017","unstructured":"Shiyu Chang, Yang Zhang, Jiliang Tang, Dawei Yin, Yi Chang, Mark A. Hasegawa-Johnson, and Thomas S. Huang. 2017. Streaming recommender systems. In Proceedings of 26th International Conference on World Wide Web, 381\u2013389."},{"key":"e_1_3_2_11_2","first-page":"75","volume-title":"Proceedings of 15th ACM International Conference on Web Search and Data Mining","author":"Chen Hanxiong","year":"2022","unstructured":"Hanxiong Chen, Yunqi Li, Shaoyun Shi, Shuchang Liu, He Zhu, and Yongfeng Zhang. 2022. Graph collaborative reasoning. In Proceedings of 15th ACM International Conference on Web Search and Data Mining, 75\u201384."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3580305.3599796"},{"key":"e_1_3_2_13_2","first-page":"1079","volume-title":"Proceedings of 22nd International Conference on Artificial Intelligence and Statistics","author":"Chi Cheung Wang","year":"2019","unstructured":"Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu. 2019. Learning to optimize under non-stationarity. In Proceedings of 22nd International Conference on Artificial Intelligence and Statistics. PMLR, 1079\u20131087."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCI.2015.2471196"},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","first-page":"2010619","DOI":"10.1007\/s11704-025-41366-5","article-title":"A survey of controllable learning: Methods and applications in information retrieval","volume":"20","author":"Shen Chenglei","year":"2026","unstructured":"Chenglei Shen, Xiao Zhang, Teng Shi, Changshuo Zhang, Guofu Xie, Jun Xu, Ming He, and Jianping Fan. 2026. A survey of controllable learning: Methods and applications in information retrieval. Frontiers of Computer Science 20 (2026), 2010619.","journal-title":"Frontiers of Computer Science"},{"key":"e_1_3_2_16_2","unstructured":"Wenqi Fan Zihuai Zhao Jiatong Li Yunqing Liu Xiaowei Mei Yiqi Wang Jiliang Tang and Qing Li. 2023. Recommender systems in the era of large language models (LLMs). arXiv:2307.02046. Retrieved from https:\/\/arxiv.org\/abs\/2307.02046"},{"key":"e_1_3_2_17_2","unstructured":"Tomer Galanti and Lior Wolf. 2020. On the modularity of hypernetworks. arXiv:2002.10006. Retrieved from https:\/\/arxiv.org\/abs\/2002.10006"},{"key":"e_1_3_2_18_2","first-page":"540","volume-title":"Proceedings of 31st ACM International Conference on Information & Knowledge Management","author":"Gao Chongming","year":"2022","unstructured":"Chongming Gao, Shijun Li, Wenqiang Lei, Jiawei Chen, Biao Li, Peng Jiang, Xiangnan He, Jiaxin Mao, and Tat-Seng Chua. 2022. KuaiRec: A fully-observed dataset and insights for evaluating recommender systems. In Proceedings of 31st ACM International Conference on Information & Knowledge Management, 540\u2013550."},{"key":"e_1_3_2_19_2","first-page":"93","volume-title":"Proceedings of 7th ACM Conference on Recommender Systems","author":"Gao Huiji","year":"2013","unstructured":"Huiji Gao, Jiliang Tang, Xia Hu, and Huan Liu. 2013. Exploring temporal effects for location recommendation on location-based social networks. In Proceedings of 7th ACM Conference on Recommender Systems, 93\u2013100."},{"key":"e_1_3_2_20_2","unstructured":"Aur\u00e9lien Garivier and Eric Moulines. 2008. On upper-confidence bound policies for non-stationary bandit problems. arXiv:0805.3415. Retrieved from https:\/\/arxiv.org\/abs\/0805.3415"},{"key":"e_1_3_2_21_2","unstructured":"David Ha Andrew Dai and Quoc V. Le. 2016. Hypernetworks. arXiv:1609.09106. Retrieved from https:\/\/arxiv.org\/abs\/1609.09106"},{"key":"e_1_3_2_22_2","unstructured":"Yanjun Han Zhengqing Zhou Zhengyuan Zhou Jose Blanchet Peter W. Glynn and Yinyu Ye. 2020. Sequential batch learning in finite-action linear contextual bandits. arXiv:2004.06321. Retrieved from https:\/\/arxiv.org\/abs\/2004.06321"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.5555\/2832747.2832852"},{"key":"e_1_3_2_24_2","unstructured":"C. Hartland S. Gelly N. Baskiotis O. Teytaud and M. Sebag. 2006. Multi-armed bandit dynamic environments and meta-bandits. arXiv:2505.13355v1. Retrieved from https:\/\/arxiv.org\/abs\/2505.13355v1"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1561\/2400000013"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","first-page":"156","DOI":"10.1016\/j.ins.2020.05.071","article-title":"New technique to alleviate the cold start problem in recommender systems using information from social media and random decision forests","volume":"536","author":"Herce-Zelaya Julio","year":"2020","unstructured":"Julio Herce-Zelaya, Carlos Porcel, Juan Bernab\u00e9-Moreno, A. Tejeda-Lorente, and Enrique Herrera-Viedma. 2020. New technique to alleviate the cold start problem in recommender systems using information from social media and random decision forests. Information Sciences 536 (2020), 156\u2013170.","journal-title":"Information Sciences"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2018.03.019"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2020.113685"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00035"},{"key":"e_1_3_2_30_2","first-page":"19971","article-title":"Efficient frameworks for generalized low-rank matrix bandit problems","volume":"35","author":"Kang Yue","year":"2022","unstructured":"Yue Kang, Cho-Jui Hsieh, and Thomas Chun Man Lee. 2022. Efficient frameworks for generalized low-rank matrix bandit problems. In Advances in Neural Information Processing Systems 35, 19971\u201319983.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_31_2","first-page":"496","volume-title":"Proceedings of 28th International Conference on Artificial Neural Networks","author":"Klocek Sylwester","year":"2019","unstructured":"Sylwester Klocek, \u0141ukasz Maziarka, Maciej Wo\u0142czyk, Jacek Tabor, Jakub Nowak, and Marek \u015amieja. 2019. Hypernetwork functional image representation. In Proceedings of 28th International Conference on Artificial Neural Networks, 496\u2013510."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics11010141"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772758"},{"key":"e_1_3_2_34_2","unstructured":"Ruyu Li Wenhao Deng Yu Cheng Zheng Yuan Jiaqi Zhang and Fajie Yuan. 2023. Exploring the upper limits of text-based collaborative filtering using large language models: Discoveries and insights. arXiv:2305.11700. Retrieved from https:\/\/arxiv.org\/abs\/2305.11700"},{"key":"e_1_3_2_35_2","unstructured":"Qijiong Liu Nuo Chen Tetsuya Sakai and Xiao-Ming Wu. 2023. A first look at LLM-powered generative news recommendation. arXiv:2305.06566. Retrieved from https:\/\/arxiv.org\/abs\/2305.06566"},{"key":"e_1_3_2_36_2","first-page":"460","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics","author":"Lu Yangyi","year":"2021","unstructured":"Yangyi Lu, Amirhossein Meisami, and Ambuj Tewari. 2021. Low-rank generalized linear bandit problems. In Proceedings of the International Conference on Artificial Intelligence and Statistics. PMLR, 460\u2013468."},{"key":"e_1_3_2_37_2","unstructured":"Hanjia Lyu Song Jiang Hanqing Zeng Yinglong Xia and Jiebo Luo. 2023. LLM-Rec: Personalized recommendation via prompting large language models. arXiv:2307.15780. Retrieved from https:\/\/arxiv.org\/abs\/2307.15780"},{"key":"e_1_3_2_38_2","unstructured":"Matthew MacKay Paul Vicol Jon Lorraine David Duvenaud and Roger Grosse. 2019. Self-tuning networks: Bilevel optimization of hyperparameters using structured best-response functions. arXiv:1903.03088. Retrieved from https:\/\/arxiv.org\/abs\/1903.03088"},{"key":"e_1_3_2_39_2","unstructured":"Eliya Nachmani and Lior Wolf. 2019. Hyper-graph-network decoders for block codes. arXiv:1909.09036. Retrieved from https:\/\/arxiv.org\/abs\/1909.09036"},{"key":"e_1_3_2_40_2","volume-title":"Proceedings of 9th International Conference on Learning Representations","author":"Navon Aviv","year":"2021","unstructured":"Aviv Navon, Aviv Shamsian, Ethan Fetaya, and Gal Chechik. 2021. Learning the pareto front with hypernetworks. In Proceedings of 9th International Conference on Learning Representations."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_42_2","article-title":"Random features for large-scale kernel machines","volume":"20","author":"Rahimi Ali","year":"2007","unstructured":"Ali Rahimi and Benjamin Recht. 2007. Random features for large-scale kernel machines. In Advances in Neural Information Processing Systems, Vol. 20 (2007).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_43_2","unstructured":"Steffen Rendle Christoph Freudenthaler Zeno Gantner and Lars Schmidt-Thieme. 2012. BPR: Bayesian personalized ranking from implicit feedback. arXiv:1205.2618. Retrieved from https:\/\/arxiv.org\/abs\/1205.2618"},{"key":"e_1_3_2_44_2","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1561\/2200000018","article-title":"Online learning and online convex optimization","volume":"4","author":"Shalev-Shwartz Shai","year":"2011","unstructured":"Shai Shalev-Shwartz. 2011. Online learning and online convex optimization. Journal of Foundations and Trends in Machine Learning 4 (2011), 107\u2013194.","journal-title":"Journal of Foundations and Trends in Machine Learning"},{"key":"e_1_3_2_45_2","first-page":"9489","volume-title":"Proceedings of 38th International Conference on Machine Learning","author":"Shamsian Aviv","year":"2021","unstructured":"Aviv Shamsian, Aviv Navon, Ethan Fetaya, and Gal Chechik. 2021. Personalized federated learning using hypernetworks. In Proceedings of 38th International Conference on Machine Learning, 9489\u20139502."},{"key":"e_1_3_2_46_2","first-page":"2239","volume-title":"Proceedings of 32nd ACM International Conference on Information and Knowledge Management","author":"Shen Chenglei","year":"2023","unstructured":"Chenglei Shen, Xiao Zhang, Wei, and Jun Xu. 2023. HyperBandit: Contextual bandit with hypernetwork for time-varying user preferences in streaming recommendation. In Proceedings of 32nd ACM International Conference on Information and Knowledge Management, 2239\u20132248."},{"key":"e_1_3_2_47_2","first-page":"343","volume-title":"Proceedings of 21st Annual Conference on Learning Theory (COLT)","author":"Slivkins Aleksandrs","year":"2008","unstructured":"Aleksandrs Slivkins and Eli Upfal. 2008. Adapting to a changing environment: The Brownian restless bandits. In Proceedings of 21st Annual Conference on Learning Theory (COLT), 343\u2013354."},{"key":"e_1_3_2_48_2","first-page":"3267","article-title":"Language modeling with recurrent highway hypernetworks","volume":"30","author":"Suarez Joseph","unstructured":"Joseph Suarez. 2017. Language modeling with recurrent highway hypernetworks. In Advances in Neural Information Processing Systems, Vol. 30, 3267\u20133276.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531889"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10489-022-03542-z"},{"key":"e_1_3_2_51_2","volume-title":"Proceedings of 8th International Conference on Learning Representations","author":"von Oswald Johannes","year":"2020","unstructured":"Johannes von Oswald, Christian Henning, Jo\u00e3o Sacramento, and Benjamin F. Grewe. 2020. Continual learning with hypernetworks. In Proceedings of 8th International Conference on Learning Representations."},{"key":"e_1_3_2_52_2","first-page":"380","volume-title":"Proceedings of 15th Knowledge Management in Organizations","author":"Wan Yongquan","year":"2021","unstructured":"Yongquan Wan, Junli Xian, and Cairong Yan. 2021. A contextual multi-armed bandit approach based on implicit feedback for online recommendation. In Proceedings of 15th Knowledge Management in Organizations, 380\u2013392."},{"key":"e_1_3_2_53_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Wang Huazheng","year":"2017","unstructured":"Huazheng Wang, Qingyun Wu, and Hongning Wang. 2017. Factorization bandits for interactive recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_2_54_2","unstructured":"Lei Wang Jingsen Zhang Xu Chen Yankai Lin Ruihua Song Wayne Xin Zhao and Ji-Rong Wen. 2023. RecAgent: A novel simulation paradigm for recommender systems. arXiv:2306.02552. Retrieved from https:\/\/arxiv.org\/abs\/2306.02552"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3220004"},{"key":"e_1_3_2_56_2","first-page":"525","volume-title":"Proceedings of 41st International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Wang Weiqing","year":"2018","unstructured":"Weiqing Wang, Hongzhi Yin, Zi Huang, Qinyong Wang, Xingzhong Du, and Quoc Viet Hung Nguyen. 2018. Streaming ranking based recommender systems. In Proceedings of 41st International ACM SIGIR Conference on Research and Development in Information Retrieval, 525\u2013534."},{"key":"e_1_3_2_57_2","first-page":"1120","volume-title":"Proceedings of 15th ACM International Conference on Web Search and Data Mining","author":"Wei Wei","year":"2022","unstructured":"Wei Wei, Chao Huang, Lianghao Xia, Yong Xu, Jiashu Zhao, and Dawei Yin. 2022. Contrastive meta learning with behavior multiplicity for recommendation. In Proceedings of 15th ACM International Conference on Web Search and Data Mining, 1120\u20131128."},{"key":"e_1_3_2_58_2","unstructured":"Wei Wei Xubin Ren Jiabin Tang Qinyong Wang Lixin Su Suqi Cheng Junfeng Wang Dawei Yin and Chao Huang. 2023. LLMRec: Large language models with graph augmentation for recommendation. arXiv:2311.00423. Retrieved from https:\/\/arxiv.org\/abs\/2311.00423"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210051"},{"key":"e_1_3_2_60_2","unstructured":"Yunjia Xi Weiwen Liu Jianghao Lin Jieming Zhu Bo Chen Ruiming Tang Weinan Zhang Rui Zhang and Yong Yu. 2023. Towards open-world recommendation with knowledge augmentation from large language models. arXiv:2306.10933. Retrieved from https:\/\/arxiv.org\/abs\/2306.10933"},{"key":"e_1_3_2_61_2","first-page":"6518","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Xu Xiao","year":"2020","unstructured":"Xiao Xu, Fang Dong, Yanghua Li, Shaojian He, and Xin Li. 2020. Contextual-bandit based personalized recommendation with time-varying user interests. In Proceedings of the AAAI Conference on Artificial Intelligence, 6518\u20136525."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.2014.2327053"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/1553374.1553524"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475259"},{"key":"e_1_3_2_65_2","first-page":"2504","volume-title":"Proceedings of 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Zhang Xiao","year":"2022","unstructured":"Xiao Zhang, Sunhao Dai, Jun Xu, Zhenhua Dong, Quanyu Dai, and Ji-Rong Wen. 2022. Counteracting user attention bias in music streaming recommendation via reward modification. In Proceedings of 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2504\u20132514."},{"key":"e_1_3_2_66_2","first-page":"41","volume-title":"Proceedings of 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Zhang Xiao","year":"2021","unstructured":"Xiao Zhang, Haonan Jia, Hanjing Su, Wenhan Wang, Jun Xu, and Ji-Rong Wen. 2021. Counterfactual reward modification for streaming recommendation with delayed feedback. In Proceedings of 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 41\u201350."},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1295"},{"key":"e_1_3_2_68_2","article-title":"Reward imputation with sketching for contextual batched bandits","volume":"36","author":"Zhang Xiao","year":"2024","unstructured":"Xiao Zhang, Ninglu Shao, Zihua Si, Jun Xu, Wenhan Wang, Hanjing Su, and Ji-Rong Wen. 2024. Reward imputation with sketching for contextual batched bandits. In Advances in Neural Information Processing Systems, Vol. 36, 2024.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_69_2","first-page":"746","volume-title":"Proceedings of 23rd International Conference on Artificial Intelligence and Statistics","author":"Zhao Peng","year":"2020","unstructured":"Peng Zhao, Lijun Zhang, Yuan Jiang, and Zhi-Hua Zhou. 2020. A simple approach for non-stationary linear bandits. In Proceedings of 23rd International Conference on Artificial Intelligence and Statistics, 746\u2013755."},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532025"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3793543","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T12:14:09Z","timestamp":1774440849000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3793543"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,25]]},"references-count":69,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3793543"],"URL":"https:\/\/doi.org\/10.1145\/3793543","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,25]]},"assertion":[{"value":"2025-03-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-11","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}