{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,7]],"date-time":"2026-03-07T19:04:15Z","timestamp":1772910255508,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":50,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,4,25]],"date-time":"2022-04-25T00:00:00Z","timestamp":1650844800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,4,25]]},"DOI":"10.1145\/3485447.3512080","type":"proceedings-article","created":{"date-parts":[[2022,4,25]],"date-time":"2022-04-25T05:11:23Z","timestamp":1650863483000},"page":"2067-2077","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["MINDSim: User Simulator for News Recommenders"],"prefix":"10.1145","author":[{"given":"Xufang","family":"Luo","sequence":"first","affiliation":[{"name":"Microsoft Research Asia, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheng","family":"Liu","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shitao","family":"Xiao","sequence":"additional","affiliation":[{"name":"Beijing University of Posts and Telecommunications, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xing","family":"Xie","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dongsheng","family":"Li","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,4,25]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Martin Arjovsky Soumith Chintala and L\u00e9on Bottou. 2017. Wasserstein generative adversarial networks. In ICML."},{"key":"e_1_3_2_1_2_1","volume-title":"A Model-Based Reinforcement Learning with Adversarial Training for Online Recommendation. NeurIPS","author":"Bai Xueying","year":"2019","unstructured":"Xueying Bai, Jian Guan, and Hongning Wang. 2019. A Model-Based Reinforcement Learning with Adversarial Training for Online Recommendation. NeurIPS (2019)."},{"key":"e_1_3_2_1_3_1","unstructured":"Hangbo Bao Li Dong Furu Wei Wenhui Wang Nan Yang Xiaodong Liu Yu Wang Jianfeng Gao Songhao Piao Ming Zhou 2020. Unilmv2: Pseudo-masked language models for unified language model pre-training. In ICML."},{"key":"e_1_3_2_1_4_1","unstructured":"Qingpeng Cai Aris Filos-Ratsikas Pingzhong Tang and Yiwei Zhang. 2018. Reinforcement Mechanism Design for e-commerce. In WWW."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Haokun Chen Xinyi Dai Han Cai Weinan Zhang Xuejian Wang Ruiming Tang Yuzhou Zhang and Yong Yu. 2019. Large-scale interactive recommendation with tree-structured policy gradient. In AAAI.","DOI":"10.1609\/aaai.v33i01.33013312"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Minmin Chen Alex Beutel Paul Covington Sagar Jain Francois Belletti and Ed\u00a0H Chi. 2019. Top-k off-policy correction for a REINFORCE recommender system. In WSDM.","DOI":"10.1145\/3289600.3290999"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Shi-Yong Chen Yang Yu Qing Da Jun Tan Hai-Kuan Huang and Hai-Hong Tang. 2018. Stabilizing reinforcement learning in dynamic environment with application to online recommendation. In KDD.","DOI":"10.1145\/3219819.3220122"},{"key":"e_1_3_2_1_8_1","unstructured":"Xinshi Chen Shuang Li Hui Li Shaohua Jiang Yuan Qi and Le Song. 2019. Generative adversarial user model for reinforcement learning based recommendation system. In ICML."},{"key":"e_1_3_2_1_9_1","volume-title":"Detecting and Examining Gender Bias in News Articles. In Companion Proceedings of the Web Conference","author":"Dacon Jamell","year":"2021","unstructured":"Jamell Dacon and Haochen Liu. 2021. Does Gender Matter in the News? Detecting and Examining Gender Bias in News Articles. In Companion Proceedings of the Web Conference 2021. 385\u2013392."},{"key":"e_1_3_2_1_10_1","unstructured":"Zihang Dai Zhilin Yang Yiming Yang Jaime\u00a0G Carbonell Quoc Le and Ruslan Salakhutdinov. 2019. Transformer-XL: Attentive Language Models beyond a Fixed-Length Context. In ACL."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Linhao Dong Shuang Xu and Bo Xu. 2018. Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition. In ICASSP.","DOI":"10.1109\/ICASSP.2018.8462506"},{"key":"e_1_3_2_1_12_1","volume-title":"The 22nd International Conference on Artificial Intelligence and Statistics. 2681\u20132690","author":"Feydy Jean","year":"2019","unstructured":"Jean Feydy, Thibault S\u00e9journ\u00e9, Fran\u00e7ois-Xavier Vialard, Shun-ichi Amari, Alain Trouve, and Gabriel Peyr\u00e9. 2019. Interpolating between Optimal Transport and MMD using Sinkhorn Divergences. In The 22nd International Conference on Artificial Intelligence and Statistics. 2681\u20132690."},{"key":"e_1_3_2_1_13_1","unstructured":"Ishaan Gulrajani Faruk Ahmed Martin Arjovsky Vincent Dumoulin and Aaron\u00a0C Courville. 2017. Improved Training of Wasserstein GANs. In NeurIPS."},{"key":"e_1_3_2_1_14_1","unstructured":"Huifeng Guo Ruiming Tang Yunming Ye Zhenguo Li and Xiuqiang He. 2017. DeepFM: a factorization-machine based neural network for CTR prediction. In IJCAI."},{"key":"e_1_3_2_1_15_1","unstructured":"Xiangnan He Lizi Liao Hanwang Zhang Liqiang Nie Xia Hu and Tat-Seng Chua. 2017. Neural collaborative filtering. In WWW."},{"key":"e_1_3_2_1_16_1","unstructured":"Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415(2016)."},{"key":"e_1_3_2_1_17_1","unstructured":"Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. In NeurIPS."},{"key":"e_1_3_2_1_18_1","unstructured":"Yujing Hu Qing Da Anxiang Zeng Yang Yu and Yinghui Xu. 2018. Reinforcement learning to rank in e-commerce search engine: Formalization analysis and application. In KDD."},{"key":"e_1_3_2_1_19_1","volume-title":"SlateQ: A tractable decomposition for reinforcement learning with recommendation sets. arXiv preprint","author":"Ie Eugene","year":"2019","unstructured":"Eugene Ie, Vihan Jain, Jing Wang, Sanmit Narvekar, Ritesh Agarwal, Rui Wu, Heng-Tze Cheng, Tushar Chandra, and Craig Boutilier. 2019. SlateQ: A tractable decomposition for reinforcement learning with recommendation sets. arXiv preprint (2019)."},{"key":"e_1_3_2_1_20_1","volume-title":"Martin Mladenov, Vihan Jain, Sanmit Narvekar, Jing Wang, Rui Wu, and Craig Boutilier.","author":"Ie Eugene","year":"2019","unstructured":"Eugene Ie, Chih wei Hsu, Martin Mladenov, Vihan Jain, Sanmit Narvekar, Jing Wang, Rui Wu, and Craig Boutilier. 2019. RecSim: A Configurable Simulation Platform for Recommender Systems. arXiv preprint arXiv:1909.04847(2019)."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recommendation. In ICDM.","DOI":"10.1109\/ICDM.2018.00035"},{"key":"e_1_3_2_1_22_1","volume-title":"A field guide to dynamical recurrent networks","author":"Kolen F","unstructured":"John\u00a0F Kolen and Stefan\u00a0C Kremer. 2001. A field guide to dynamical recurrent networks. John Wiley & Sons."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Lihong Li Wei Chu John Langford and Xuanhui Wang. 2011. Unbiased Offline Evaluation of Contextual-Bandit-Based News Article Recommendation Algorithms. In WSDM.","DOI":"10.1145\/1935826.1935878"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"crossref","unstructured":"Jianxun Lian Fuzheng Zhang Xing Xie and Guangzhong Sun. 2018. Towards Better Representation Learning for Personalized News Recommendation: a Multi-Channel Deep Fusion Approach.. In IJCAI. 3805\u20133811.","DOI":"10.24963\/ijcai.2018\/529"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1719970.1719976"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"crossref","unstructured":"Zheng Liu Yu Xing Fangzhao Wu Mingxiao An and Xing Xie. 2019. Hi-Fi Ark: Deep User Representation via High-Fidelity Archive Network.. In IJCAI.","DOI":"10.24963\/ijcai.2019\/424"},{"key":"e_1_3_2_1_27_1","unstructured":"Andrew\u00a0L Maas Awni\u00a0Y Hannun and Andrew\u00a0Y Ng. 2013. Rectifier nonlinearities improve neural network acoustic models. In ICML."},{"key":"e_1_3_2_1_28_1","volume-title":"Human-level control through deep reinforcement learning. Nature 518, 7540","author":"Mnih Volodymyr","year":"2015","unstructured":"Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei\u00a0A Rusu, Joel Veness, Marc\u00a0G Bellemare, Alex Graves, Martin Riedmiller, Andreas\u00a0K Fidjeland, Georg Ostrovski, 2015. Human-level control through deep reinforcement learning. Nature 518, 7540 (2015), 529."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Shumpei Okura Yukihiro Tagami Shingo Ono and Akira Tajima. 2017. Embedding-based news recommendation for millions of users. In KDD.","DOI":"10.1145\/3097983.3098108"},{"key":"e_1_3_2_1_30_1","volume-title":"Patrick Forre, Marcello Carioni, Samarth Bhargav, Max Welling, Tim Genewein, and Frank Nielsen.","author":"Patrini Giorgio","year":"2020","unstructured":"Giorgio Patrini, Rianne van\u00a0den Berg, Patrick Forre, Marcello Carioni, Samarth Bhargav, Max Welling, Tim Genewein, and Frank Nielsen. 2020. Sinkhorn autoencoders. In UAI."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Gabriel Peyr\u00e9 Marco Cuturi 2019. Computational optimal transport: With applications to data science. Foundations and Trends\u00ae in Machine Learning 11 5-6(2019) 355\u2013607.","DOI":"10.1561\/2200000073"},{"key":"e_1_3_2_1_32_1","unstructured":"Marc\u2019Aurelio Ranzato Sumit Chopra Michael Auli and Wojciech Zaremba. 2015. Sequence level training with recurrent neural networks. arXiv preprint arXiv:1511.06732(2015)."},{"key":"e_1_3_2_1_33_1","volume-title":"Factorization machines","author":"Rendle Steffen","unstructured":"Steffen Rendle. 2010. Factorization machines. In ICDM. IEEE."},{"key":"e_1_3_2_1_34_1","volume-title":"BPR: Bayesian personalized ranking from implicit feedback. In UAI.","author":"Rendle Steffen","year":"2009","unstructured":"Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian personalized ranking from implicit feedback. In UAI."},{"key":"e_1_3_2_1_35_1","volume-title":"Recogym: A reinforcement learning environment for the problem of product recommendation in online advertising. arXiv preprint arXiv:1808.00720(2018).","author":"Rohde David","year":"2018","unstructured":"David Rohde, Stephen Bonner, Travis Dunlop, Flavian Vasile, and Alexandros Karatzoglou. 2018. Recogym: A reinforcement learning environment for the problem of product recommendation in online advertising. arXiv preprint arXiv:1808.00720(2018)."},{"key":"e_1_3_2_1_36_1","unstructured":"John Schulman Sergey Levine Pieter Abbeel Michael Jordan and Philipp Moritz. 2015. Trust region policy optimization. In ICML. PMLR."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"crossref","unstructured":"Wenjie Shang Yang Yu Qingyang Li Zhiwei Qin Yiping Meng and Jieping Ye. 2019. Environment reconstruction with hidden confounders for reinforcement learning based recommendation. In KDD.","DOI":"10.1145\/3292500.3330933"},{"key":"e_1_3_2_1_38_1","volume-title":"Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning. In AAAI.","author":"Shi Jing-Cheng","year":"2019","unstructured":"Jing-Cheng Shi, Yang Yu, Qing Da, Shi-Yong Chen, and An-Xiang Zeng. 2019. Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning. In AAAI."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-1-55860-141-3.50030-4"},{"key":"e_1_3_2_1_40_1","volume-title":"Introduction to reinforcement learning. Vol.\u00a0135","author":"Sutton S","unstructured":"Richard\u00a0S Sutton, Andrew\u00a0G Barto, 1998. Introduction to reinforcement learning. Vol.\u00a0135. MIT press Cambridge."},{"key":"e_1_3_2_1_41_1","volume-title":"Attention is All you Need. NeurIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan\u00a0N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. NeurIPS (2017)."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3406522.3446019"},{"key":"e_1_3_2_1_43_1","unstructured":"et\u00a0al. Wayne Xin\u00a0Zhao. 2020. RecBole: Towards a Unified Comprehensive and Efficient Framework for Recommendation Algorithms. arXiv preprint arXiv:2011.01731(2020)."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/s42486-020-00044-0"},{"key":"e_1_3_2_1_45_1","volume-title":"MIND: A Large-scale Dataset for News Recommendation. In ACL. https:\/\/msnews.github.io\/.","author":"Wu Fangzhao","year":"2020","unstructured":"Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, and Ming Zhou. 2020. MIND: A Large-scale Dataset for News Recommendation. In ACL. https:\/\/msnews.github.io\/."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"crossref","unstructured":"Shitao Xiao Zheng Liu Yingxia Shao Tao Di and Xing Xie. 2021. Training Large-Scale News Recommenders with Pretrained Language Models in the Loop. arXiv preprint arXiv:2102.09268(2021).","DOI":"10.1145\/3534678.3539120"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"crossref","unstructured":"Shuo Zhang and Krisztian Balog. 2020. Evaluating Conversational Recommender Systems via User Simulation. In KDD.","DOI":"10.1145\/3394486.3403202"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"crossref","unstructured":"Xiangyu Zhao Liang Zhang Zhuoye Ding Long Xia Jiliang Tang and Dawei Yin. 2018. Recommendations with negative feedback via pairwise deep reinforcement learning. In KDD.","DOI":"10.1145\/3219819.3219886"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178876.3185994"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"crossref","unstructured":"Lixin Zou Long Xia Pan Du Zhuo Zhang Ting Bai Weidong Liu Jian-Yun Nie and Dawei Yin. 2020. Pseudo Dyna-Q: A reinforcement learning framework for interactive recommendation. In WSDM.","DOI":"10.1145\/3336191.3371801"}],"event":{"name":"WWW '22: The ACM Web Conference 2022","location":"Virtual Event, Lyon France","acronym":"WWW '22","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web"]},"container-title":["Proceedings of the ACM Web Conference 2022"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3485447.3512080","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3485447.3512080","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:07Z","timestamp":1750188607000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3485447.3512080"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,25]]},"references-count":50,"alternative-id":["10.1145\/3485447.3512080","10.1145\/3485447"],"URL":"https:\/\/doi.org\/10.1145\/3485447.3512080","relation":{},"subject":[],"published":{"date-parts":[[2022,4,25]]},"assertion":[{"value":"2022-04-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}