{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,12]],"date-time":"2025-12-12T13:07:39Z","timestamp":1765544859698,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":47,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,17]],"date-time":"2022-10-17T00:00:00Z","timestamp":1665964800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,17]]},"DOI":"10.1145\/3511808.3557094","type":"proceedings-article","created":{"date-parts":[[2022,10,16]],"date-time":"2022-10-16T01:29:57Z","timestamp":1665883797000},"page":"3555-3564","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Learning List-wise Representation in Reinforcement Learning for Ads Allocation with Multiple Auxiliary Tasks"],"prefix":"10.1145","author":[{"given":"Ze","family":"Wang","sequence":"first","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guogang","family":"Liao","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaowen","family":"Shi","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoxu","family":"Wu","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chuheng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongkang","family":"Wang","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xingxing","family":"Wang","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dong","family":"Wang","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,17]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Unsupervised state representation learning in atari. arXiv preprint arXiv:1906.08226","author":"Anand Ankesh","year":"2019","unstructured":"Ankesh Anand , Evan Racah , Sherjil Ozair , Yoshua Bengio , Marc-Alexandre C\u00f4t\u00e9 , and R Devon Hjelm . 2019. Unsupervised state representation learning in atari. arXiv preprint arXiv:1906.08226 ( 2019 ). Ankesh Anand, Evan Racah, Sherjil Ozair, Yoshua Bengio, Marc-Alexandre C\u00f4t\u00e9, and R Devon Hjelm. 2019. Unsupervised state representation learning in atari. arXiv preprint arXiv:1906.08226 (2019)."},{"key":"e_1_3_2_2_2_1","volume-title":"Ziyu Wang, and Nando de Freitas.","author":"Aytar Yusuf","year":"2018","unstructured":"Yusuf Aytar , Tobias Pfaff , David Budden , Tom Le Paine , Ziyu Wang, and Nando de Freitas. 2018 . Playing hard exploration games by watching youtube. arXiv preprint arXiv:1805.11592 (2018). Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, and Nando de Freitas. 2018. Playing hard exploration games by watching youtube. arXiv preprint arXiv:1805.11592 (2018)."},{"key":"e_1_3_2_2_3_1","volume-title":"Yan","author":"Carrion Carlos","year":"2021","unstructured":"Carlos Carrion , Zenan Wang , Harikesh Nair , Xianghong Luo , Yulin Lei , Xiliang Lin , Wenlong Chen , Qiyu Hu , Changping Peng , Yongjun Bao , and Weipeng P . Yan . 2021 . Blending Advertising with Organic Content in E-Commerce: A Virtual Bids Optimization Approach. ArXiv , Vol. abs\/ 2105 .13556 (2021). Carlos Carrion, Zenan Wang, Harikesh Nair, Xianghong Luo, Yulin Lei, Xiliang Lin, Wenlong Chen, Qiyu Hu, Changping Peng, Yongjun Bao, and Weipeng P. Yan. 2021. Blending Advertising with Organic Content in E-Commerce: A Virtual Bids Optimization Approach. ArXiv, Vol. abs\/2105.13556 (2021)."},{"volume-title":"2018 IEEE. In RSJ International Conference on Intelligent Robots and Systems (IROS). 1577--1584","author":"Dwibedi Debidatta","key":"e_1_3_2_2_4_1","unstructured":"Debidatta Dwibedi , Jonathan Tompson , Corey Lynch , and Pierre Sermanet . [n.,d.]. Learning actionable representations from visual observations . In 2018 IEEE. In RSJ International Conference on Intelligent Robots and Systems (IROS). 1577--1584 . Debidatta Dwibedi, Jonathan Tompson, Corey Lynch, and Pierre Sermanet. [n.,d.]. Learning actionable representations from visual observations. In 2018 IEEE. In RSJ International Conference on Intelligent Robots and Systems (IROS). 1577--1584."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178876.3186165"},{"key":"e_1_3_2_2_6_1","volume-title":"Revisit Recommender System in the Permutation Prospective. ArXiv","author":"Feng Yufei","year":"2057","unstructured":"Yufei Feng , Yu Gong , Fei Sun , Qingwen Liu , and Wenwu Ou. 2021. Revisit Recommender System in the Permutation Prospective. ArXiv , Vol. abs\/ 2102 .1 2057 (2021). Yufei Feng, Yu Gong, Fei Sun, Qingwen Liu, and Wenwu Ou. 2021. Revisit Recommender System in the Permutation Prospective. ArXiv, Vol. abs\/2102.12057 (2021)."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2016.7487173"},{"key":"e_1_3_2_2_8_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Lavet Vincent Francc","year":"2019","unstructured":"Vincent Francc ois- Lavet , Yoshua Bengio , Doina Precup , and Joelle Pineau . 2019 . Combined reinforcement learning via abstract representations . In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33 . 3582--3589. Vincent Francc ois-Lavet, Yoshua Bengio, Doina Precup, and Joelle Pineau. 2019. Combined reinforcement learning via abstract representations. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 3582--3589."},{"key":"e_1_3_2_2_9_1","first-page":"1605","article-title":"An Empirical Analysis of Search Engine Advertising","volume":"55","author":"Ghose A.","year":"2009","unstructured":"A. Ghose and Sha Yang . 2009 . An Empirical Analysis of Search Engine Advertising : Sponsored Search in Electronic Markets. Manag. Sci. , Vol. 55 (2009), 1605 -- 1622 . A. Ghose and Sha Yang. 2009. An Empirical Analysis of Search Engine Advertising: Sponsored Search in Electronic Markets. Manag. Sci., Vol. 55 (2009), 1605--1622.","journal-title":"Sponsored Search in Electronic Markets. Manag. Sci."},{"key":"e_1_3_2_2_10_1","volume-title":"International Conference on Machine Learning. PMLR, 3875--3886","author":"Guo Zhaohan Daniel","year":"2020","unstructured":"Zhaohan Daniel Guo , Bernardo Avila Pires , Bilal Piot , Jean-Bastien Grill , Florent Altch\u00e9 , R\u00e9mi Munos , and Mohammad Gheshlaghi Azar . 2020 . Bootstrap latent-predictive representations for multitask reinforcement learning . In International Conference on Machine Learning. PMLR, 3875--3886 . Zhaohan Daniel Guo, Bernardo Avila Pires, Bilal Piot, Jean-Bastien Grill, Florent Altch\u00e9, R\u00e9mi Munos, and Mohammad Gheshlaghi Azar. 2020. Bootstrap latent-predictive representations for multitask reinforcement learning. In International Conference on Machine Learning. PMLR, 3875--3886."},{"key":"e_1_3_2_2_11_1","volume-title":"Recurrent world models facilitate policy evolution. Advances in neural information processing systems","author":"Ha David","year":"2018","unstructured":"David Ha and J\u00fcrgen Schmidhuber . 2018. Recurrent world models facilitate policy evolution. Advances in neural information processing systems , Vol. 31 ( 2018 ). David Ha and J\u00fcrgen Schmidhuber. 2018. Recurrent world models facilitate policy evolution. Advances in neural information processing systems, Vol. 31 (2018)."},{"key":"e_1_3_2_2_12_1","volume-title":"Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603","author":"Hafner Danijar","year":"2019","unstructured":"Danijar Hafner , Timothy Lillicrap , Jimmy Ba , and Mohammad Norouzi . 2019a. Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603 ( 2019 ). Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. 2019a. Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603 (2019)."},{"key":"e_1_3_2_2_13_1","volume-title":"International conference on machine learning. PMLR, 2555--2565","author":"Hafner Danijar","year":"2019","unstructured":"Danijar Hafner , Timothy Lillicrap , Ian Fischer , Ruben Villegas , David Ha , Honglak Lee , and James Davidson . 2019 b. Learning latent dynamics for planning from pixels . In International conference on machine learning. PMLR, 2555--2565 . Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. 2019b. Learning latent dynamics for planning from pixels. In International conference on machine learning. PMLR, 2555--2565."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33013796"},{"key":"e_1_3_2_2_15_1","volume-title":"Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu.","author":"Jaderberg Max","year":"2016","unstructured":"Max Jaderberg , Volodymyr Mnih , Wojciech Marian Czarnecki , Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu. 2016 . Reinforcement learning with unsupervised auxiliary tasks. arXiv preprint arXiv:1611.05397 (2016). Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu. 2016. Reinforcement learning with unsupervised auxiliary tasks. arXiv preprint arXiv:1611.05397 (2016)."},{"key":"e_1_3_2_2_16_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba . 2014 . Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014). Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2944789.2944791"},{"key":"e_1_3_2_2_18_1","volume-title":"International Conference on Machine Learning. PMLR, 5639--5650","author":"Laskin Michael","year":"2020","unstructured":"Michael Laskin , Aravind Srinivas , and Pieter Abbeel . 2020 . Curl: Contrastive unsupervised representations for reinforcement learning . In International Conference on Machine Learning. PMLR, 5639--5650 . Michael Laskin, Aravind Srinivas, and Pieter Abbeel. 2020. Curl: Contrastive unsupervised representations for reinforcement learning. In International Conference on Machine Learning. PMLR, 5639--5650."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3411952"},{"key":"e_1_3_2_2_20_1","volume-title":"Deep Page-Level Interest Network in Reinforcement Learning for Ads Allocation. arXiv preprint arXiv:2204.00377","author":"Liao Guogang","year":"2022","unstructured":"Guogang Liao , Xiaowen Shi , Ze Wang , Xiaoxu Wu , Chuheng Zhang , Yongkang Wang , Xingxing Wang , and Dong Wang . 2022. Deep Page-Level Interest Network in Reinforcement Learning for Ads Allocation. arXiv preprint arXiv:2204.00377 ( 2022 ). Guogang Liao, Xiaowen Shi, Ze Wang, Xiaoxu Wu, Chuheng Zhang, Yongkang Wang, Xingxing Wang, and Dong Wang. 2022. Deep Page-Level Interest Network in Reinforcement Learning for Ads Allocation. arXiv preprint arXiv:2204.00377 (2022)."},{"key":"e_1_3_2_2_21_1","volume-title":"Cross DQN: Cross Deep Q Network for Ads Allocation in Feed. arXiv preprint arXiv:2109.04353","author":"Liao Guogang","year":"2021","unstructured":"Guogang Liao , Ze Wang , Xiaoxu Wu , Xiaowen Shi , Chuheng Zhang , Yongkang Wang , Xingxing Wang , and Dong Wang . 2021. Cross DQN: Cross Deep Q Network for Ads Allocation in Feed. arXiv preprint arXiv:2109.04353 ( 2021 ). Guogang Liao, Ze Wang, Xiaoxu Wu, Xiaowen Shi, Chuheng Zhang, Yongkang Wang, Xingxing Wang, and Dong Wang. 2021. Cross DQN: Cross Deep Q Network for Ads Allocation in Feed. arXiv preprint arXiv:2109.04353 (2021)."},{"key":"e_1_3_2_2_22_1","volume-title":"Return-based contrastive representation learning for reinforcement learning. arXiv preprint arXiv:2102.10960","author":"Liu Guoqing","year":"2021","unstructured":"Guoqing Liu , Chuheng Zhang , Li Zhao , Tao Qin , Jinhua Zhu , Jian Li , Nenghai Yu , and Tie-Yan Liu . 2021. Return-based contrastive representation learning for reinforcement learning. arXiv preprint arXiv:2102.10960 ( 2021 ). Guoqing Liu, Chuheng Zhang, Li Zhao, Tao Qin, Jinhua Zhu, Jian Li, Nenghai Yu, and Tie-Yan Liu. 2021. Return-based contrastive representation learning for reinforcement learning. arXiv preprint arXiv:2102.10960 (2021)."},{"key":"e_1_3_2_2_23_1","volume-title":"Thang Doan, Philip Bachman, and R Devon Hjelm.","author":"Mazoure Bogdan","year":"2020","unstructured":"Bogdan Mazoure , Remi Tachet des Combes , Thang Doan, Philip Bachman, and R Devon Hjelm. 2020 . Deep reinforcement and infomax learning. arXiv preprint arXiv:2006.07217 (2020). Bogdan Mazoure, Remi Tachet des Combes, Thang Doan, Philip Bachman, and R Devon Hjelm. 2020. Deep reinforcement and infomax learning. arXiv preprint arXiv:2006.07217 (2020)."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1561\/0400000057"},{"key":"e_1_3_2_2_25_1","unstructured":"Piotr Mirowski Razvan Pascanu Fabio Viola Hubert Soyer Andrew J Ballard Andrea Banino Misha Denil Ross Goroshin Laurent Sifre Koray Kavukcuoglu etal 2016. Learning to navigate in complex environments. arXiv preprint arXiv:1611.03673 (2016).  Piotr Mirowski Razvan Pascanu Fabio Viola Hubert Soyer Andrew J Ballard Andrea Banino Misha Denil Ross Goroshin Laurent Sifre Koray Kavukcuoglu et al. 2016. Learning to navigate in complex environments. arXiv preprint arXiv:1611.03673 (2016)."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"crossref","unstructured":"Volodymyr Mnih Koray Kavukcuoglu David Silver Andrei A Rusu Joel Veness Marc G Bellemare Alex Graves Martin Riedmiller Andreas K Fidjeland Georg Ostrovski etal 2015. Human-level control through deep reinforcement learning. nature Vol. 518 7540 (2015) 529--533.  Volodymyr Mnih Koray Kavukcuoglu David Silver Andrei A Rusu Joel Veness Marc G Bellemare Alex Graves Martin Riedmiller Andreas K Fidjeland Georg Ostrovski et al. 2015. Human-level control through deep reinforcement learning. nature Vol. 518 7540 (2015) 529--533.","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2020.07.038"},{"key":"e_1_3_2_2_28_1","volume-title":"Representation learning-assisted click-through rate prediction. arXiv preprint arXiv:1906.04365","author":"Ouyang Wentao","year":"2019","unstructured":"Wentao Ouyang , Xiuwu Zhang , Shukui Ren , Chao Qi , Zhaojie Liu , and Yanlong Du. 2019. Representation learning-assisted click-through rate prediction. arXiv preprint arXiv:1906.04365 ( 2019 ). Wentao Ouyang, Xiuwu Zhang, Shukui Ren, Chao Qi, Zhaojie Liu, and Yanlong Du. 2019. Representation learning-assisted click-through rate prediction. arXiv preprint arXiv:1906.04365 (2019)."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3412728"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8462891"},{"key":"e_1_3_2_2_31_1","volume-title":"Loss is its own reward: Self-supervision for reinforcement learning. arXiv preprint arXiv:1612.07307","author":"Shelhamer Evan","year":"2016","unstructured":"Evan Shelhamer , Parsa Mahmoudieh , Max Argus , and Trevor Darrell . 2016. Loss is its own reward: Self-supervision for reinforcement learning. arXiv preprint arXiv:1612.07307 ( 2016 ). Evan Shelhamer, Parsa Mahmoudieh, Max Argus, and Trevor Darrell. 2016. Loss is its own reward: Self-supervision for reinforcement learning. arXiv preprint arXiv:1612.07307 (2016)."},{"key":"e_1_3_2_2_32_1","volume-title":"Improved deep metric learning with multi-class n-pair loss objective. Advances in neural information processing systems","author":"Sohn Kihyuk","year":"2016","unstructured":"Kihyuk Sohn . 2016. Improved deep metric learning with multi-class n-pair loss objective. Advances in neural information processing systems , Vol. 29 ( 2016 ). Kihyuk Sohn. 2016. Improved deep metric learning with multi-class n-pair loss objective. Advances in neural information processing systems, Vol. 29 (2016)."},{"key":"e_1_3_2_2_33_1","unstructured":"Richard S Sutton Andrew G Barto etal 1998. Introduction to reinforcement learning. Vol. 135. MIT press Cambridge.  Richard S Sutton Andrew G Barto et al. 1998. Introduction to reinforcement learning. Vol. 135. MIT press Cambridge."},{"key":"e_1_3_2_2_34_1","volume-title":"Representation learning with contrastive predictive coding. arXiv e-prints","author":"den Oord Aaron Van","year":"2018","unstructured":"Aaron Van den Oord , Yazhe Li , and Oriol Vinyals . 2018. Representation learning with contrastive predictive coding. arXiv e-prints ( 2018 ), arXiv--1807. Aaron Van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv e-prints (2018), arXiv--1807."},{"key":"e_1_3_2_2_35_1","volume-title":"Advances in Neural Information Processing Systems","volume":"33","author":"van der Pol Elise","year":"2020","unstructured":"Elise van der Pol , Daniel Worrall , Herke van Hoof , Frans Oliehoek , and Max Welling . 2020 . MDP homomorphic networks: Group symmetries in reinforcement learning . Advances in Neural Information Processing Systems , Vol. 33 (2020). Elise van der Pol, Daniel Worrall, Herke van Hoof, Frans Oliehoek, and Max Welling. 2020. MDP homomorphic networks: Group symmetries in reinforcement learning. Advances in Neural Information Processing Systems, Vol. 33 (2020)."},{"key":"e_1_3_2_2_36_1","volume-title":"Attention is all you need. arXiv preprint arXiv:1706.03762","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , Lukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. arXiv preprint arXiv:1706.03762 ( 2017 ). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762 (2017)."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"crossref","unstructured":"B. Wang Zhaonan Li Jie Tang Kuo Zhang Songcan Chen and Liyun Ru. 2011. Learning to Advertise: How Many Ads Are Enough?. In PAKDD.  B. Wang Zhaonan Li Jie Tang Kuo Zhang Songcan Chen and Liyun Ru. 2011. Learning to Advertise: How Many Ads Are Enough?. In PAKDD.","DOI":"10.1007\/978-3-642-20847-8_42"},{"key":"e_1_3_2_2_38_1","volume-title":"International conference on machine learning. PMLR","author":"Wang Ziyu","year":"2016","unstructured":"Ziyu Wang , Tom Schaul , Matteo Hessel , Hado Hasselt , Marc Lanctot , and Nando Freitas . 2016 . Dueling network architectures for deep reinforcement learning . In International conference on machine learning. PMLR , 1995--2003. Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas. 2016. Dueling network architectures for deep reinforcement learning. In International conference on machine learning. PMLR, 1995--2003."},{"key":"e_1_3_2_2_39_1","volume-title":"Generator and Critic: A Deep Reinforcement Learning Approach for Slate Re-ranking in E-commerce. ArXiv","author":"Wei Jianxiong","year":"2020","unstructured":"Jianxiong Wei , Anxiang Zeng , Yueqiu Wu , Pengxin Guo , Q. Hua , and Qingpeng Cai . 2020. Generator and Critic: A Deep Reinforcement Learning Approach for Slate Re-ranking in E-commerce. ArXiv , Vol. abs\/ 2005 .12206 ( 2020 ). Jianxiong Wei, Anxiang Zeng, Yueqiu Wu, Pengxin Guo, Q. Hua, and Qingpeng Cai. 2020. Generator and Critic: A Deep Reinforcement Learning Approach for Slate Re-ranking in E-commerce. ArXiv, Vol. abs\/2005.12206 (2020)."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i5.16580"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403391"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3184558.3191584"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"crossref","unstructured":"Mengchen Zhao Z. Li Bo An Haifeng Lu Yifan Yang and Chen Chu. 2018. Impression Allocation for Combating Fraud in E-commerce Via Deep Reinforcement Learning with Action Norm Penalty. In IJCAI.  Mengchen Zhao Z. Li Bo An Haifeng Lu Yifan Yang and Chen Chu. 2018. Impression Allocation for Combating Fraud in E-commerce Via Deep Reinforcement Learning with Action Norm Penalty. In IJCAI.","DOI":"10.24963\/ijcai.2018\/548"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i1.16156"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403384"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219823"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2021.115825"}],"event":{"name":"CIKM '22: The 31st ACM International Conference on Information and Knowledge Management","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","SIGIR ACM Special Interest Group on Information Retrieval"],"location":"Atlanta GA USA","acronym":"CIKM '22"},"container-title":["Proceedings of the 31st ACM International Conference on Information &amp; Knowledge Management"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3511808.3557094","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3511808.3557094","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:56Z","timestamp":1750188656000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3511808.3557094"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,17]]},"references-count":47,"alternative-id":["10.1145\/3511808.3557094","10.1145\/3511808"],"URL":"https:\/\/doi.org\/10.1145\/3511808.3557094","relation":{},"subject":[],"published":{"date-parts":[[2022,10,17]]},"assertion":[{"value":"2022-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}