{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T00:10:51Z","timestamp":1773015051487,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":27,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,4,19]],"date-time":"2021-04-19T00:00:00Z","timestamp":1618790400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,4,19]]},"DOI":"10.1145\/3442381.3449862","type":"proceedings-article","created":{"date-parts":[[2021,6,3]],"date-time":"2021-06-03T19:00:27Z","timestamp":1622746827000},"page":"1040-1052","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Match Plan Generation in Web Search with Parameterized Action Reinforcement Learning"],"prefix":"10.1145","author":[{"given":"Ziyan","family":"Luo","sequence":"first","affiliation":[{"name":"Microsoft and University of California, San Diego, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Linfeng","family":"Zhao","sequence":"additional","affiliation":[{"name":"Microsoft and Northeastern University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Cheng","sequence":"additional","affiliation":[{"name":"Microsoft and North China Electric Power University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sihao","family":"Chen","sequence":"additional","affiliation":[{"name":"Microsoft and University of California, Berkeley, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qi","family":"Chen","sequence":"additional","affiliation":[{"name":"Microsoft, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hui","family":"Xue","sequence":"additional","affiliation":[{"name":"Microsoft, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haidong","family":"Wang","sequence":"additional","affiliation":[{"name":"Microsoft, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chuanjie","family":"Liu","sequence":"additional","affiliation":[{"name":"Microsoft, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mao","family":"Yang","sequence":"additional","affiliation":[{"name":"Microsoft, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lintao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Microsoft, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,6,3]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Marcin Andrychowicz Filip Wolski Alex Ray Jonas Schneider Rachel Fong Peter Welinder Bob McGrew Josh Tobin OpenAI\u00a0Pieter Abbeel and Wojciech Zaremba. 2017. Hindsight experience replay. In Advances in neural information processing systems. 5048\u20135058.  Marcin Andrychowicz Filip Wolski Alex Ray Jonas Schneider Rachel Fong Peter Welinder Bob McGrew Josh Tobin OpenAI\u00a0Pieter Abbeel and Wojciech Zaremba. 2017. Hindsight experience replay. In Advances in neural information processing systems. 5048\u20135058."},{"key":"e_1_3_2_1_2_1","unstructured":"Craig\u00a0J Bester Steven\u00a0D James and George\u00a0D Konidaris. 2019. Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces. arXiv preprint arXiv:1905.04388(2019).  Craig\u00a0J Bester Steven\u00a0D James and George\u00a0D Konidaris. 2019. Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces. arXiv preprint arXiv:1905.04388(2019)."},{"key":"e_1_3_2_1_3_1","unstructured":"Petros Christodoulou. 2019. Soft Actor-Critic for Discrete Action Settings. arxiv:cs.LG\/1910.07207  Petros Christodoulou. 2019. Soft Actor-Critic for Discrete Action Settings. arxiv:cs.LG\/1910.07207"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3015022.3015026"},{"key":"e_1_3_2_1_5_1","unstructured":"Scott Fujimoto Herke van Hoof and David Meger. 2018. Addressing function approximation error in actor-critic methods. arXiv preprint arXiv:1802.09477(2018).  Scott Fujimoto Herke van Hoof and David Meger. 2018. Addressing function approximation error in actor-critic methods. arXiv preprint arXiv:1802.09477(2018)."},{"key":"e_1_3_2_1_6_1","volume-title":"International Conference on Machine Learning. 1861\u20131870","author":"Haarnoja Tuomas","year":"2018","unstructured":"Tuomas Haarnoja , Aurick Zhou , Pieter Abbeel , and Sergey Levine . 2018 . Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor . In International Conference on Machine Learning. 1861\u20131870 . Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In International Conference on Machine Learning. 1861\u20131870."},{"key":"e_1_3_2_1_7_1","unstructured":"Tuomas Haarnoja Aurick Zhou Kristian Hartikainen George Tucker Sehoon Ha Jie Tan Vikash Kumar Henry Zhu Abhishek Gupta Pieter Abbeel 2018. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905(2018).  Tuomas Haarnoja Aurick Zhou Kristian Hartikainen George Tucker Sehoon Ha Jie Tan Vikash Kumar Henry Zhu Abhishek Gupta Pieter Abbeel 2018. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905(2018)."},{"key":"e_1_3_2_1_8_1","volume-title":"Deep Reinforcement Learning in Parameterized Action Space. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings.","author":"J.","unstructured":"Matthew\u00a0 J. Hausknecht and Peter Stone. 2016 . Deep Reinforcement Learning in Parameterized Action Space. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings. Matthew\u00a0J. Hausknecht and Peter Stone. 2016. Deep Reinforcement Learning in Parameterized Action Space. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings."},{"key":"e_1_3_2_1_9_1","unstructured":"Nicolas Heess Jonathan\u00a0J Hunt Timothy\u00a0P Lillicrap and David Silver. 2015. Memory-based control with recurrent neural networks. arXiv preprint arXiv:1512.04455(2015).  Nicolas Heess Jonathan\u00a0J Hunt Timothy\u00a0P Lillicrap and David Silver. 2015. Memory-based control with recurrent neural networks. arXiv preprint arXiv:1512.04455(2015)."},{"key":"e_1_3_2_1_10_1","volume-title":"Long short-term memory. Neural computation 9, 8","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber . 1997. Long short-term memory. Neural computation 9, 8 ( 1997 ), 1735\u20131780. Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural computation 9, 8 (1997), 1735\u20131780."},{"key":"e_1_3_2_1_11_1","volume-title":"Planning and acting in partially observable stochastic domains. Artificial intelligence 101, 1-2","author":"Kaelbling Leslie\u00a0Pack","year":"1998","unstructured":"Leslie\u00a0Pack Kaelbling , Michael\u00a0 L Littman , and Anthony\u00a0 R Cassandra . 1998. Planning and acting in partially observable stochastic domains. Artificial intelligence 101, 1-2 ( 1998 ), 99\u2013134. Leslie\u00a0Pack Kaelbling, Michael\u00a0L Littman, and Anthony\u00a0R Cassandra. 1998. Planning and acting in partially observable stochastic domains. Artificial intelligence 101, 1-2 (1998), 99\u2013134."},{"key":"e_1_3_2_1_12_1","unstructured":"Steven Kapturowski Georg Ostrovski John Quan Remi Munos and Will Dabney. 2018. Recurrent experience replay in distributed reinforcement learning. (2018).  Steven Kapturowski Georg Ostrovski John Quan Remi Munos and Will Dabney. 2018. Recurrent experience replay in distributed reinforcement learning. (2018)."},{"key":"e_1_3_2_1_13_1","unstructured":"Timothy\u00a0P Lillicrap Jonathan\u00a0J Hunt Alexander Pritzel Nicolas Heess Tom Erez Yuval Tassa David Silver and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971(2015).  Timothy\u00a0P Lillicrap Jonathan\u00a0J Hunt Alexander Pritzel Nicolas Heess Tom Erez Yuval Tassa David Silver and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971(2015)."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/3016100.3016169"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1148170.1148246"},{"key":"e_1_3_2_1_16_1","unstructured":"Volodymyr Mnih Koray Kavukcuoglu David Silver Alex Graves Ioannis Antonoglou Daan Wierstra and Martin Riedmiller. 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602(2013).  Volodymyr Mnih Koray Kavukcuoglu David Silver Alex Graves Ioannis Antonoglou Daan Wierstra and Martin Riedmiller. 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602(2013)."},{"key":"e_1_3_2_1_17_1","volume-title":"Human-level control through deep reinforcement learning. Nature 518, 7540","author":"Mnih Volodymyr","year":"2015","unstructured":"Volodymyr Mnih , Koray Kavukcuoglu , David Silver , Andrei\u00a0 A Rusu , Joel Veness , Marc\u00a0 G Bellemare , Alex Graves , Martin Riedmiller , Andreas\u00a0 K Fidjeland , Georg Ostrovski , 2015. Human-level control through deep reinforcement learning. Nature 518, 7540 ( 2015 ), 529. Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei\u00a0A Rusu, Joel Veness, Marc\u00a0G Bellemare, Alex Graves, Martin Riedmiller, Andreas\u00a0K Fidjeland, Georg Ostrovski, 2015. Human-level control through deep reinforcement learning. Nature 518, 7540 (2015), 529."},{"key":"e_1_3_2_1_18_1","volume-title":"Overcoming Exploration in Reinforcement Learning with Demonstrations. In 2018 IEEE International Conference on Robotics and Automation (ICRA). 6292\u20136299","author":"Nair Ashvin","year":"2018","unstructured":"Ashvin Nair , Bob McGrew , Marcin Andrychowicz , Wojciech Zaremba , and Pieter Abbeel . 2018 . Overcoming Exploration in Reinforcement Learning with Demonstrations. In 2018 IEEE International Conference on Robotics and Automation (ICRA). 6292\u20136299 . https:\/\/doi.org\/10.1109\/ICRA.2018.8463162 ISSN: 2577-087X. Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel. 2018. Overcoming Exploration in Reinforcement Learning with Demonstrations. In 2018 IEEE International Conference on Robotics and Automation (ICRA). 6292\u20136299. https:\/\/doi.org\/10.1109\/ICRA.2018.8463162 ISSN: 2577-087X."},{"key":"e_1_3_2_1_19_1","unstructured":"Matthias Plappert Rein Houthooft Prafulla Dhariwal Szymon Sidor Richard\u00a0Y Chen Xi Chen Tamim Asfour Pieter Abbeel and Marcin Andrychowicz. 2017. Parameter space noise for exploration. arXiv preprint arXiv:1706.01905(2017).  Matthias Plappert Rein Houthooft Prafulla Dhariwal Szymon Sidor Richard\u00a0Y Chen Xi Chen Tamim Asfour Pieter Abbeel and Marcin Andrychowicz. 2017. Parameter space noise for exploration. arXiv preprint arXiv:1706.01905(2017)."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210127"},{"key":"e_1_3_2_1_21_1","unstructured":"Tom Schaul John Quan Ioannis Antonoglou and David Silver. 2015. Prioritized experience replay. arXiv preprint arXiv:1511.05952(2015).  Tom Schaul John Quan Ioannis Antonoglou and David Silver. 2015. Prioritized experience replay. arXiv preprint arXiv:1511.05952(2015)."},{"key":"e_1_3_2_1_22_1","unstructured":"David Silver Guy Lever Nicolas Heess Thomas Degris Daan Wierstra and Martin Riedmiller. 2014. Deterministic policy gradient algorithms.  David Silver Guy Lever Nicolas Heess Thomas Degris Daan Wierstra and Martin Riedmiller. 2014. Deterministic policy gradient algorithms."},{"key":"e_1_3_2_1_23_1","volume-title":"Reinforcement learning: An introduction","author":"Sutton S","unstructured":"Richard\u00a0 S Sutton and Andrew\u00a0 G Barto . 2018. Reinforcement learning: An introduction . MIT press . Richard\u00a0S Sutton and Andrew\u00a0G Barto. 2018. Reinforcement learning: An introduction. MIT press."},{"key":"e_1_3_2_1_24_1","volume-title":"Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards. arXiv:1707.08817 [cs] (Oct","author":"Vecerik Mel","year":"2018","unstructured":"Mel Vecerik , Todd Hester , Jonathan Scholz , Fumin Wang , Olivier Pietquin , Bilal Piot , Nicolas Heess , Thomas Roth\u00f6rl , Thomas Lampe , and Martin Riedmiller . 2018. Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards. arXiv:1707.08817 [cs] (Oct . 2018 ). arXiv:1707.08817. Mel Vecerik, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Roth\u00f6rl, Thomas Lampe, and Martin Riedmiller. 2018. Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards. arXiv:1707.08817 [cs] (Oct. 2018). arXiv:1707.08817."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.58337"},{"key":"e_1_3_2_1_26_1","unstructured":"Jiechao Xiong Qing Wang Zhuoran Yang Peng Sun Lei Han Yang Zheng Haobo Fu Tong Zhang Ji Liu and Han Liu. 2018. Parametrized deep q-networks learning: Reinforcement learning with discrete-continuous hybrid action space. arXiv preprint arXiv:1810.06394(2018).  Jiechao Xiong Qing Wang Zhuoran Yang Peng Sun Lei Han Yang Zheng Haobo Fu Tong Zhang Ji Liu and Han Liu. 2018. Parametrized deep q-networks learning: Reinforcement learning with discrete-continuous hybrid action space. arXiv preprint arXiv:1810.06394(2018)."},{"key":"e_1_3_2_1_27_1","volume-title":"Inverted files for text search engines. ACM computing surveys (CSUR) 38, 2","author":"Zobel Justin","year":"2006","unstructured":"Justin Zobel and Alistair Moffat . 2006. Inverted files for text search engines. ACM computing surveys (CSUR) 38, 2 ( 2006 ), 6. Justin Zobel and Alistair Moffat. 2006. Inverted files for text search engines. ACM computing surveys (CSUR) 38, 2 (2006), 6."}],"event":{"name":"WWW '21: The Web Conference 2021","location":"Ljubljana Slovenia","acronym":"WWW '21","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web"]},"container-title":["Proceedings of the Web Conference 2021"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3442381.3449862","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3442381.3449862","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:24:31Z","timestamp":1750195471000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3442381.3449862"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,19]]},"references-count":27,"alternative-id":["10.1145\/3442381.3449862","10.1145\/3442381"],"URL":"https:\/\/doi.org\/10.1145\/3442381.3449862","relation":{},"subject":[],"published":{"date-parts":[[2021,4,19]]},"assertion":[{"value":"2021-06-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}