{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T23:20:55Z","timestamp":1780356055107,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":40,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,4,20]],"date-time":"2020-04-20T00:00:00Z","timestamp":1587340800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,4,20]]},"DOI":"10.1145\/3366423.3380248","type":"proceedings-article","created":{"date-parts":[[2020,5,4]],"date-time":"2020-05-04T08:11:44Z","timestamp":1588579904000},"page":"1785-1795","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["Adversarial Cooperative Imitation Learning for Dynamic Treatment Regimes\u2731"],"prefix":"10.1145","author":[{"given":"Lu","family":"Wang","sequence":"first","affiliation":[{"name":"East China Normal University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenchao","family":"Yu","sequence":"additional","affiliation":[{"name":"NEC Laboratories America Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaofeng","family":"He","sequence":"additional","affiliation":[{"name":"East China Normal University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei","family":"Cheng","sequence":"additional","affiliation":[{"name":"NEC Laboratories America Inc"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Martin Renqiang","family":"Ren","sequence":"additional","affiliation":[{"name":"NEC Laboratories America Inc"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei","family":"Wang","sequence":"additional","affiliation":[{"name":"University of California Los Angeles"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bo","family":"Zong","sequence":"additional","affiliation":[{"name":"NEC Laboratories America Inc"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haifeng","family":"Chen","sequence":"additional","affiliation":[{"name":"NEC Laboratories America Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongyuan","family":"Zha","sequence":"additional","affiliation":[{"name":"Georgia Tech"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,4,20]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015330.1015430"},{"key":"e_1_3_2_1_2_1","unstructured":"Jacek\u00a0M Bajor and Thomas\u00a0A Lasko. 2016. Predicting medications from diagnostic codes with recurrent neural networks. (2016).  Jacek\u00a0M Bajor and Thomas\u00a0A Lasko. 2016. Predicting medications from diagnostic codes with recurrent neural networks. (2016)."},{"key":"e_1_3_2_1_3_1","volume-title":"The use of reinforcement learning algorithms to meet the challenges of an artificial pancreas. Expert review of medical devices 10, 5","author":"Bothe K","year":"2013","unstructured":"Melanie\u00a0 K Bothe , Luke Dickens , Katrin Reichel , Arn Tellmann , Bjoern Ellger , Martin Westphal , and Ahmed\u00a0 A Faisal . 2013. The use of reinforcement learning algorithms to meet the challenges of an artificial pancreas. Expert review of medical devices 10, 5 ( 2013 ), 661\u2013673. Melanie\u00a0K Bothe, Luke Dickens, Katrin Reichel, Arn Tellmann, Bjoern Ellger, Martin Westphal, and Ahmed\u00a0A Faisal. 2013. The use of reinforcement learning algorithms to meet the challenges of an artificial pancreas. Expert review of medical devices 10, 5 (2013), 661\u2013673."},{"key":"e_1_3_2_1_4_1","unstructured":"Lars Buesing Theophane Weber Sebastien Racaniere SM Eslami Danilo Rezende David\u00a0P Reichert Fabio Viola Frederic Besse Karol Gregor Demis Hassabis 2018. Learning and querying fast generative models for reinforcement learning. arXiv preprint arXiv:1802.03006(2018).  Lars Buesing Theophane Weber Sebastien Racaniere SM Eslami Danilo Rezende David\u00a0P Reichert Fabio Viola Frederic Besse Karol Gregor Demis Hassabis 2018. Learning and querying fast generative models for reinforcement learning. arXiv preprint arXiv:1802.03006(2018)."},{"key":"e_1_3_2_1_5_1","volume-title":"Dynamic treatment regimes. Annual review of statistics and its application 1","author":"Chakraborty Bibhas","year":"2014","unstructured":"Bibhas Chakraborty and Susan\u00a0 A Murphy . 2014. Dynamic treatment regimes. Annual review of statistics and its application 1 ( 2014 ), 447\u2013464. Bibhas Chakraborty and Susan\u00a0A Murphy. 2014. Dynamic treatment regimes. Annual review of statistics and its application 1 (2014), 447\u2013464."},{"key":"e_1_3_2_1_6_1","volume-title":"Generative Adversarial User Model for Reinforcement Learning Based Recommendation System. In International Conference on Machine Learning. 1052\u20131061","author":"Chen Xinshi","year":"2019","unstructured":"Xinshi Chen , Shuang Li , Hui Li , Shaohua Jiang , Yuan Qi , and Le Song . 2019 . Generative Adversarial User Model for Reinforcement Learning Based Recommendation System. In International Conference on Machine Learning. 1052\u20131061 . Xinshi Chen, Shuang Li, Hui Li, Shaohua Jiang, Yuan Qi, and Le Song. 2019. Generative Adversarial User Model for Reinforcement Learning Based Recommendation System. In International Conference on Machine Learning. 1052\u20131061."},{"key":"e_1_3_2_1_7_1","volume-title":"Machine Learning for Healthcare Conference. 301\u2013318","author":"Choi Edward","year":"2016","unstructured":"Edward Choi , Mohammad\u00a0Taha Bahadori , Andy Schuetz , Walter\u00a0 F Stewart , and Jimeng Sun . 2016 . Doctor ai: Predicting clinical events via recurrent neural networks . In Machine Learning for Healthcare Conference. 301\u2013318 . Edward Choi, Mohammad\u00a0Taha Bahadori, Andy Schuetz, Walter\u00a0F Stewart, and Jimeng Sun. 2016. Doctor ai: Predicting clinical events via recurrent neural networks. In Machine Learning for Healthcare Conference. 301\u2013318."},{"key":"e_1_3_2_1_8_1","volume-title":"Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings. In International Conference on Machine Learning. 1008\u20131017","author":"Co-Reyes John","year":"2018","unstructured":"John Co-Reyes , YuXuan Liu , Abhishek Gupta , Benjamin Eysenbach , Pieter Abbeel , and Sergey Levine . 2018 . Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings. In International Conference on Machine Learning. 1008\u20131017 . John Co-Reyes, YuXuan Liu, Abhishek Gupta, Benjamin Eysenbach, Pieter Abbeel, and Sergey Levine. 2018. Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings. In International Conference on Machine Learning. 1008\u20131017."},{"key":"e_1_3_2_1_9_1","unstructured":"Yan Duan Marcin Andrychowicz Bradly Stadie OpenAI\u00a0Jonathan Ho Jonas Schneider Ilya Sutskever Pieter Abbeel and Wojciech Zaremba. 2017. One-shot imitation learning. In Advances in neural information processing systems. 1087\u20131098.  Yan Duan Marcin Andrychowicz Bradly Stadie OpenAI\u00a0Jonathan Ho Jonas Schneider Ilya Sutskever Pieter Abbeel and Wojciech Zaremba. 2017. One-shot imitation learning. In Advances in neural information processing systems. 1087\u20131098."},{"key":"e_1_3_2_1_10_1","unstructured":"Miroslav Dud\u00edk John Langford and Lihong Li. 2011. Doubly Robust Policy Evaluation and Learning. In ICML. 1097\u20131104.  Miroslav Dud\u00edk John Langford and Lihong Li. 2011. Doubly Robust Policy Evaluation and Learning. In ICML. 1097\u20131104."},{"key":"e_1_3_2_1_11_1","volume-title":"International Conference on Machine Learning. 49\u201358","author":"Finn Chelsea","year":"2016","unstructured":"Chelsea Finn , Sergey Levine , and Pieter Abbeel . 2016 . Guided cost learning: Deep inverse optimal control via policy optimization . In International Conference on Machine Learning. 49\u201358 . Chelsea Finn, Sergey Levine, and Pieter Abbeel. 2016. Guided cost learning: Deep inverse optimal control via policy optimization. In International Conference on Machine Learning. 49\u201358."},{"key":"e_1_3_2_1_12_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in neural information processing systems. 2672\u20132680.  Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in neural information processing systems. 2672\u20132680."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2011.5979757"},{"key":"e_1_3_2_1_14_1","volume-title":"Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates","author":"Gu Shixiang","unstructured":"Shixiang Gu , Ethan Holly , Timothy Lillicrap , and Sergey Levine . 2017. Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates . In ICRA. IEEE , 3389\u20133396. Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine. 2017. Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In ICRA. IEEE, 3389\u20133396."},{"key":"e_1_3_2_1_15_1","unstructured":"Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. In Advances in neural information processing systems. 4565\u20134573.  Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. In Advances in neural information processing systems. 4565\u20134573."},{"key":"e_1_3_2_1_16_1","unstructured":"Nan Jiang and Lihong Li. 2015. Doubly robust off-policy value evaluation for reinforcement learning. arXiv preprint arXiv:1511.03722(2015).  Nan Jiang and Lihong Li. 2015. Doubly robust off-policy value evaluation for reinforcement learning. arXiv preprint arXiv:1511.03722(2015)."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3220095"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3272021"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Alistair\u00a0EW Johnson Tom\u00a0J Pollard Lu Shen H\u00a0Lehman Li-wei Mengling Feng Mohammad Ghassemi Benjamin Moody Peter Szolovits Leo\u00a0Anthony Celi and Roger\u00a0G Mark. 2016. MIMIC-III a freely accessible critical care database. Scientific data 3(2016) 160035.  Alistair\u00a0EW Johnson Tom\u00a0J Pollard Lu Shen H\u00a0Lehman Li-wei Mengling Feng Mohammad Ghassemi Benjamin Moody Peter Szolovits Leo\u00a0Anthony Celi and Roger\u00a0G Mark. 2016. MIMIC-III a freely accessible critical care database. Scientific data 3(2016) 160035.","DOI":"10.1038\/sdata.2016.35"},{"key":"e_1_3_2_1_20_1","unstructured":"Diederik\u00a0P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114(2013).  Diederik\u00a0P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114(2013)."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-018-0213-5"},{"key":"e_1_3_2_1_22_1","volume-title":"Infogail: Interpretable imitation learning from visual demonstrations. In Advances in Neural Information Processing Systems. 3812\u20133822.","author":"Li Yunzhu","year":"2017","unstructured":"Yunzhu Li , Jiaming Song , and Stefano Ermon . 2017 . Infogail: Interpretable imitation learning from visual demonstrations. In Advances in Neural Information Processing Systems. 3812\u20133822. Yunzhu Li, Jiaming Song, and Stefano Ermon. 2017. Infogail: Interpretable imitation learning from visual demonstrations. In Advances in Neural Information Processing Systems. 3812\u20133822."},{"key":"e_1_3_2_1_23_1","first-page":"2579","article-title":"Visualizing data using t-SNE","author":"van\u00a0der Maaten Laurens","year":"2008","unstructured":"Laurens van\u00a0der Maaten and Geoffrey Hinton . 2008 . Visualizing data using t-SNE . Journal of machine learning research 9 , Nov (2008), 2579 \u2013 2605 . Laurens van\u00a0der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of machine learning research 9, Nov (2008), 2579\u20132605.","journal-title":"Journal of machine learning research 9"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1111\/1467-9868.00389"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1991.3.1.88"},{"key":"e_1_3_2_1_26_1","unstructured":"Doina Precup Richard\u00a0S. Sutton and Sanjoy Dasgupta. 2001. Off-policy temporal-difference learning with function approximation. In ICML. 417\u2013424.  Doina Precup Richard\u00a0S. Sutton and Sanjoy Dasgupta. 2001. Off-policy temporal-difference learning with function approximation. In ICML. 417\u2013424."},{"key":"e_1_3_2_1_27_1","unstructured":"Aniruddh Raghu Matthieu Komorowski Imran Ahmed Leo Celi Peter Szolovits and Marzyeh Ghassemi. 2017. Deep reinforcement learning for sepsis treatment. arXiv preprint arXiv:1711.09602(2017).  Aniruddh Raghu Matthieu Komorowski Imran Ahmed Leo Celi Peter Szolovits and Marzyeh Ghassemi. 2017. Deep reinforcement learning for sepsis treatment. arXiv preprint arXiv:1711.09602(2017)."},{"key":"e_1_3_2_1_28_1","unstructured":"Aniruddh Raghu Matthieu Komorowski Leo\u00a0Anthony Celi Peter Szolovits and Marzyeh Ghassemi. 2017. Continuous state-space models for optimal sepsis treatment-a deep reinforcement learning approach. arXiv preprint arXiv:1705.08422(2017).  Aniruddh Raghu Matthieu Komorowski Leo\u00a0Anthony Celi Peter Szolovits and Marzyeh Ghassemi. 2017. Continuous state-space models for optimal sepsis treatment-a deep reinforcement learning approach. arXiv preprint arXiv:1705.08422(2017)."},{"key":"e_1_3_2_1_29_1","volume-title":"Proceedings of the thirteenth international conference on artificial intelligence and statistics. 661\u2013668","author":"Ross St\u00e9phane","year":"2010","unstructured":"St\u00e9phane Ross and Drew Bagnell . 2010 . Efficient reductions for imitation learning . In Proceedings of the thirteenth international conference on artificial intelligence and statistics. 661\u2013668 . St\u00e9phane Ross and Drew Bagnell. 2010. Efficient reductions for imitation learning. In Proceedings of the thirteenth international conference on artificial intelligence and statistics. 661\u2013668."},{"key":"e_1_3_2_1_30_1","volume-title":"Proceedings of the fourteenth international conference on artificial intelligence and statistics. 627\u2013635","author":"Ross St\u00e9phane","year":"2011","unstructured":"St\u00e9phane Ross , Geoffrey Gordon , and Drew Bagnell . 2011 . A reduction of imitation learning and structured prediction to no-regret online learning . In Proceedings of the fourteenth international conference on artificial intelligence and statistics. 627\u2013635 . St\u00e9phane Ross, Geoffrey Gordon, and Drew Bagnell. 2011. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics. 627\u2013635."},{"key":"e_1_3_2_1_31_1","volume-title":"Individualized sepsis treatment using reinforcement learning. Nature medicine 24, 11","author":"Saria Suchi","year":"2018","unstructured":"Suchi Saria . 2018. Individualized sepsis treatment using reinforcement learning. Nature medicine 24, 11 ( 2018 ), 1641. Suchi Saria. 2018. Individualized sepsis treatment using reinforcement learning. Nature medicine 24, 11 (2018), 1641."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1001\/jama.2016.0287"},{"key":"e_1_3_2_1_33_1","volume-title":"Introduction to reinforcement learning. Vol.\u00a02","author":"Sutton S","unstructured":"Richard\u00a0 S Sutton , Andrew\u00a0 G Barto , 1998. Introduction to reinforcement learning. Vol.\u00a02 . MIT press Cambridge . Richard\u00a0S Sutton, Andrew\u00a0G Barto, 1998. Introduction to reinforcement learning. Vol.\u00a02. MIT press Cambridge."},{"key":"e_1_3_2_1_34_1","volume-title":"Proceedings of Learning, Inference and Control of Multi-Agent Systems (at NIPS 2016)","author":"Pol Elise Van\u00a0der","year":"2016","unstructured":"Elise Van\u00a0der Pol and Frans\u00a0 A Oliehoek . 2016 . Coordinated deep reinforcement learners for traffic light control . Proceedings of Learning, Inference and Control of Multi-Agent Systems (at NIPS 2016) (2016). Elise Van\u00a0der Pol and Frans\u00a0A Oliehoek. 2016. Coordinated deep reinforcement learners for traffic light control. Proceedings of Learning, Inference and Control of Multi-Agent Systems (at NIPS 2016) (2016)."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219961"},{"key":"e_1_3_2_1_36_1","unstructured":"Markus Wulfmeier Peter Ondruska and Ingmar Posner. 2015. Maximum entropy deep inverse reinforcement learning. arXiv preprint arXiv:1507.04888(2015).  Markus Wulfmeier Peter Ondruska and Ingmar Posner. 2015. Maximum entropy deep inverse reinforcement learning. arXiv preprint arXiv:1507.04888(2015)."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098109"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178876.3185994"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1108\/17563781211255862"},{"key":"e_1_3_2_1_40_1","unstructured":"Brian\u00a0D Ziebart Andrew Maas J\u00a0Andrew Bagnell and Anind\u00a0K Dey. 2008. Maximum entropy inverse reinforcement learning. (2008).  Brian\u00a0D Ziebart Andrew Maas J\u00a0Andrew Bagnell and Anind\u00a0K Dey. 2008. Maximum entropy inverse reinforcement learning. (2008)."}],"event":{"name":"WWW '20: The Web Conference 2020","location":"Taipei Taiwan","acronym":"WWW '20","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web"]},"container-title":["Proceedings of The Web Conference 2020"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3366423.3380248","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3366423.3380248","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:33:01Z","timestamp":1750199581000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3366423.3380248"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,4,20]]},"references-count":40,"alternative-id":["10.1145\/3366423.3380248","10.1145\/3366423"],"URL":"https:\/\/doi.org\/10.1145\/3366423.3380248","relation":{},"subject":[],"published":{"date-parts":[[2020,4,20]]},"assertion":[{"value":"2020-04-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}