{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T16:45:52Z","timestamp":1784306752531,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":14,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,7,25]],"date-time":"2020-07-25T00:00:00Z","timestamp":1595635200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,7,25]]},"DOI":"10.1145\/3397271.3401238","type":"proceedings-article","created":{"date-parts":[[2020,7,25]],"date-time":"2020-07-25T07:50:08Z","timestamp":1595663408000},"page":"1553-1556","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":23,"title":["Video Recommendation with Multi-gate Mixture of Experts Soft Actor Critic"],"prefix":"10.1145","author":[{"given":"Dingcheng","family":"Li","sequence":"first","affiliation":[{"name":"Baidu Research, Seattle, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xu","family":"Li","sequence":"additional","affiliation":[{"name":"Baidu Research, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Wang","sequence":"additional","affiliation":[{"name":"Baidu Inc., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ping","family":"Li","sequence":"additional","affiliation":[{"name":"Baidu Research, Bellevue, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,7,25]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557040"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2959100.2959190"},{"key":"e_1_3_2_1_3_1","volume-title":"Proceedings of the Thirty-Fourth Conference on Uncertainty in Artificial Intelligence (UAI)","author":"Derman Esther","year":"2018","unstructured":"Esther Derman , Daniel J. Mankowitz , Timothy A. Mann , and Shie Mannor . 2018 . Soft-Robust Actor-Critic Policy-Gradient . In Proceedings of the Thirty-Fourth Conference on Uncertainty in Artificial Intelligence (UAI) . Monterey, CA, 208--218. Esther Derman, Daniel J. Mankowitz, Timothy A. Mann, and Shie Mannor. 2018. Soft-Robust Actor-Critic Policy-Gradient. In Proceedings of the Thirty-Fourth Conference on Uncertainty in Artificial Intelligence (UAI). Monterey, CA, 208--218."},{"key":"e_1_3_2_1_4_1","volume-title":"Proceedings of the 33nd International Conference on Machine Learning (ICML)","author":"Duan Yan","year":"2016","unstructured":"Yan Duan , Xi Chen , Rein Houthooft , John Schulman , and Pieter Abbeel . 2016 . Benchmarking Deep Reinforcement Learning for Continuous Control . In Proceedings of the 33nd International Conference on Machine Learning (ICML) . New York City, NY, 1329--1338. Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel. 2016. Benchmarking Deep Reinforcement Learning for Continuous Control. In Proceedings of the 33nd International Conference on Machine Learning (ICML). New York City, NY, 1329--1338."},{"key":"e_1_3_2_1_5_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning (ICML)","author":"Fujimoto Scott","year":"2018","unstructured":"Scott Fujimoto , Herke van Hoof , and David Meger . 2018 . Addressing Function Approximation Error in Actor-Critic Methods . In Proceedings of the 35th International Conference on Machine Learning (ICML) . Stockholmsmassan, Sweden , 1856--1865. Scott Fujimoto, Herke van Hoof, and David Meger. 2018. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the 35th International Conference on Machine Learning (ICML). Stockholmsmassan, Sweden, 1856--1865."},{"key":"e_1_3_2_1_6_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning (ICML)","author":"Haarnoja Tuomas","year":"2018","unstructured":"Tuomas Haarnoja , Aurick Zhou , Pieter Abbeel , and Sergey Levine . 2018 . Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor . In Proceedings of the 35th International Conference on Machine Learning (ICML) . Stockholmsmassan, Sweden , 1856--1865. Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning (ICML). Stockholmsmassan, Sweden, 1856--1865."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1952.10483446"},{"key":"e_1_3_2_1_9_1","volume-title":"Proceedings of the 33nd International Conference on Machine Learning (ICML)","author":"Jiang Nan","year":"2016","unstructured":"Nan Jiang and Lihong Li . 2016 . Doubly Robust Off-policy Value Evaluation for Reinforcement Learning . In Proceedings of the 33nd International Conference on Machine Learning (ICML) . New York City, NY, 652--661. Nan Jiang and Lihong Li. 2016. Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. In Proceedings of the 33nd International Conference on Machine Learning (ICML). New York City, NY, 652--661."},{"key":"e_1_3_2_1_10_1","volume-title":"Proceedings of the 4th International Conference on Learning Representations (ICLR)","author":"Lillicrap Timothy P.","year":"2016","unstructured":"Timothy P. Lillicrap , Jonathan J. Hunt , Alexander Pritzel , Nicolas Heess , Tom Erez , Yuval Tassa , David Silver , and Daan Wierstra . 2016 . Continuous control with deep reinforcement learning . In Proceedings of the 4th International Conference on Learning Representations (ICLR) . San Juan, Puerto Rico. Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016. Continuous control with deep reinforcement learning. In Proceedings of the 4th International Conference on Learning Representations (ICLR). San Juan, Puerto Rico."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611976236.25"},{"key":"e_1_3_2_1_12_1","volume-title":"Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD)","author":"Ma Jiaqi","year":"1930","unstructured":"Jiaqi Ma , Zhe Zhao , Xinyang Yi , Jilin Chen , Lichan Hong , and Ed H. Chi . 2018. Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts . In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD) . London, UK , 1930 --1939. Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi. 2018. Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD). London, UK, 1930--1939."},{"key":"e_1_3_2_1_13_1","volume-title":"Proceedings of the 31th International Conference on Machine Learning (ICML)","author":"Silver David","unstructured":"David Silver , Guy Lever , Nicolas Heess , Thomas Degris , Daan Wierstra , and Martin A. Riedmiller . 2014. Deterministic Policy Gradient Algorithms . In Proceedings of the 31th International Conference on Machine Learning (ICML) . Beijing, China, 387--395. David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin A. Riedmiller. 2014. Deterministic Policy Gradient Algorithms. In Proceedings of the 31th International Conference on Machine Learning (ICML). Beijing, China, 387--395."},{"key":"e_1_3_2_1_14_1","volume-title":"Proceedings of the 13th ACM Conference on Recommender Systems (RecSys)","author":"Zhao Zhe","unstructured":"Zhe Zhao , Lichan Hong , Li Wei , Jilin Chen , Aniruddh Nath , Shawn Andrews , Aditee Kumthekar , Maheswaran Sathiamoorthy , Xinyang Yi , and Ed H. Chi . 2019. Recommending what video to watch next: a multitask ranking system . In Proceedings of the 13th ACM Conference on Recommender Systems (RecSys) . Copenhagen, Denmark, 43--51. Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, and Ed H. Chi. 2019. Recommending what video to watch next: a multitask ranking system. In Proceedings of the 13th ACM Conference on Recommender Systems (RecSys). Copenhagen, Denmark, 43--51."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3336191.3371801"}],"event":{"name":"SIGIR '20: The 43rd International ACM SIGIR conference on research and development in Information Retrieval","location":"Virtual Event China","acronym":"SIGIR '20","sponsor":["SIGIR ACM Special Interest Group on Information Retrieval"]},"container-title":["Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3397271.3401238","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3397271.3401238","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:41:44Z","timestamp":1750200104000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3397271.3401238"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,7,25]]},"references-count":14,"alternative-id":["10.1145\/3397271.3401238","10.1145\/3397271"],"URL":"https:\/\/doi.org\/10.1145\/3397271.3401238","relation":{},"subject":[],"published":{"date-parts":[[2020,7,25]]},"assertion":[{"value":"2020-07-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}