{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:46:54Z","timestamp":1750308414983,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":19,"publisher":"ACM","license":[{"start":{"date-parts":[[2018,5,28]],"date-time":"2018-05-28T00:00:00Z","timestamp":1527465600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2018,5,28]]},"DOI":"10.1145\/3194085.3194088","type":"proceedings-article","created":{"date-parts":[[2018,7,17]],"date-time":"2018-07-17T16:16:43Z","timestamp":1531844203000},"page":"16-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Distributed deep reinforcement learning on the cloud for autonomous driving"],"prefix":"10.1145","author":[{"given":"Mitchell","family":"Spryn","sequence":"first","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aditya","family":"Sharma","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dhawal","family":"Parkar","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Madhur","family":"Shrimal","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,5,28]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. (2015). https:\/\/www.tensorflow.org\/ Software available from tensorflow.org.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. (2015). https:\/\/www.tensorflow.org\/ Software available from tensorflow.org."},{"key":"e_1_3_2_1_2_1","unstructured":"Mariusz Bojarski Davide Del Testa Daniel Dworakowski Bernhard Firner Beat Flepp Prasoon Goyal Lawrence D. Jackel Mathew Monfort Urs Muller Jiakai Zhang Xin Zhang Jake Zhao and Karol Zieba. 2016. End to End Learning for Self-Driving Cars. CoRR abs\/1604.07316 (2016). arXiv:1604.07316 http:\/\/arxiv.org\/abs\/1604.07316  Mariusz Bojarski Davide Del Testa Daniel Dworakowski Bernhard Firner Beat Flepp Prasoon Goyal Lawrence D. Jackel Mathew Monfort Urs Muller Jiakai Zhang Xin Zhang Jake Zhao and Karol Zieba. 2016. End to End Learning for Self-Driving Cars. CoRR abs\/1604.07316 (2016). arXiv:1604.07316 http:\/\/arxiv.org\/abs\/1604.07316"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.312"},{"key":"e_1_3_2_1_4_1","unstructured":"Fran\u00e7ois Chollet et al. 2015. Keras. https:\/\/github.com\/keras-team\/keras. (2015).  Fran\u00e7ois Chollet et al. 2015. Keras. https:\/\/github.com\/keras-team\/keras. (2015)."},{"key":"e_1_3_2_1_5_1","unstructured":"Jeffrey Dean Greg Corrado Rajat Monga Kai Chen Matthieu Devin Mark Mao Marc'aurelio Ranzato Andrew Senior Paul Tucker Ke Yang Quoc V. Le and Andrew Y. Ng. 2012. Large Scale Distributed Deep Networks. In Advances in Neural Information Processing Systems 25 F. Pereira C. J. C. Burges L. Bottou and K. Q. Weinberger (Eds.). Curran Associates Inc. 1223--1231. http:\/\/papers.nips.cc\/paper\/4687-large-scale-distributed-deep-networks.pdf   Jeffrey Dean Greg Corrado Rajat Monga Kai Chen Matthieu Devin Mark Mao Marc'aurelio Ranzato Andrew Senior Paul Tucker Ke Yang Quoc V. Le and Andrew Y. Ng. 2012. Large Scale Distributed Deep Networks. In Advances in Neural Information Processing Systems 25 F. Pereira C. J. C. Burges L. Bottou and K. Q. Weinberger (Eds.). Curran Associates Inc. 1223--1231. http:\/\/papers.nips.cc\/paper\/4687-large-scale-distributed-deep-networks.pdf"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Nidhi Kalra and Susan M. Paddock. 2016. Driving to Safety: How Many Miles of Driving Would It Take to Demonstrate Autonomous Vehicle Reliability? (2016). https:\/\/www.rand.org\/pubs\/research_reports\/RR1478.html  Nidhi Kalra and Susan M. Paddock. 2016. Driving to Safety: How Many Miles of Driving Would It Take to Demonstrate Autonomous Vehicle Reliability? (2016). https:\/\/www.rand.org\/pubs\/research_reports\/RR1478.html","DOI":"10.7249\/RR1478"},{"key":"e_1_3_2_1_7_1","unstructured":"Francisco S Melo. {n. d.}. Convergence of Q-learning: A simple proof. ({n. d.}).  Francisco S Melo. {n. d.}. Convergence of Q-learning: A simple proof. ({n. d.})."},{"volume-title":"Playing Atari With Deep Reinforcement Learning. In NIPS Deep Learning Workshop.","year":"2013","author":"Mnih Volodymyr","key":"e_1_3_2_1_8_1"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_1_10_1","unstructured":"Arun Nair Praveen Srinivasan Sam Blackwell Cagdas Alcicek Rory Fearon Alessandro De Maria Vedavyas Panneershelvam Mustafa Suleyman Charles Beattie Stig Petersen Shane Legg Volodymyr Mnih Koray Kavukcuoglu and David Silver. 2015. Massively Parallel Methods for Deep Reinforcement Learning. CoRR abs\/1507.04296 (2015). arXiv:1507.04296 http:\/\/arxiv.org\/abs\/1507.04296  Arun Nair Praveen Srinivasan Sam Blackwell Cagdas Alcicek Rory Fearon Alessandro De Maria Vedavyas Panneershelvam Mustafa Suleyman Charles Beattie Stig Petersen Shane Legg Volodymyr Mnih Koray Kavukcuoglu and David Silver. 2015. Massively Parallel Methods for Deep Reinforcement Learning. CoRR abs\/1507.04296 (2015). arXiv:1507.04296 http:\/\/arxiv.org\/abs\/1507.04296"},{"key":"e_1_3_2_1_11_1","unstructured":"A. E. Sallab M. Abdou E. Perot and S. Yogamani. 2016. End-to-end deep reinforcement learning for lane keeping assist. (2016). https:\/\/arxiv.org\/pdf\/1612.04340.pdf  A. E. Sallab M. Abdou E. Perot and S. Yogamani. 2016. End-to-end deep reinforcement learning for lane keeping assist. (2016). https:\/\/arxiv.org\/pdf\/1612.04340.pdf"},{"key":"e_1_3_2_1_12_1","unstructured":"Shital Shah Debadeepta Dey Chris Lovett and Ashish Kapoor. 2017. AirSim: High-Fidelity Visual and Physical Simulation for Autonomous Vehicles. In Field and Service Robotics. arXiv:arXiv:1705.05065 https:\/\/arxiv.org\/abs\/1705.05065  Shital Shah Debadeepta Dey Chris Lovett and Ashish Kapoor. 2017. AirSim: High-Fidelity Visual and Physical Simulation for Autonomous Vehicles. In Field and Service Robotics. arXiv:arXiv:1705.05065 https:\/\/arxiv.org\/abs\/1705.05065"},{"key":"e_1_3_2_1_13_1","unstructured":"Shai Shalev-Shwartz Shaked Shammah and Amnon Shashua. 2016. Safe Multi-Agent Reinforcement Learning for Autonomous Driving. CoRR abs\/1610.03295 (2016). arXiv:1610.03295 http:\/\/arxiv.org\/abs\/1610.03295  Shai Shalev-Shwartz Shaked Shammah and Amnon Shashua. 2016. Safe Multi-Agent Reinforcement Learning for Autonomous Driving. CoRR abs\/1610.03295 (2016). arXiv:1610.03295 http:\/\/arxiv.org\/abs\/1610.03295"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"crossref","unstructured":"David Silver Julian Schrittwieser Karen Simonyan Ioannis Antonoglou Aja Huang Arthur Guez Thomas Hubert Lucas Baker Matthew Lai Adrian Bolton Yutian Chen Timothy Lillicrap Fan Hui Laurent Sifre George van den Driessche Thore Graepel and Demis Hassabis. 2017. Mastering the game of Go without human knowledge. 550 (10 2017) 354--359.  David Silver Julian Schrittwieser Karen Simonyan Ioannis Antonoglou Aja Huang Arthur Guez Thomas Hubert Lucas Baker Matthew Lai Adrian Bolton Yutian Chen Timothy Lillicrap Fan Hui Laurent Sifre George van den Driessche Thore Graepel and Demis Hassabis. 2017. Mastering the game of Go without human knowledge. 550 (10 2017) 354--359.","DOI":"10.1038\/nature24270"},{"key":"e_1_3_2_1_15_1","unstructured":"Matthew E. Taylor and Peter Stone. 2009. Transfer Learning for Reinforcement Learning Domains: A Survey. J. Mach. Learn. Res. 10 (Dec. 2009) 1633--1685. http:\/\/dl.acm.org\/citation.cfm?id=1577069.1755839   Matthew E. Taylor and Peter Stone. 2009. Transfer Learning for Reinforcement Learning Domains: A Survey. J. Mach. Learn. Res. 10 (Dec. 2009) 1633--1685. http:\/\/dl.acm.org\/citation.cfm?id=1577069.1755839"},{"key":"e_1_3_2_1_16_1","unstructured":"Christopher John Cornish Hellaby Watkins. 1989. Learning from Delayed Rewards. Ph.D. Dissertation. King's College Cambridge UK. http:\/\/www.cs.rhul.ac.uk\/~chrisw\/new_thesis.pdf  Christopher John Cornish Hellaby Watkins. 1989. Learning from Delayed Rewards. Ph.D. Dissertation. King's College Cambridge UK. http:\/\/www.cs.rhul.ac.uk\/~chrisw\/new_thesis.pdf"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-016-0043-6"},{"key":"e_1_3_2_1_19_1","unstructured":"Wei Zhang Suyog Gupta Xiangru Lian and Ji Liu. 2015. Staleness-aware Async-SGD for Distributed Deep Learning. CoRR abs\/1511.05950 (2015). http:\/\/arxiv.org\/abs\/1511.05950   Wei Zhang Suyog Gupta Xiangru Lian and Ji Liu. 2015. Staleness-aware Async-SGD for Distributed Deep Learning. CoRR abs\/1511.05950 (2015). http:\/\/arxiv.org\/abs\/1511.05950"}],"event":{"name":"ICSE '18: 40th International Conference on Software Engineering","sponsor":["SIGSOFT ACM Special Interest Group on Software Engineering","IEEE-CS Computer Society"],"location":"Gothenburg Sweden","acronym":"ICSE '18"},"container-title":["Proceedings of the 1st International Workshop on Software Engineering for AI in Autonomous Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3194085.3194088","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3194085.3194088","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:49:15Z","timestamp":1750268955000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3194085.3194088"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,5,28]]},"references-count":19,"alternative-id":["10.1145\/3194085.3194088","10.1145\/3194085"],"URL":"https:\/\/doi.org\/10.1145\/3194085.3194088","relation":{},"subject":[],"published":{"date-parts":[[2018,5,28]]},"assertion":[{"value":"2018-05-28","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}