{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T17:45:27Z","timestamp":1784137527221,"version":"3.55.0"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2019,2,13]],"date-time":"2019-02-13T00:00:00Z","timestamp":1550016000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["1414119,1430145,1718135"],"award-info":[{"award-number":["1414119,1430145,1718135"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Cyber-Phys. Syst."],"published-print":{"date-parts":[[2019,4,30]]},"abstract":"<jats:p>Autopilot systems are typically composed of an \u201cinner loop\u201d providing stability and control, whereas an \u201couter loop\u201d is responsible for mission-level objectives, such as way-point navigation. Autopilot systems for unmanned aerial vehicles are predominately implemented using Proportional-Integral-Derivative\u00a0(PID) control systems, which have demonstrated exceptional performance in stable environments. However, more sophisticated control is required to operate in unpredictable and harsh environments. Intelligent flight control systems is an active area of research addressing limitations of PID control most recently through the use of reinforcement learning\u00a0(RL), which has had success in other applications, such as robotics. Yet previous work has focused primarily on using RL at the mission-level controller. In this work, we investigate the performance and accuracy of the inner control loop providing attitude control when using intelligent flight control systems trained with state-of-the-art RL algorithms\u2014Deep Deterministic Policy Gradient, Trust Region Policy Optimization, and Proximal Policy Optimization. To investigate these unknowns, we first developed an open source high-fidelity simulation environment to train a flight controller attitude control of a quadrotor through RL. We then used our environment to compare their performance to that of a PID controller to identify if using RL is appropriate in high-precision, time-critical flight control.<\/jats:p>","DOI":"10.1145\/3301273","type":"journal-article","created":{"date-parts":[[2019,2,14]],"date-time":"2019-02-14T19:36:17Z","timestamp":1550172977000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":409,"title":["Reinforcement Learning for UAV Attitude Control"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3982-482X","authenticated-orcid":false,"given":"William","family":"Koch","sequence":"first","affiliation":[{"name":"Boston University, Boston, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Renato","family":"Mancuso","sequence":"additional","affiliation":[{"name":"Boston University, Boston, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Richard","family":"West","sequence":"additional","affiliation":[{"name":"Boston University, Boston, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Azer","family":"Bestavros","sequence":"additional","affiliation":[{"name":"Boston University, Boston, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,2,13]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Retrieved","year":"2018"},{"key":"e_1_2_1_2_1","volume-title":"Retrieved","year":"2018"},{"key":"e_1_2_1_3_1","volume-title":"Retrieved","author":"Robotics Foundation Open Source","year":"2018"},{"key":"e_1_2_1_4_1","volume-title":"Retrieved","author":"Copter APM","year":"2018"},{"key":"e_1_2_1_5_1","volume-title":"Retrieved","year":"2018"},{"key":"e_1_2_1_6_1","volume-title":"Ng","author":"Abbeel Pieter","year":"2007"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 2001 IEEE International Conference on Robotics and Automation (ICRA\u201901)","volume":"2","author":"Andrew Bagnell J."},{"key":"e_1_2_1_8_1","unstructured":"John Blitzer Koby Crammer Alex Kulesza Fernando Pereira and Jennifer Wortman. 2008. Learning bounds for domain adaptation. In Advances in Neural Information Processing Systems. 129--136.   John Blitzer Koby Crammer Alex Kulesza Fernando Pereira and Jennifer Wortman. 2008. Learning bounds for domain adaptation. In Advances in Neural Information Processing Systems. 129--136."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICUMT.2016.7765223"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2004.1302409"},{"key":"e_1_2_1_11_1","unstructured":"Greg Brockman Vicki Cheung Ludwig Pettersson Jonas Schneider John Schulman Jie Tang etal 2016. Openai gym. arXiv:1606.01540.  Greg Brockman Vicki Cheung Ludwig Pettersson Jonas Schneider John Schulman Jie Tang et al. 2016. Openai gym. arXiv:1606.01540."},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 2014 AAAI Spring Symposium Series.","author":"Dewey Daniel","year":"2014"},{"key":"e_1_2_1_13_1","volume-title":"GitHub. Retrieved","author":"Dhariwal Prafulla","year":"2017"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2009.2034145"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/MMAR.2013.6669989"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICAC.2016.29"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCS.2011.941961"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2017.2720851"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/645300.648380"},{"key":"e_1_2_1_20_1","volume-title":"Retrieved","author":"Karpathy Andrej","year":"2018"},{"key":"e_1_2_1_21_1","volume-title":"Ng","author":"Kim H. Jin","year":"2004"},{"key":"e_1_2_1_22_1","volume-title":"GitHub. Retrieved","author":"Koch William","year":"2018"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the 2004 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201904)","volume":"3","author":"Koenig Nathan"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the JANAFF Interagency Propulsion Committee Meeting","author":"KrishnaKumar Kalmanje","year":"2002"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12555-009-0311-8"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1080\/002071700411304"},{"key":"e_1_2_1_27_1","unstructured":"Timothy P. Lillicrap Jonathan J. Hunt Alexander Pritzel Nicolas Heess Tom Erez Yuval Tassa etal 2015. Continuous control with deep reinforcement learning. arXiv:1509.02971.  Timothy P. Lillicrap Jonathan J. Hunt Alexander Pritzel Nicolas Heess Tom Erez Yuval Tassa et al. 2015. Continuous control with deep reinforcement learning. arXiv:1509.02971."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2006.282433"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/DASC.2016.7778103"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/1667965.1667969"},{"key":"e_1_2_1_31_1","unstructured":"Volodymyr Mnih Koray Kavukcuoglu David Silver Alex Graves Ioannis Antonoglou Daan Wierstra etal 2013. Playing Atari with deep reinforcement learning. arXiv:1312.5602.  Volodymyr Mnih Koray Kavukcuoglu David Silver Alex Graves Ioannis Antonoglou Daan Wierstra et al. 2013. Playing Atari with deep reinforcement learning. arXiv:1312.5602."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASE.2017.2651109"},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the International Conference on Machine Learning. 1889--1897","author":"Schulman John","year":"2015"},{"key":"e_1_2_1_34_1","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv:1707.06347.  John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv:1707.06347."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/1830483.1830693"},{"key":"e_1_2_1_36_1","volume-title":"Barto","author":"Sutton Richard S.","year":"1998"},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the 2001 American Control Conference","volume":"6","author":"Wang Le Yi","year":"2001"},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the 2005 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201905)","author":"Waslander Steven Lake"},{"key":"e_1_2_1_39_1","volume-title":"Design of Model-Reference Adaptive Control Systems for Aircraft. MIT Instrumentation Laboratory","author":"Whitaker H. Philip"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.2514\/6.2005-6995"},{"key":"e_1_2_1_41_1","first-page":"11","article-title":"Optimum settings for automatic controllers","volume":"64","author":"Ziegler John G.","year":"1942","journal-title":"Transactions of the ASME"}],"container-title":["ACM Transactions on Cyber-Physical Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3301273","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3301273","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3301273","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:53:59Z","timestamp":1750204439000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3301273"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,2,13]]},"references-count":41,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2019,4,30]]}},"alternative-id":["10.1145\/3301273"],"URL":"https:\/\/doi.org\/10.1145\/3301273","relation":{},"ISSN":["2378-962X","2378-9638"],"issn-type":[{"value":"2378-962X","type":"print"},{"value":"2378-9638","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,2,13]]},"assertion":[{"value":"2018-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-02-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}