{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,26]],"date-time":"2026-08-26T15:21:53Z","timestamp":1787757713772,"version":"build-2784847793"},"publisher-location":"New York, NY, USA","reference-count":51,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,11,13]],"date-time":"2019-11-13T00:00:00Z","timestamp":1573603200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,11,13]]},"DOI":"10.1145\/3360322.3360849","type":"proceedings-article","created":{"date-parts":[[2019,11,5]],"date-time":"2019-11-05T13:56:15Z","timestamp":1572962175000},"page":"316-325","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":121,"title":["Gnu-RL"],"prefix":"10.1145","author":[{"given":"Bingqing","family":"Chen","sequence":"first","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zicheng","family":"Cai","sequence":"additional","affiliation":[{"name":"Dell Technologies, Austin, TX, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mario","family":"Berg\u00e9s","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,11,13]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CDC.2012.6425995"},{"key":"e_1_3_2_1_2_1","unstructured":"Brandon Amos Ivan Jimenez Jacob Sacks Byron Boots and J. Zico Kolter. 2018. Differentiable MPC for End-to-end Planning and Control. (2018).  Brandon Amos Ivan Jimenez Jacob Sacks Byron Boots and J. Zico Kolter. 2018. Differentiable MPC for End-to-end Planning and Control. (2018)."},{"key":"e_1_3_2_1_3_1","volume-title":"Proceedings of the 34th International Conference on Machine Learning-Volume 70","author":"Amos Brandon","year":"2017","unstructured":"Brandon Amos and J Zico Kolter . 2017 . Optnet: Differentiable optimization as a layer in neural networks . In Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 136--145. Brandon Amos and J Zico Kolter. 2017. Optnet: Differentiable optimization as a layer in neural networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70. JMLR. org, 136--145."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2011.2161242"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCS.2016.2535913"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/SIU.2018.8404287"},{"key":"e_1_3_2_1_7_1","first-page":"1","article-title":"Automatic differentiation in machine learning: a survey","volume":"18","author":"Baydin Atilim Gunes","year":"2018","unstructured":"Atilim Gunes Baydin , Barak A Pearlmutter , Alexey Andreyevich Radul , and Jeffrey Mark Siskind . 2018 . Automatic differentiation in machine learning: a survey . Journal of Marchine Learning Research 18 (2018), 1 -- 43 . Atilim Gunes Baydin, Barak A Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind. 2018. Automatic differentiation in machine learning: a survey. Journal of Marchine Learning Research 18 (2018), 1--43.","journal-title":"Journal of Marchine Learning Research"},{"key":"e_1_3_2_1_8_1","volume-title":"Second international conference on building energy and environment. 979--986","author":"Bengea S","year":"2012","unstructured":"S Bengea , A Kelman , Francesco Borrelli , Russell Taylor , and Satish Narayanan . 2012 . Model predictive control for mid-size commercial building hvac: Implementation, results and energy savings . In Second international conference on building energy and environment. 979--986 . S Bengea, A Kelman, Francesco Borrelli, Russell Taylor, and Satish Narayanan. 2012. Model predictive control for mid-size commercial building hvac: Implementation, results and energy savings. In Second international conference on building energy and environment. 979--986."},{"key":"e_1_3_2_1_9_1","volume-title":"Openai gym. arXiv preprint arXiv:1606.01540","author":"Brockman Greg","year":"2016","unstructured":"Greg Brockman , Vicki Cheung , Ludwig Pettersson , Jonas Schneider , John Schulman , Jie Tang , and Wojciech Zaremba . 2016. Openai gym. arXiv preprint arXiv:1606.01540 ( 2016 ). Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016. Openai gym. arXiv preprint arXiv:1606.01540 (2016)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.segan.2016.02.002"},{"key":"e_1_3_2_1_11_1","volume-title":"Reinforcement learning for energy conservation and comfort in buildings. Building and environment 42, 7","author":"Dalamagkidis Konstantinos","year":"2007","unstructured":"Konstantinos Dalamagkidis , Denia Kolokotsa , Konstantinos Kalaitzakis , and George S Stavrakakis . 2007. Reinforcement learning for energy conservation and comfort in buildings. Building and environment 42, 7 ( 2007 ), 2686--2698. Konstantinos Dalamagkidis, Denia Kolokotsa, Konstantinos Kalaitzakis, and George S Stavrakakis. 2007. Reinforcement learning for energy conservation and comfort in buildings. Building and environment 42, 7 (2007), 2686--2698."},{"key":"e_1_3_2_1_12_1","unstructured":"Richard Evans and Jim Gao. 2016. DeepMind AI reduces Google data centre cooling bill by 40%. https:\/\/deepmind.com\/blog\/deepmind-ai-reduces-google-data-centre-cooling-bill-40\/. Accessed: 2019-06-19.  Richard Evans and Jim Gao. 2016. DeepMind AI reduces Google data centre cooling bill by 40%. https:\/\/deepmind.com\/blog\/deepmind-ai-reduces-google-data-centre-cooling-bill-40\/. Accessed: 2019-06-19."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.enbuild.2012.06.016"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCA.2015.7320826"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11694"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.egypro.2019.01.494"},{"key":"e_1_3_2_1_19_1","unstructured":"Sham M Kakade. 2002. A natural policy gradient. In Advances in neural information processing systems. 1531--1538.  Sham M Kakade. 2002. A natural policy gradient. In Advances in neural information processing systems. 1531--1538."},{"key":"e_1_3_2_1_20_1","unstructured":"Sham Machandranath Kakade etal 2003. On the sample complexity of reinforcement learning. Ph.D. Dissertation.  Sham Machandranath Kakade et al. 2003. On the sample complexity of reinforcement learning. Ph.D. Dissertation."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.energy.2017.12.019"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.buildenv.2016.05.034"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.rser.2014.05.056"},{"key":"e_1_3_2_1_24_1","volume-title":"Transforming cooling optimization for green data center via deep reinforcement learning. arXiv preprint arXiv:1709.05077","author":"Li Yuanlong","year":"2017","unstructured":"Yuanlong Li , Yonggang Wen , Kyle Guan , and Dacheng Tao . 2017. Transforming cooling optimization for green data center via deep reinforcement learning. arXiv preprint arXiv:1709.05077 ( 2017 ). Yuanlong Li, Yonggang Wen, Kyle Guan, and Dacheng Tao. 2017. Transforming cooling optimization for green data center via deep reinforcement learning. arXiv preprint arXiv:1709.05077 (2017)."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.enbuild.2005.06.001"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1115\/1.2710491"},{"key":"e_1_3_2_1_27_1","volume-title":"Wiley Encyclopedia of Electrical and Electronics Engineering (1999), 1--19.","author":"Ljung Lennart","unstructured":"Lennart Ljung . 1999. System identification . Wiley Encyclopedia of Electrical and Electronics Engineering (1999), 1--19. Lennart Ljung. 1999. System identification. Wiley Encyclopedia of Electrical and Electronics Engineering (1999), 1--19."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.apenergy.2014.12.019"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.enbuild.2014.03.057"},{"key":"e_1_3_2_1_30_1","volume-title":"International conference on machine learning. 1928--1937","author":"Mnih Volodymyr","year":"2016","unstructured":"Volodymyr Mnih , Adria Puigdomenech Badia , Mehdi Mirza , Alex Graves , Timothy Lillicrap , Tim Harley , David Silver , and Koray Kavukcuoglu . 2016 . Asynchronous methods for deep reinforcement learning . In International conference on machine learning. 1928--1937 . Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016. Asynchronous methods for deep reinforcement learning. In International conference on machine learning. 1928--1937."},{"key":"e_1_3_2_1_31_1","volume-title":"Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602","author":"Mnih Volodymyr","year":"2013","unstructured":"Volodymyr Mnih , Koray Kavukcuoglu , David Silver , Alex Graves , Ioannis Antonoglou , Daan Wierstra , and Martin Riedmiller . 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 ( 2013 ). Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 (2013)."},{"key":"e_1_3_2_1_32_1","volume-title":"How Google's AlphaGo beat a Go world champion. The Atlantic 28","author":"Moyer Christopher","year":"2016","unstructured":"Christopher Moyer . 2016. How Google's AlphaGo beat a Go world champion. The Atlantic 28 ( 2016 ). Christopher Moyer. 2016. How Google's AlphaGo beat a Go world champion. The Atlantic 28 (2016)."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2017.8"},{"key":"e_1_3_2_1_34_1","volume-title":"Deep Reinforcement Learning for Optimal Control of Space Heating. arXiv preprint arXiv:1805.03777","author":"Nagy Adam","year":"2018","unstructured":"Adam Nagy , Hussain Kazmi , Farah Cheaib , and Johan Driesen . 2018. Deep Reinforcement Learning for Optimal Control of Space Heating. arXiv preprint arXiv:1805.03777 ( 2018 ). Adam Nagy, Hussain Kazmi, Farah Cheaib, and Johan Driesen. 2018. Deep Reinforcement Learning for Optimal Control of Space Heating. arXiv preprint arXiv:1805.03777 (2018)."},{"key":"e_1_3_2_1_35_1","unstructured":"OSIsoft. [n. d.]. Tools: PI DataLink. https:\/\/www.osisoft.com\/pi-system\/picapabilities\/pi-system-tools\/pi-datalink\/. Accessed: 2019-06-19.  OSIsoft. [n. d.]. Tools: PI DataLink. https:\/\/www.osisoft.com\/pi-system\/picapabilities\/pi-system-tools\/pi-datalink\/. Accessed: 2019-06-19."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.buildenv.2018.10.028"},{"key":"e_1_3_2_1_37_1","unstructured":"Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in pytorch. (2017).  Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in pytorch. (2017)."},{"key":"e_1_3_2_1_38_1","volume-title":"IEEE International Conference on Automatic Computing: Feedback Computing","volume":"16","author":"Peng Kuo Shiuan","year":"2016","unstructured":"Kuo Shiuan Peng and Clayton T Morrison . 2016 . Model Predictive Prior Reinforcement Learning for a Heat Pump Thermostat . In IEEE International Conference on Automatic Computing: Feedback Computing , Vol. 16 . Kuo Shiuan Peng and Clayton T Morrison. 2016. Model Predictive Prior Reinforcement Learning for a Heat Pump Thermostat. In IEEE International Conference on Automatic Computing: Feedback Computing, Vol. 16."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.enbuild.2012.10.024"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCA.2011.6044402"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1080\/09613218.2016.1139885"},{"key":"e_1_3_2_1_42_1","volume-title":"International Conference on Machine Learning. 1889--1897","author":"Schulman John","year":"2015","unstructured":"John Schulman , Sergey Levine , Pieter Abbeel , Michael Jordan , and Philipp Moritz . 2015 . Trust region policy optimization . In International Conference on Machine Learning. 1889--1897 . John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015. Trust region policy optimization. In International Conference on Machine Learning. 1889--1897."},{"key":"e_1_3_2_1_43_1","volume-title":"Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347","author":"Schulman John","year":"2017","unstructured":"John Schulman , Filip Wolski , Prafulla Dhariwal , Alec Radford , and Oleg Klimov . 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 ( 2017 ). John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017)."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364910369715"},{"key":"e_1_3_2_1_45_1","unstructured":"Dark Sky. [n. d.]. Dark Sky API. https:\/\/darksky.net\/dev. Accessed: 2019-06-19.  Dark Sky. [n. d.]. Dark Sky API. https:\/\/darksky.net\/dev. Accessed: 2019-06-19."},{"key":"e_1_3_2_1_46_1","unstructured":"Energy Star. 2018. Portfolio Manager Technical Reference: Climate and Weather. https:\/\/portfoliomanager.energystar.gov\/pdf\/reference\/Climate%20and%20Weather.pdf. Accessed: 2019-06-19.  Energy Star. 2018. Portfolio Manager Technical Reference: Climate and Weather. https:\/\/portfoliomanager.energystar.gov\/pdf\/reference\/Climate%20and%20Weather.pdf. Accessed: 2019-06-19."},{"key":"e_1_3_2_1_47_1","volume-title":"Reinforcement learning: An introduction","author":"Sutton Richard S","unstructured":"Richard S Sutton and Andrew G Barto . 2018. Reinforcement learning: An introduction . MIT press . Richard S Sutton and Andrew G Barto. 2018. Reinforcement learning: An introduction. MIT press."},{"key":"e_1_3_2_1_48_1","volume-title":"Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning 4, 2","author":"Tieleman Tijmen","year":"2012","unstructured":"Tijmen Tieleman and Geoffrey Hinton . 2012. Lecture 6.5-rmsprop : Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning 4, 2 ( 2012 ), 26--31. Tijmen Tieleman and Geoffrey Hinton. 2012. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning 4, 2 (2012), 26--31."},{"key":"e_1_3_2_1_49_1","volume-title":"Michelle Yeo, Alireza Makhzani, Heinrich K\u00fcttler, John Agapiou, Julian Schrittwieser, et al.","author":"Vinyals Oriol","year":"2017","unstructured":"Oriol Vinyals , Timo Ewalds , Sergey Bartunov , Petko Georgiev , Alexander Sasha Vezhnevets , Michelle Yeo, Alireza Makhzani, Heinrich K\u00fcttler, John Agapiou, Julian Schrittwieser, et al. 2017 . Starcraft ii: A new challenge for reinforcement learning. arXiv preprint arXiv:1708.04782 (2017). Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich K\u00fcttler, John Agapiou, Julian Schrittwieser, et al. 2017. Starcraft ii: A new challenge for reinforcement learning. arXiv preprint arXiv:1708.04782 (2017)."},{"key":"e_1_3_2_1_50_1","volume-title":"Sample efficient actor-critic with experience replay. arXiv preprint arXiv:1611.01224","author":"Wang Ziyu","year":"2016","unstructured":"Ziyu Wang , Victor Bapst , Nicolas Heess , Volodymyr Mnih , Remi Munos , Koray Kavukcuoglu , and Nando de Freitas . 2016. Sample efficient actor-critic with experience replay. arXiv preprint arXiv:1611.01224 ( 2016 ). Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas. 2016. Sample efficient actor-critic with experience replay. arXiv preprint arXiv:1611.01224 (2016)."},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"crossref","unstructured":"Stephen Wilcox and William Marion. 2008. Users manual for TMY3 data sets. National Renewable Energy Laboratory Golden CO.  Stephen Wilcox and William Marion. 2008. Users manual for TMY3 data sets. National Renewable Energy Laboratory Golden CO.","DOI":"10.2172\/928611"},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.apenergy.2015.07.050"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3276774.3276775"}],"event":{"name":"BuildSys '19: The 6th ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation","location":"New York NY USA","acronym":"BuildSys '19","sponsor":["SIGEnergy ACM Special Interest Group on Energy Systems and Informatics"]},"container-title":["Proceedings of the 6th ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3360322.3360849","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3360322.3360849","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:44:29Z","timestamp":1750189469000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3360322.3360849"}},"subtitle":["A Precocial Reinforcement Learning Solution for Building HVAC Control Using a Differentiable MPC Policy"],"short-title":[],"issued":{"date-parts":[[2019,11,13]]},"references-count":51,"alternative-id":["10.1145\/3360322.3360849","10.1145\/3360322"],"URL":"https:\/\/doi.org\/10.1145\/3360322.3360849","relation":{},"subject":[],"published":{"date-parts":[[2019,11,13]]},"assertion":[{"value":"2019-11-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}