{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T16:51:38Z","timestamp":1784911898653,"version":"3.55.0"},"reference-count":54,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,3,3]],"date-time":"2023-03-03T00:00:00Z","timestamp":1677801600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Systems"],"abstract":"<jats:p>To conserve building energy, optimal operation of a building\u2019s energy systems, especially heating, ventilation and air-conditioning (HVAC) systems, is important. This study focuses on the optimization of the central chiller plant, which accounts for a large portion of the HVAC system\u2019s energy consumption. Classic optimal control methods for central chiller plants are mostly based on system performance models which takes much effort and cost to establish. In addition, inevitable model error could cause control risk to the applied system. To mitigate the model dependency of HVAC optimal control, reinforcement learning (RL) algorithms have been drawing attention in the HVAC control domain due to its model-free feature. Currently, the RL-based optimization of central chiller plants faces several challenges: (1) existing model-free control methods based on RL typically adopt single-agent scheme, which brings high training cost and long training period when optimizing multiple controllable variables for large-scaled systems; (2) multi-agent scheme could overcome the former problem, but it also requires a proper coordination mechanism to harmonize the potential conflicts among all involved RL agents; (3) previous agent coordination frameworks (identified by distributed control or decentralized control) are mainly designed for model-based control methods instead of model-free controllers. To tackle the problems above, this article proposes a multi-agent, model-free optimal control approach for central chiller plants. This approach utilizes game theory and the RL algorithm SARSA for agent coordination and learning, respectively. A data-driven system model is set up using measured field data of a real HVAC system for simulation. The simulation case study results suggest that the energy saving performance (both short- and long-term) of the proposed approach (over 10% in a cooling season compared to the rule-based baseline controller) is close to the classic multi-agent reinforcement learning (MARL) algorithm WoLF-PHC; moreover, the proposed approach\u2019s nature of few pending parameters makes it more feasible and robust for engineering practices than the WoLF-PHC algorithm.<\/jats:p>","DOI":"10.3390\/systems11030136","type":"journal-article","created":{"date-parts":[[2023,3,3]],"date-time":"2023-03-03T02:56:50Z","timestamp":1677812210000},"page":"136","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["Multi-Agent Optimal Control for Central Chiller Plants Using Reinforcement Learning and Game Theory"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2119-7722","authenticated-orcid":false,"given":"Shunian","family":"Qiu","sequence":"first","affiliation":[{"name":"School of Civil Engineering and Architecture, Zhejiang University of Science and Technology, Hangzhou 310023, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhenhai","family":"Li","sequence":"additional","affiliation":[{"name":"School of Mechanical Engineering, Tongji University, Shanghai 200092, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5506-9388","authenticated-orcid":false,"given":"Zhihong","family":"Pang","sequence":"additional","affiliation":[{"name":"Department of Construction Management, Louisiana State University, Patrick F. Taylor Hall 3315-D, Baton Rouge, LA 70803, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhengwei","family":"Li","sequence":"additional","affiliation":[{"name":"School of Mechanical Engineering, Tongji University, Shanghai 200092, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2231-8314","authenticated-orcid":false,"given":"Yinying","family":"Tao","sequence":"additional","affiliation":[{"name":"School of Design and Fashion, Zhejiang University of Science and Technology, Hangzhou 310023, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,3]]},"reference":[{"key":"ref_1","unstructured":"Delmastro, C., De Bienassis, T., Goodson, T., Lane, K., Le Marois, J.-B., Martinez-Gordon, R., and Husek, M. (2022). Buildings: Tracking Progress 2022, International Energy Agency."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1080\/10789669.2008.10390991","article-title":"Supervisory and Optimal Control of Building HVAC Systems: A Review","volume":"14","author":"Wang","year":"2008","journal-title":"Hvac R Res."},{"key":"ref_3","unstructured":"Commercial Buildings Energy Consumption Survey (CBECS) (2012). 2012 CBECS Survey Data."},{"key":"ref_4","unstructured":"Taylor, S.T. (2017). Fundamentals of Design and Control of Central Chilled-Water Plants, ASHRAE Learning Institute."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Qiu, S., Li, Z., Li, Z., and Wu, Q. (2022). Comparative Evaluation of Different Multi-Agent Reinforcement Learning Mechanisms in Condenser Water System Control. Buildings, 12.","DOI":"10.3390\/buildings12081092"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1016\/j.epsr.2003.10.012","article-title":"A novel energy conservation method\u2014Optimal chiller loading","volume":"69","author":"Chang","year":"2004","journal-title":"Electr. Power Syst. Res."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1016\/j.egypro.2017.07.375","article-title":"Decentralized control of parallel-connected chillers","volume":"122","author":"Dai","year":"2017","journal-title":"Energy Procedia"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"203","DOI":"10.1016\/j.enbuild.2014.07.072","article-title":"Stochastic chiller sequencing control","volume":"84","author":"Li","year":"2014","journal-title":"Energy Build."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"261","DOI":"10.1016\/j.enbuild.2019.06.016","article-title":"Cooling load forecasting-based predictive optimisation for chiller plants","volume":"198","author":"Wang","year":"2019","journal-title":"Energy Build."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"102246","DOI":"10.1016\/j.jobe.2021.102246","article-title":"Data mining approach for improving the optimal control of HVAC systems: An event-driven strategy","volume":"39","author":"Wang","year":"2021","journal-title":"J. Build. Eng."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1444","DOI":"10.1016\/j.apenergy.2019.01.170","article-title":"Online chiller loading strategy based on the near-optimal performance map for energy conservation","volume":"238","author":"Wang","year":"2019","journal-title":"Appl. Energy"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"104159","DOI":"10.1016\/j.jobe.2022.104159","article-title":"Real-time optimal control of HVAC systems: Model accuracy and optimization reward","volume":"50","author":"Hou","year":"2022","journal-title":"J. Build. Eng."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"111694","DOI":"10.1016\/j.enbuild.2021.111694","article-title":"Chilled water temperature resetting using model-free reinforcement learning: Engineering application","volume":"255","author":"Qiu","year":"2022","journal-title":"Energy Build."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"540","DOI":"10.1016\/j.enbuild.2013.08.050","article-title":"An optimal control strategy with enhanced robustness for air-conditioning systems considering model and measurement uncertainties","volume":"67","author":"Zhu","year":"2013","journal-title":"Energy Build."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"100020","DOI":"10.1016\/j.egyai.2020.100020","article-title":"Reinforcement learning for whole-building HVAC control and demand response","volume":"2","author":"Azuatalam","year":"2020","journal-title":"Energy AI"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1080\/10789669.2003.10391069","article-title":"Evaluation of Reinforcement Learning Control for Thermal Energy Storage Systems","volume":"9","author":"Henze","year":"2003","journal-title":"HVAC R Res."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"115036","DOI":"10.1016\/j.apenergy.2020.115036","article-title":"Reinforcement learning for building controls: The opportunities and challenges","volume":"269","author":"Wang","year":"2020","journal-title":"Appl. Energy"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1016\/j.enbuild.2005.06.001","article-title":"Experimental analysis of simulated reinforcement learning control for active and passive building thermal storage inventory. Part 2: Results and analysis","volume":"38","author":"Liu","year":"2006","journal-title":"Energy Build."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"109747","DOI":"10.1016\/j.buildenv.2022.109747","article-title":"Towards self-learning control of HVAC systems with the consideration of dynamic occupancy patterns: Application of model-free deep reinforcement learning","volume":"226","author":"Haghighat","year":"2022","journal-title":"Build. Environ."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1016\/S0378-7788(00)00114-6","article-title":"EnergyPlus: Creating a new-generation building energy simulation program","volume":"33","author":"Crawley","year":"2001","journal-title":"Energy Build."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"981","DOI":"10.1016\/j.energy.2016.08.081","article-title":"Electricity demand response in China: Status, feasible market schemes and pilots","volume":"114","author":"Li","year":"2016","journal-title":"Energy"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"110490","DOI":"10.1016\/j.enbuild.2020.110490","article-title":"Application of two promising Reinforcement Learning algorithms for load shifting in a cooling supply system","volume":"229","author":"Schreiber","year":"2020","journal-title":"Energy Build."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"124857","DOI":"10.1016\/j.energy.2022.124857","article-title":"A multi-step predictive deep reinforcement learning algorithm for HVAC control systems in smart buildings","volume":"259","author":"Liu","year":"2022","journal-title":"Energy"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"112778","DOI":"10.1016\/j.enbuild.2023.112778","article-title":"Reinforcement learning control strategy for differential pressure setpoint in large-scale multi-source looped district cooling system","volume":"282","author":"Wang","year":"2023","journal-title":"Energy Build."},{"key":"ref_25","unstructured":"Qiu, S., Li, Z., and Li, Z. (2021, January 28\u201330). Model-Free Optimal Control Method for Chilled Water Pumps Based on Multi-objective Optimization: Engineering Application. Proceedings of the 2021 ASHRAE Virtual Conference, Phoenix, AZ, USA."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"112284","DOI":"10.1016\/j.enbuild.2022.112284","article-title":"Optimal control method of HVAC based on multi-agent deep reinforcement learning","volume":"270","author":"Fu","year":"2022","journal-title":"Energy Build."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"124817","DOI":"10.1016\/j.jclepro.2020.124817","article-title":"A decentralized peer-to-peer control scheme for heating and cooling trading in distributed energy systems","volume":"285","author":"Li","year":"2021","journal-title":"J. Clean. Prod."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"104498","DOI":"10.1016\/j.jobe.2022.104498","article-title":"A general multi agent-based distributed framework for optimal control of building HVAC systems","volume":"52","author":"Wang","year":"2022","journal-title":"J. Build. Eng."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"115371","DOI":"10.1016\/j.apenergy.2020.115371","article-title":"A multi-agent based distributed approach for optimal control of multi-zone ventilation systems considering indoor air quality and energy use","volume":"275","author":"Li","year":"2020","journal-title":"Appl. Energy"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"103919","DOI":"10.1016\/j.autcon.2021.103919","article-title":"An event-driven multi-agent based distributed optimal control strategy for HVAC systems in IoT-enabled smart buildings","volume":"132","author":"Li","year":"2021","journal-title":"Autom. Constr."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1015","DOI":"10.1007\/s12273-021-0869-5","article-title":"A non-cooperative game-based distributed optimization method for chiller plant control","volume":"15","author":"Li","year":"2022","journal-title":"Build. Simul."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"105689","DOI":"10.1016\/j.jobe.2022.105689","article-title":"Deep clustering of cooperative multi-agent reinforcement learning to optimize multi chiller HVAC systems for smart buildings energy management","volume":"65","author":"Homod","year":"2023","journal-title":"J. Build. Eng."},{"key":"ref_33","unstructured":"Zhang, K., Yang, Z., and Baar, T. (2019). Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms. arXiv."},{"key":"ref_34","unstructured":"Fudenberg, D., and Tirole, J. (1991). Game Theory, The MIT Press. [1st ed.]."},{"key":"ref_35","unstructured":"Myerson, R.B. (1997). Game Theory: Analysis of Conflict, Harvard University Press."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"119198","DOI":"10.1016\/j.apenergy.2022.119198","article-title":"An online robust sequencing control strategy for identical chillers using a probabilistic approach concerning flow measurement uncertainties","volume":"317","author":"Sun","year":"2022","journal-title":"Appl. Energy"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1016\/j.enbuild.2016.09.049","article-title":"Event-driven optimization of complex HVAC systems","volume":"133","author":"Wang","year":"2016","journal-title":"Energy Build."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"2177","DOI":"10.1016\/j.enbuild.2008.06.010","article-title":"A novel approach for optimal chiller loading using particle swarm optimization","volume":"40","author":"Ardakani","year":"2008","journal-title":"Energy Build."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"1730","DOI":"10.1016\/j.applthermaleng.2008.08.004","article-title":"Optimal chiller loading by particle swarm algorithm for reducing energy consumption","volume":"29","author":"Lee","year":"2009","journal-title":"Appl. Therm. Eng."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"147","DOI":"10.1016\/j.enbuild.2004.06.002","article-title":"Optimal chiller loading by genetic algorithm for reducing energy consumption","volume":"37","author":"Chang","year":"2005","journal-title":"Energy Build."},{"key":"ref_41","first-page":"2","article-title":"Near-optimal control of cooling towers for chilled-water systems","volume":"96","author":"Braun","year":"1990","journal-title":"ASHRAE Trans."},{"key":"ref_42","first-page":"67","article-title":"Integrated Multi-objective Optimization of Predictive Maintenance and Production Scheduling: Perspective from Lead Time Constraints","volume":"1","author":"Zhao","year":"2022","journal-title":"J. Intell. Manag. Decis."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"149","DOI":"10.1016\/j.enbuild.2019.05.006","article-title":"Stochastic optimized chiller operation strategy based on multi-objective optimization considering measurement uncertainty","volume":"195","author":"Qiu","year":"2019","journal-title":"Energy Build."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1017\/S0269888912000057","article-title":"Independent reinforcement learners in cooperative Markov games: A survey regarding coordination problems","volume":"27","author":"Matignon","year":"2012","journal-title":"Knowl. Eng. Rev."},{"key":"ref_45","unstructured":"Rummery, G., and Niranjan, M. (1994). On-Line Q-Learning Using Connectionist Systems, University of Cambridge, Department of Engineering. Technical Report CUED\/F-INFENG\/TR 166."},{"key":"ref_46","first-page":"2825","article-title":"Scikit-learn: Machine Learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Tao, J.Y., and Li, D.S. (2006). Cooperative Strategy Learning in Multi-Agent Environment with Continuous State Space, IEEE.","DOI":"10.1109\/ICMLC.2006.258352"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"977","DOI":"10.1016\/j.energy.2018.04.042","article-title":"Smart generation control based on multi-agent reinforcement learning with the idea of the time tunnel","volume":"153","author":"Xi","year":"2018","journal-title":"Energy"},{"key":"ref_49","unstructured":"Lauer, M. (July, January 29). An algorithm for distributed reinforcement learning in cooperative multiagent systems. Proceedings of the 17th International Conference on Machine Learning, Stanford, CA, USA."},{"key":"ref_50","unstructured":"Cohen, W.W., and Hirsh, H. (1994). Machine Learning Proceedings 1994, Morgan Kaufmann."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Srinivasan, D., and Jain, L.C. (2010). Multi-Agent Reinforcement Learning: An Overview, in Innovations in Multi-Agent Systems and Applications\u20141, Springer.","DOI":"10.1007\/978-3-642-14435-6"},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"215","DOI":"10.1016\/S0004-3702(02)00121-2","article-title":"Multiagent learning using a variable learning rate","volume":"136","author":"Bowling","year":"2002","journal-title":"Artif. Intell."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1016\/j.enconman.2015.06.030","article-title":"A novel multi-agent decentralized win or learn fast policy hill-climbing with eligibility trace algorithm for smart generation control of interconnected complex power grids","volume":"103","author":"Xi","year":"2015","journal-title":"Energy Convers. Manag."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"1100","DOI":"10.1080\/23744731.2020.1757328","article-title":"Model-free optimal chiller loading method based on Q-learning","volume":"26","author":"Qiu","year":"2020","journal-title":"Sci. Technol. Built Environ."}],"container-title":["Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2079-8954\/11\/3\/136\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:46:52Z","timestamp":1760122012000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2079-8954\/11\/3\/136"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,3]]},"references-count":54,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["systems11030136"],"URL":"https:\/\/doi.org\/10.3390\/systems11030136","relation":{},"ISSN":["2079-8954"],"issn-type":[{"value":"2079-8954","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,3]]}}}