{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T11:07:13Z","timestamp":1775041633731,"version":"3.50.1"},"reference-count":36,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2012,4,1]],"date-time":"2012-04-01T00:00:00Z","timestamp":1333238400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001602","name":"Science Foundation Ireland","doi-asserted-by":"publisher","award":["03\/CE2\/I303_1"],"award-info":[{"award-number":["03\/CE2\/I303_1"]}],"id":[{"id":"10.13039\/501100001602","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Lero-The Irish Software Engineering Research Centre"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Auton. Adapt. Syst."],"published-print":{"date-parts":[[2012,4]]},"abstract":"<jats:p>This article describes Distributed W-Learning (DWL), a reinforcement learning-based algorithm for collaborative agent-based optimization of pervasive systems. DWL supports optimization towards multiple heterogeneous policies and addresses the challenges arising from the heterogeneity of the agents that are charged with implementing them. DWL learns and exploits the dependencies between agents and between policies to improve overall system performance. Instead of always executing the locally-best action, agents learn how their actions affect their immediate neighbors and execute actions suggested by neighboring agents if their importance exceeds the local action's importance when scaled using a predefined or learned collaboration coefficient. We have evaluated DWL in a simulation of an Urban Traffic Control (UTC) system, a canonical example of the large-scale pervasive systems that we are addressing. We show that DWL outperforms widely deployed fixed-time and simple adaptive UTC controllers under a variety of traffic loads and patterns. Our results also confirm that enabling collaboration between agents is beneficial as is the ability for agents to learn the degree to which it is appropriate for them to collaborate. These results suggest that DWL is a suitable basis for optimization in other large-scale systems with similar characteristics.<\/jats:p>","DOI":"10.1145\/2168260.2168271","type":"journal-article","created":{"date-parts":[[2012,5,1]],"date-time":"2012-05-01T13:43:38Z","timestamp":1335879818000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":22,"title":["Autonomic multi-policy optimization in pervasive systems"],"prefix":"10.1145","volume":"7","author":[{"given":"Ivana","family":"Dusparic","sequence":"first","affiliation":[{"name":"Trinity College Dublin"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vinny","family":"Cahill","sequence":"additional","affiliation":[{"name":"Trinity College Dublin"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,5,4]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1061\/(ASCE)0733-947X(2003)129:3(278)"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10458-004-6975-9"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1287\/moor.27.4.819.297"},{"key":"e_1_2_1_4_1","doi-asserted-by":"crossref","unstructured":"Cuayahuitl H. Renals S. Lemon O. and Shimodaira H. 2006. Learning multi-goal dialogue strategies using reinforcement learning with reduced state-action spaces. Int. J. Game Theory 547--565. Cuayahuitl H. Renals S. Lemon O. and Shimodaira H. 2006. Learning multi-goal dialogue strategies using reinforcement learning with reduced state-action spaces. Int. J. Game Theory 547--565.","DOI":"10.21437\/Interspeech.2006-149"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143872"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1017\/S0269888906000956"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/SASO.2009.23"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-02704-8_9"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2004.838180"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the ICML-2002 the 19th International Conference on Machine Learning. 227--234","author":"Guestrin C.","unstructured":"Guestrin , C. , Lagoudakis , M. , and Parr , R . 2002. Coordinated reinforcement learning . In Proceedings of the ICML-2002 the 19th International Conference on Machine Learning. 227--234 . Guestrin, C., Lagoudakis, M., and Parr, R. 2002. Coordinated reinforcement learning. In Proceedings of the ICML-2002 the 19th International Conference on Machine Learning. 227--234."},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 2002 Congress. IEEE Computer Society","author":"Hoar R.","year":"1910","unstructured":"Hoar , R. , Penner , J. , and Jacob , C . 2002. Evolutionary swarm traffic: If ant roads had traffic lights. In (CEC'02) Proceedings of the Evolutionary Computation (CEC '02) . Proceedings of the 2002 Congress. IEEE Computer Society , Washington, DC , 1910 --1915. Hoar, R., Penner, J., and Jacob, C. 2002. Evolutionary swarm traffic: If ant roads had traffic lights. In (CEC'02) Proceedings of the Evolutionary Computation (CEC '02). Proceedings of the 2002 Congress. IEEE Computer Society, Washington, DC, 1910--1915."},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the 4th International Conference on Simulation of Adaptive Behavior. MIT Press, 135--144","author":"Humphrys M.","year":"1996","unstructured":"Humphrys , M. 1996 a. Action selection methods using reinforcement learning . In Proceedings of the 4th International Conference on Simulation of Adaptive Behavior. MIT Press, 135--144 . Humphrys, M. 1996a. Action selection methods using reinforcement learning. In Proceedings of the 4th International Conference on Simulation of Adaptive Behavior. MIT Press, 135--144."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1329125.1329241"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2003.1160055"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the IEEE Symposium on Computational Intelligence and Games (CIG). 29--36","author":"Kok J. R.","unstructured":"Kok , J. R. , 't Hoen , P. J. , Bakker , B. , and Vlassis , N . 2005. Utile coordination: Learning interdependencies among cooperative agents . In Proceedings of the IEEE Symposium on Computational Intelligence and Games (CIG). 29--36 . Kok, J. R., 't Hoen, P. J., Bakker, B., and Vlassis, N. 2005. Utile coordination: Learning interdependencies among cooperative agents. In Proceedings of the IEEE Symposium on Computational Intelligence and Games (CIG). 29--36."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 1st International Conference on Autonomic Computing (ICAC'04)","author":"Littman M. L.","unstructured":"Littman , M. L. , Ravi , N. , Fenson , E. , and Howard , R . 2004. Reinforcement learning for autonomic network repair . In Proceedings of the 1st International Conference on Autonomic Computing (ICAC'04) . IEEE Computer Society, Washington, DC, 284--285. Littman, M. L., Ravi, N., Fenson, E., and Howard, R. 2004. Reinforcement learning for autonomic network repair. In Proceedings of the 1st International Conference on Autonomic Computing (ICAC'04). IEEE Computer Society, Washington, DC, 284--285."},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 8th International Conference on Autonomous Agents and Multi-Agent Systems.","author":"Melo F.","unstructured":"Melo , F. and Veloso , M . 2009. Learning of coordination: Exploiting sparse interactions in multiagent systems . In Proceedings of the 8th International Conference on Autonomous Agents and Multi-Agent Systems. Melo, F. and Veloso, M. 2009. Learning of coordination: Exploiting sparse interactions in multiagent systems. In Proceedings of the 8th International Conference on Autonomous Agents and Multi-Agent Systems."},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the European Simulation and Modelling Conference. 128--135","author":"Oliveira E.","unstructured":"Oliveira , E. and Duarte , N . 2005. Making way for emergency vehicles . In Proceedings of the European Simulation and Modelling Conference. 128--135 . Oliveira, E. and Duarte, N. 2005. Making way for emergency vehicles. In Proceedings of the European Simulation and Modelling Conference. 128--135."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCGRID.2008.33"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-69295-9_19"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1142680.1142682"},{"key":"e_1_2_1_24_1","unstructured":"Richter S. 2006. Learning traffic control - Towards practical traffic control using policy gradients. Tech. rep. Albert-Ludwigs-Universitat Freiburg. Richter S. 2006. Learning traffic control - Towards practical traffic control using policy gradients. Tech. rep. Albert-Ludwigs-Universitat Freiburg."},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Richter S. Aberdeen D. and Yu J. 2007. Natural actor-critic for road traffic optimisation. Adv. Neural Inf. Process. Syst. 19. The MIT Press Cambridge MA. Richter S. Aberdeen D. and Yu J. 2007. Natural actor-critic for road traffic optimisation. Adv. Neural Inf. Process. Syst. 19. The MIT Press Cambridge MA.","DOI":"10.7551\/mitpress\/7503.003.0151"},{"key":"e_1_2_1_26_1","volume-title":"13th International IEEE Conference on Intelligent Transportation System (ITSC '10)","author":"Salkham A.","unstructured":"Salkham , A. and Cahill , V . 2010. Soilse: A decentralized approach to optimization of fluctuating urban traffic using reinforcement learning . In 13th International IEEE Conference on Intelligent Transportation System (ITSC '10) . Salkham, A. and Cahill, V. 2010. Soilse: A decentralized approach to optimization of fluctuating urban traffic using reinforcement learning. In 13th International IEEE Conference on Intelligent Transportation System (ITSC '10)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/WIIAT.2008.88"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the 16th International Conference on Machine Learning. Morgan Kaufmann, 371--378","author":"Schneider J.","unstructured":"Schneider , J. , Wong , W.-K. , Moore , A. , and Riedmiller , M . 1999. Distributed value functions . In Proceedings of the 16th International Conference on Machine Learning. Morgan Kaufmann, 371--378 . Schneider, J., Wong, W.-K., Moore, A., and Riedmiller, M. 1999. Distributed value functions. In Proceedings of the 16th International Conference on Machine Learning. Morgan Kaufmann, 371--378."},{"key":"e_1_2_1_29_1","volume-title":"Reinforcement Learning: An Introduction. A Bradford Book","author":"Suton R. S.","year":"1998","unstructured":"Suton , R. S. and Barto , A. G . 1998 . Reinforcement Learning: An Introduction. A Bradford Book . The MIT Press , Cambridge, MA . Suton, R. S. and Barto, A. G. 1998. Reinforcement Learning: An Introduction. A Bradford Book. The MIT Press, Cambridge, MA."},{"key":"e_1_2_1_30_1","first-page":"2","article-title":"Multiagent systems","volume":"19","author":"Sycara K.","year":"1998","unstructured":"Sycara , K. 1998 . Multiagent systems . AI Mag. 19 , 2 . Sycara, K. 1998. Multiagent systems. AI Mag. 19, 2.","journal-title":"AI Mag."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-1-55860-307-3.50049-6"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/MIC.2007.21"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.5555\/1018409.1018780"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICAC.2005.65"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICAC.2006.1662383"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"e_1_2_1_37_1","unstructured":"Wiering M. van Veenen J. Vreeken J. and Koopman A. 2004. Intelligent traffic light control. Tech. rep. Institute of Information and Computing Sciences Utrecht University. Wiering M. van Veenen J. Vreeken J. and Koopman A. 2004. Intelligent traffic light control. Tech. rep. Institute of Information and Computing Sciences Utrecht University."},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the International Conference on Machine Learning and Cybernetics. 1482--1486","author":"Yang Z.","unstructured":"Yang , Z. , Chen , X. , Tang , Y. , and Sun , J . 2005. Intelligent cooperation control of urban traffic networks . In Proceedings of the International Conference on Machine Learning and Cybernetics. 1482--1486 . Yang, Z., Chen, X., Tang, Y., and Sun, J. 2005. Intelligent cooperation control of urban traffic networks. In Proceedings of the International Conference on Machine Learning and Cybernetics. 1482--1486."}],"container-title":["ACM Transactions on Autonomous and Adaptive Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2168260.2168271","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2168260.2168271","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T10:52:06Z","timestamp":1750243926000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2168260.2168271"}},"subtitle":["Overview and evaluation"],"short-title":[],"issued":{"date-parts":[[2012,4]]},"references-count":36,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2012,4]]}},"alternative-id":["10.1145\/2168260.2168271"],"URL":"https:\/\/doi.org\/10.1145\/2168260.2168271","relation":{},"ISSN":["1556-4665","1556-4703"],"issn-type":[{"value":"1556-4665","type":"print"},{"value":"1556-4703","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,4]]},"assertion":[{"value":"2010-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-05-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}