{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T06:50:32Z","timestamp":1777704632865,"version":"3.51.4"},"reference-count":36,"publisher":"SAGE Publications","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IFS"],"published-print":{"date-parts":[[2020,12,4]]},"abstract":"<jats:p>Industry has always been in the pursuit of becoming more economically efficient and the current focus has been to reduce human labour using modern technologies. Even with cutting edge technologies, which range from packaging robots to AI for fault detection, there is still some ambiguity on the aims of some new systems, namely, whether they are automated or autonomous. In this paper, we indicate the distinctions between automated and autonomous systems as well as review the current literature and identify the core challenges for creating learning mechanisms of autonomous agents. We discuss using different types of extended realities, such as digital twins, how to train reinforcement learning agents to learn specific tasks through generalisation. Once generalisation is achieved, we discuss how these can be used to develop self-learning agents. We then introduce self-play scenarios and how they can be used to teach self-learning agents through a supportive environment that focuses on how the agents can adapt to different environments. We introduce an initial prototype of our ideas by solving a multi-armed bandit problem using two \u03b5-greedy algorithms. Further, we discuss future applications in the industrial management realm and propose a modular architecture for improving the decision-making process via autonomous agents.<\/jats:p>","DOI":"10.3233\/jifs-189161","type":"journal-article","created":{"date-parts":[[2020,8,14]],"date-time":"2020-08-14T13:39:03Z","timestamp":1597412343000},"page":"8427-8439","source":"Crossref","is-referenced-by-count":12,"title":["Autonomous Industrial Management via Reinforcement Learning"],"prefix":"10.1177","volume":"39","author":[{"given":"Leonardo","family":"Espinosa-Leal","sequence":"first","affiliation":[{"name":"Department of Business Management and Analytics, Arcada University of Applied Sciences, Helsinki, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anthony","family":"Chapman","sequence":"additional","affiliation":[{"name":"Computing Department, University of Aberdeen, Aberdeen, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Magnus","family":"Westerlund","sequence":"additional","affiliation":[{"name":"Department of Business Management and Analytics, Arcada University of Applied Sciences, Helsinki, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"key":"10.3233\/JIFS-189161_ref1","unstructured":"Schwab K. , The Fourth Industrial Revolution, Crown Publishing Group, New York, NY, USA, 2017. ISBN ISBN 1524758868, 9781524758868."},{"issue":"1","key":"10.3233\/JIFS-189161_ref2","doi-asserted-by":"crossref","first-page":"283","DOI":"10.3233\/JIFS-179403","article-title":"Framework of industrial networking sensing system based on edge computing and artificial intelligence","volume":"38","author":"Lu","year":"2020","journal-title":"Journal of Intelligent & Fuzzy Systems"},{"issue":"4","key":"10.3233\/JIFS-189161_ref4","doi-asserted-by":"crossref","first-page":"2197","DOI":"10.3233\/JIFS-171186","article-title":"Hierarchical control strategy towards safe driving of autonomous vehicles","volume":"34","author":"Chen","year":"2018","journal-title":"Journal of Intelligent & Fuzzy Systems"},{"key":"10.3233\/JIFS-189161_ref5","unstructured":"Rolls-Royce, Rolls-Royce and Finferries demonstrate world\u2019s first Fully Autonomous Ferry, 2018, Accessed: 2019-07-04."},{"issue":"2","key":"10.3233\/JIFS-189161_ref8","doi-asserted-by":"crossref","first-page":"416","DOI":"10.1109\/TNNLS.2015.2411671","article-title":"A combined adaptive neural network and nonlinear model predictive control for multirate networked industrial process control","volume":"27","author":"Wang","year":"2016","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"issue":"11","key":"10.3233\/JIFS-189161_ref10","doi-asserted-by":"crossref","first-page":"1016","DOI":"10.1016\/j.ifacol.2018.08.474","article-title":"Digital Twin in manufacturing: A categorical literature review and classification","volume":"51","author":"Kritzinger","year":"2018","journal-title":"IFAC-PapersOnLine"},{"key":"10.3233\/JIFS-189161_ref12","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-84882-809-4"},{"issue":"1","key":"10.3233\/JIFS-189161_ref13","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1016\/j.compag.2010.02.003","article-title":"Conceptual model of a future farm management information system","volume":"72","author":"S\u00f8rensen","year":"2010","journal-title":"Computers and Electronics in Agriculture"},{"issue":"7","key":"10.3233\/JIFS-189161_ref15","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1145\/176789.176795","article-title":"KidSim: programming agents without a programming language","volume":"37","author":"Smith","year":"1994","journal-title":"Communications of the ACM"},{"key":"10.3233\/JIFS-189161_ref16","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1016\/j.trf.2015.04.014","article-title":"Public opinion on automated driving: Results of an international questionnaire among 5000 respondents","volume":"32","author":"Kyriakidis","year":"2015","journal-title":"Transportation Research part F: Traffic Psychology and Behaviour"},{"key":"10.3233\/JIFS-189161_ref17","unstructured":"SAE On-Road Automated Vehicle Standards Committee, Taxonomy and definitions for terms related to on-road motor vehicle automated driving systems, SAE International (2014)."},{"key":"10.3233\/JIFS-189161_ref18","unstructured":"Carlsson C. , Fedrizzi M. and Full\u00e9r R. , Fuzzy logic in management, Vol. 66, Springer Science & Business Media, 2012."},{"issue":"1","key":"10.3233\/JIFS-189161_ref19","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1214\/aoms\/1177704255","article-title":"A definition of subjective probability","volume":"34","author":"Anscombe","year":"1963","journal-title":"Annals of Mathematical Statistics"},{"issue":"7553","key":"10.3233\/JIFS-189161_ref23","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"10.3233\/JIFS-189161_ref24","doi-asserted-by":"publisher","DOI":"10.1145\/3197768.3201525"},{"issue":"3","key":"10.3233\/JIFS-189161_ref26","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1109\/LSP.2017.2657381","article-title":"Deep convolutional neural networks and data augmentation for environmental sound classification","volume":"24","author":"Salamon","year":"2017","journal-title":"IEEE Signal Processing Letters"},{"issue":"1","key":"10.3233\/JIFS-189161_ref27","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1162\/089976600300015961","article-title":"Reinforcement learning in continuous time and space","volume":"12","author":"Doya","year":"2000","journal-title":"Neural Computation"},{"key":"10.3233\/JIFS-189161_ref29","unstructured":"Bellman R. , Dynamic Programming, Princeton University Press, 1957."},{"key":"10.3233\/JIFS-189161_ref30","doi-asserted-by":"crossref","first-page":"237","DOI":"10.1613\/jair.301","article-title":"Reinforcement learning: A survey","volume":"4","author":"Kaelbling","year":"1996","journal-title":"Journal of Artificial Intelligence Research"},{"issue":"10","key":"10.3233\/JIFS-189161_ref33","doi-asserted-by":"crossref","first-page":"1345","DOI":"10.1109\/TKDE.2009.191","article-title":"A survey on transfer learning","volume":"22","author":"Pan","year":"2010","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"10.3233\/JIFS-189161_ref34","doi-asserted-by":"crossref","unstructured":"Lake B.M. , Ullman T.D. , Tenenbaum J.B. and Gershman S.J. , Building machines that learn and think like people, Behavioral and Brain Sciences 40 (2017).","DOI":"10.1017\/S0140525X16001837"},{"issue":"1","key":"10.3233\/JIFS-189161_ref35","first-page":"1334","article-title":"End-to-end training of deep visuomotor policies","volume":"17","author":"Levine","year":"2016","journal-title":"The Journal of Machine Learning Research"},{"key":"10.3233\/JIFS-189161_ref42","unstructured":"Sutton R.S. and Barto A.G. , Reinforcement learning: An introduction, Cambridge, MA: MIT Press (2011)."},{"issue":"7540","key":"10.3233\/JIFS-189161_ref43","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"10.3233\/JIFS-189161_ref61","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1016\/j.artint.2018.01.002","article-title":"Autonomous agents modelling other agents: A comprehensive survey and open problems","volume":"258","author":"Albrecht","year":"2018","journal-title":"Artificial Intelligence"},{"key":"10.3233\/JIFS-189161_ref63","unstructured":"Wooldridge M. , An introduction to multiagent systems, John Wiley & Sons, 2009."},{"key":"10.3233\/JIFS-189161_ref64","unstructured":"Michalewicz Z. , Schmidt M. , Michalewicz M. and Chiriac C. , Adaptive business intelligence, Springer, 2006."},{"issue":"3","key":"10.3233\/JIFS-189161_ref66","doi-asserted-by":"publisher","first-page":"285","DOI":"10.1023\/A:1010071910869","article-title":"The Gaia Methodology for Agent-Oriented Analysis and Design","volume":"3","author":"Wooldridge","year":"2000","journal-title":"Autonomous Agents and Multi-Agent Systems"},{"issue":"6","key":"10.3233\/JIFS-189161_ref68","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1016\/S0306-4379(02)00012-1","article-title":"Towards requirementsdriven information systems engineering: the Tropos project","volume":"27","author":"Castro","year":"2002","journal-title":"Information systems"},{"key":"10.3233\/JIFS-189161_ref69","doi-asserted-by":"crossref","unstructured":"Bordini R.H. , H\u00fcbner J.F. and Wooldridge M. , Programming multi-agent systems in AgentSpeak using Jason, Vol. 8, John Wiley & Sons, 2007.","DOI":"10.1002\/9780470061848"},{"issue":"6","key":"10.3233\/JIFS-189161_ref73","doi-asserted-by":"crossref","first-page":"1299","DOI":"10.1080\/00207540110118640","article-title":"Global supply chain management: a reinforcement learning approach","volume":"40","author":"Pontrandolfo","year":"2002","journal-title":"International Journal of Production Research"},{"issue":"1\u20133","key":"10.3233\/JIFS-189161_ref75","doi-asserted-by":"crossref","first-page":"519","DOI":"10.1016\/S0925-5273(98)00114-5","article-title":"Models for warehouse management: Classification and examples","volume":"59","author":"Van den Berg","year":"1999","journal-title":"International Journal of Production Economics"},{"issue":"3","key":"10.3233\/JIFS-189161_ref76","doi-asserted-by":"crossref","first-page":"476","DOI":"10.1109\/TII.2011.2158834","article-title":"Optimizing warehouse forklift dispatching using a sensor network and stochastic learning","volume":"7","author":"Estanjini","year":"2011","journal-title":"IEEE Transactions on Industrial Informatics"},{"issue":"3\u20134","key":"10.3233\/JIFS-189161_ref77","doi-asserted-by":"crossref","first-page":"197","DOI":"10.1002\/nav.21481","article-title":"A least squares temporal difference actor\u2013critic algorithm with applications to warehouse management","volume":"59","author":"Estanjini","year":"2012","journal-title":"Naval Research Logistics (NRL)"},{"issue":"3","key":"10.3233\/JIFS-189161_ref81","doi-asserted-by":"crossref","first-page":"567","DOI":"10.1016\/j.ifacol.2015.06.141","article-title":"About the importance of autonomy and digital twins for the future of manufacturing","volume":"48","author":"Rosen","year":"2015","journal-title":"IFAC-PapersOnLine"},{"key":"10.3233\/JIFS-189161_ref83","unstructured":"Wymore A.W. , Model-based systems engineering, CRC press, 1993."}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/JIFS-189161","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:41:40Z","timestamp":1777455700000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/JIFS-189161"}},"subtitle":["Towards Self-Learning Agents for Decision-Making"],"editor":[{"given":"Vijayakumar","family":"Varadarajan","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]},{"given":"Piet","family":"Kommers","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]},{"given":"Vincenzo","family":"Piuri","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]},{"given":"V.","family":"Subramaniyaswamy","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2020,12,4]]},"references-count":36,"journal-issue":{"issue":"6"},"URL":"https:\/\/doi.org\/10.3233\/jifs-189161","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,12,4]]}}}