{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,12]],"date-time":"2026-07-12T10:20:55Z","timestamp":1783851655158,"version":"3.55.0"},"reference-count":64,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2022,2,9]],"date-time":"2022-02-09T00:00:00Z","timestamp":1644364800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,2,9]],"date-time":"2022-02-09T00:00:00Z","timestamp":1644364800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100010661","name":"Horizon 2020 Framework Programme","doi-asserted-by":"crossref","award":["886977"],"award-info":[{"award-number":["886977"]}],"id":[{"id":"10.13039\/100010661","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Auton Robot"],"published-print":{"date-parts":[[2022,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper presents a learning-based method that uses simulation data to learn an object manipulation task using two model-free reinforcement learning (RL) algorithms. The learning performance is compared across on-policy and off-policy algorithms: Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC). In order to accelerate the learning process, the fine-tuning procedure is proposed that demonstrates the continuous adaptation of on-policy RL to new environments, allowing the learned policy to adapt and execute the (partially) modified task. A dense reward function is designed for the task to enable an efficient learning of the agent. A grasping task involving a Franka Emika Panda manipulator is considered as the reference task to be learned. The learned control policy is demonstrated to be generalizable across multiple object geometries and initial robot\/parts configurations. The approach is finally tested on a real Franka Emika Panda robot, showing the possibility to transfer the learned behavior from simulation. Experimental results show 100% of successful grasping tasks, making the proposed approach applicable to real applications.<\/jats:p>","DOI":"10.1007\/s10514-022-10034-z","type":"journal-article","created":{"date-parts":[[2022,2,9]],"date-time":"2022-02-09T16:02:46Z","timestamp":1644422566000},"page":"483-498","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":63,"title":["Continuous control actions learning and adaptation for robotic manipulation through reinforcement learning"],"prefix":"10.1007","volume":"46","author":[{"given":"Asad Ali","family":"Shahid","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dario","family":"Piga","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Francesco","family":"Braghin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4427-536X","authenticated-orcid":false,"given":"Loris","family":"Roveda","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,2,9]]},"reference":[{"key":"10034_CR1","doi-asserted-by":"crossref","unstructured":"Abbeel, P., Coates, A., Quigley, M., & Ng, A. Y. (2007). An application of reinforcement learning to aerobatic helicopter flight. In Advances in neural information processing systems (pp. 1\u20138).","DOI":"10.7551\/mitpress\/7503.003.0006"},{"key":"10034_CR2","unstructured":"Achiam, J. (2018). Spinning up in deep reinforcement learning. https:\/\/spinningup.openai.com\/en\/latest\/algorithms\/sac.html"},{"key":"10034_CR3","unstructured":"Boularias, A., Kober, J., & Peters, J. (2011). Relative entropy inverse reinforcement learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics (pp. 182\u2013189)."},{"key":"10034_CR4","doi-asserted-by":"crossref","unstructured":"Chebotar, Y., Kalakrishnan, M., Yahya, A., Li, A., Schaal, S., & Levine, S. (2017). Path integral guided policy search. In 2017 IEEE international conference on robotics and automation (ICRA) (pp. 3381\u20133388). IEEE.","DOI":"10.1109\/ICRA.2017.7989384"},{"key":"10034_CR5","doi-asserted-by":"crossref","unstructured":"Cui, F., Cui, Q., & Song, Y. (2020). A survey on learning-based approaches for modeling and classification of human-machine dialog systems. IEEE Transactions on Neural Networks and Learning Systems","DOI":"10.1109\/TNNLS.2020.2985588"},{"key":"10034_CR6","unstructured":"Deisenroth, M., & Rasmussen, C. E. (2011). Pilco: A model-based and data-efficient approach to policy search. In Proceedings of the 28th international conference on machine learning (ICML-11) (pp. 465\u2013472)."},{"key":"10034_CR7","doi-asserted-by":"crossref","unstructured":"Deisenroth, M. P., Neumann, G., & Peters, J. (2013). A survey on policy search for robotics. Now publishers.","DOI":"10.1109\/ICRA.2014.6907421"},{"key":"10034_CR8","unstructured":"Duan, Y., Chen, X., Houthooft, R., Schulman, J., & Abbeel, P. (2016). Benchmarking deep reinforcement learning for continuous control. In International conference on machine learning (pp. 1329\u20131338)."},{"key":"10034_CR9","unstructured":"Fan, L., Zhu, Y., Zhu, J., Liu, Z., Zeng, O., Gupta, A., Creus-Costa, J., Savarese, S., & Fei-Fei, L. (2018). Surreal: Open-source reinforcement learning framework and robot manipulation benchmark. In Conference on robot learning (pp. 767\u2013782)."},{"key":"10034_CR10","doi-asserted-by":"crossref","unstructured":"Gu, S., Holly, E., Lillicrap, T., & Levine, S. (2017). Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In 2017 IEEE international conference on robotics and automation (ICRA) (pp. 3389\u20133396). IEEE.","DOI":"10.1109\/ICRA.2017.7989385"},{"key":"10034_CR11","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018a). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. arXiv preprint arXiv:180101290."},{"key":"10034_CR12","unstructured":"Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et\u00a0al. (2018b). Soft actor-critic algorithms and applications. arXiv preprint arXiv:181205905"},{"key":"10034_CR13","unstructured":"Heess, N., TB, D., Sriram, S., Lemmon, J., Merel, J., Wayne, G., Tassa, Y., Erez, T., Wang, Z., Eslami, S., et\u00a0al. (2017). Emergence of locomotion behaviours in rich environments. arXiv preprint arXiv:170702286"},{"key":"10034_CR14","unstructured":"Julian, R., Swanson, B., Sukhatme, G. S., Levine, S., Finn, C., & Hausman, K. (2020). Never stop learning: The effectiveness of fine-tuning in robotic reinforcement learning. arXiv preprint arXiv:200410190."},{"key":"10034_CR15","doi-asserted-by":"crossref","unstructured":"Kalakrishnan, M., Righetti, L., Pastor, P., & Schaal, S. (2011). Learning force control policies for compliant manipulation. In 2011 IEEE\/RSJ international conference on intelligent robots and systems (pp. 4639\u20134644). IEEE.","DOI":"10.1109\/IROS.2011.6095096"},{"key":"10034_CR16","unstructured":"Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et\u00a0al. (2018). Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation. arXiv preprint arXiv:180610293"},{"key":"10034_CR17","doi-asserted-by":"crossref","unstructured":"Kearney, K. T., Presenza, D., Sacc\u00e0, F., & Wright, P. (2018). Key challenges for developing a socially assistive robotic (SAR) solution for the health sector. In 2018 IEEE 23rd international workshop on computer aided modeling and design of communication links and networks (CAMAD) (pp. 1\u20137). IEEE","DOI":"10.1109\/CAMAD.2018.8515005"},{"key":"10034_CR18","unstructured":"Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:14126980."},{"key":"10034_CR19","unstructured":"Konda, V. R., & Tsitsiklis, J. N. (2000). Actor-critic algorithms. In Advances in neural information processing systems (pp. 1008\u20131014)."},{"key":"10034_CR20","doi-asserted-by":"crossref","unstructured":"Kormushev, P., Calinon, S., & Caldwell, D. G. (2010). Robot motor skill coordination with em-based reinforcement learning. In 2010 IEEE\/RSJ international conference on intelligent robots and systems (pp. 3232\u20133237). IEEE.","DOI":"10.1109\/IROS.2010.5649089"},{"key":"10034_CR21","doi-asserted-by":"crossref","unstructured":"Kornblith, S., Shlens, J., & Le, Q. V. (2019). Do better imagenet models transfer better? In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 2661\u20132671).","DOI":"10.1109\/CVPR.2019.00277"},{"key":"10034_CR22","doi-asserted-by":"crossref","unstructured":"Lasi, H., Fettke, P., Kemper, H. G., Feld, T., & Hoffmann, M. (2014). Industry 4.0. Business & Information Systems Engineering, 6(4), 239\u2013242.","DOI":"10.1007\/s12599-014-0334-4"},{"key":"10034_CR23","unstructured":"Lazaric, A., Restelli, M., & Bonarini, A. (2008). Reinforcement learning in continuous action spaces through sequential monte carlo methods. In Advances in neural information processing systems (pp. 833\u2013840)."},{"key":"10034_CR24","unstructured":"Lee, Y., Yang, J., & Lim, J. J. (2019). Learning to coordinate manipulation skills via skill behavior diversification. In International conference on learning representations."},{"issue":"4\u20135","key":"10034_CR25","doi-asserted-by":"publisher","first-page":"705","DOI":"10.1177\/0278364914549607","volume":"34","author":"I Lenz","year":"2015","unstructured":"Lenz, I., Lee, H., & Saxena, A. (2015). Deep learning for detecting robotic grasps. The International Journal of Robotics Research, 34(4\u20135), 705\u2013724.","journal-title":"The International Journal of Robotics Research"},{"key":"10034_CR26","unstructured":"Levine, S., & Abbeel, P. (2014). Learning neural network policies with guided policy search under unknown dynamics. In Advances in neural information processing systems (pp. 1071\u20131079)."},{"issue":"1","key":"10034_CR27","first-page":"1334","volume":"17","author":"S Levine","year":"2016","unstructured":"Levine, S., Finn, C., Darrell, T., & Abbeel, P. (2016). End-to-end training of deep visuomotor policies. The Journal of Machine Learning Research, 17(1), 1334\u20131373.","journal-title":"The Journal of Machine Learning Research"},{"issue":"4\u20135","key":"10034_CR28","doi-asserted-by":"publisher","first-page":"421","DOI":"10.1177\/0278364917710318","volume":"37","author":"S Levine","year":"2018","unstructured":"Levine, S., Pastor, P., Krizhevsky, A., Ibarz, J., & Quillen, D. (2018). Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection. The International Journal of Robotics Research, 37(4\u20135), 421\u2013436.","journal-title":"The International Journal of Robotics Research"},{"key":"10034_CR29","doi-asserted-by":"publisher","unstructured":"Levine, S., Wagener, N., & Abbeel, P. (2015). Learning contact-rich manipulation skills with guided policy search. In Proceedings\u2014IEEE international conference on robotics and automation, 2015. https:\/\/doi.org\/10.1109\/ICRA.2015.7138994","DOI":"10.1109\/ICRA.2015.7138994"},{"key":"10034_CR30","unstructured":"Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv preprint arXiv:150902971"},{"key":"10034_CR31","doi-asserted-by":"crossref","unstructured":"Mahler, J., Matl, M., Liu, X., Li, A., Gealy, D., & Goldberg, K. (2017). Dex-net 3.0: Computing robust robot vacuum suction grasp targets in point clouds using a new analytic model and deep learning. arXiv preprint arXiv:170906670.","DOI":"10.1109\/ICRA.2018.8460887"},{"key":"10034_CR32","doi-asserted-by":"crossref","unstructured":"Mart\u00edn-Mart\u00edn, R., Lee, M. A., Gardner, R., Savarese, S., Bohg, J., & Garg, A. (2019). Variable impedance control in end-effector space: An action space for reinforcement learning in contact-rich tasks. arXiv preprint arXiv:190608880","DOI":"10.1109\/IROS40897.2019.8968201"},{"issue":"7540","key":"10034_CR33","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529\u2013533.","journal-title":"Nature"},{"key":"10034_CR34","doi-asserted-by":"crossref","unstructured":"Morrison, D., Corke, P., & Leitner, J. (2018). Closing the loop for robotic grasping: A real-time, generative grasp synthesis approach. arXiv preprint arXiv:180405172","DOI":"10.15607\/RSS.2018.XIV.021"},{"issue":"3","key":"10034_CR35","doi-asserted-by":"publisher","first-page":"263","DOI":"10.1177\/0278364912472380","volume":"32","author":"K M\u00fclling","year":"2013","unstructured":"M\u00fclling, K., Kober, J., Kroemer, O., & Peters, J. (2013). Learning to select and generalize striking movements in robot table tennis. The International Journal of Robotics Research, 32(3), 263\u2013279.","journal-title":"The International Journal of Robotics Research"},{"key":"10034_CR36","doi-asserted-by":"crossref","unstructured":"Pastor, P., Kalakrishnan, M., Chitta, S., Theodorou, E., & Schaal, S. (2011). Skill learning and task outcome prediction for manipulation. In 2011 IEEE international conference on robotics and automation (pp. 3828\u20133834). IEEE.","DOI":"10.1109\/ICRA.2011.5980200"},{"key":"10034_CR37","first-page":"8026","volume":"32","author":"A Paszke","year":"2019","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019). Pytorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems, 32, 8026\u20138037.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"7\u20139","key":"10034_CR38","doi-asserted-by":"publisher","first-page":"1180","DOI":"10.1016\/j.neucom.2007.11.026","volume":"71","author":"J Peters","year":"2008","unstructured":"Peters, J., & Schaal, S. (2008). Natural actor-critic. Neurocomputing, 71(7\u20139), 1180\u20131190.","journal-title":"Neurocomputing"},{"key":"10034_CR39","doi-asserted-by":"crossref","unstructured":"Quillen, D., Jang, E., Nachum, O., Finn, C., Ibarz, J., & Levine, S. (2018). Deep reinforcement learning for vision-based robotic grasping: A simulated comparative evaluation of off-policy methods. In 2018 IEEE international conference on robotics and automation (ICRA) (pp. 6284\u20136291). IEEE.","DOI":"10.1109\/ICRA.2018.8461039"},{"key":"10034_CR40","doi-asserted-by":"crossref","unstructured":"Rahmatizadeh, R., Abolghasemi, P., B\u00f6l\u00f6ni, L., & Levine, S. (2018). Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration. In 2018 IEEE international conference on robotics and automation (ICRA). (pp. 3758\u20133765). IEEE.","DOI":"10.1109\/ICRA.2018.8461076"},{"key":"10034_CR41","doi-asserted-by":"crossref","unstructured":"Rajan, K., & Saffiotti, A. (2017). Towards a science of integrated ai and robotics","DOI":"10.1016\/j.artint.2017.03.003"},{"key":"10034_CR42","doi-asserted-by":"crossref","unstructured":"Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., & Levine, S. (2017). Learning complex dexterous manipulation with deep reinforcement learning and demonstrations. arXiv preprint arXiv:170910087.","DOI":"10.15607\/RSS.2018.XIV.049"},{"key":"10034_CR43","doi-asserted-by":"publisher","DOI":"10.1016\/j.conengprac.2020.104488","volume":"101","author":"L Roveda","year":"2020","unstructured":"Roveda, L., Forgione, M., & Piga, D. (2020). Robot control parameters auto-tuning in trajectory tracking applications. Control Engineering Practice, 101, 104488.","journal-title":"Control Engineering Practice"},{"key":"10034_CR44","doi-asserted-by":"crossref","unstructured":"Roveda, L., Maroni, M., Mazzuchelli, L., Praolini, L., Bucca, G., & Piga, D. (2021). Enhancing object detection performance through sensor pose definition with bayesian optimization. In 2021 IEEE international workshop on metrology for Industry 4.0 & IoT (MetroInd4. 0&IoT) (pp. 699\u2013703). IEEE.","DOI":"10.1109\/MetroInd4.0IoT51437.2021.9488517"},{"key":"10034_CR45","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M., & Moritz, P. (2015a). Trust region policy optimization. In International conference on machine learning (pp. 1889\u20131897)."},{"key":"10034_CR46","unstructured":"Schulman, J., Moritz, P., Levine, S., Jordan, M., & Abbeel, P. (2015b). High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:150602438."},{"key":"10034_CR47","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:170706347."},{"issue":"2","key":"10034_CR48","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1016\/J.ENG.2016.02.010","volume":"2","author":"HG Seif","year":"2016","unstructured":"Seif, H. G., & Hu, X. (2016). Autonomous driving in the icity-hd maps as a key challenge of the automotive industry. Engineering, 2(2), 159\u2013162.","journal-title":"Engineering"},{"key":"10034_CR49","unstructured":"Shahid, A. A. (2020). Github repository: Intelligent-task-learning. https:\/\/github.com\/Asad-Shahid\/Intelligent-Task-Learning"},{"issue":"21","key":"10034_CR50","doi-asserted-by":"publisher","first-page":"10227","DOI":"10.3390\/app112110227","volume":"11","author":"AA Shahid","year":"2021","unstructured":"Shahid, A. A., Sesin, J. S. V., Pecioski, D., Braghin, F., Piga, D., & Roveda, L. (2021). Decentralized multi-agent control of a manipulator in continuous task learning. Applied Sciences, 11(21), 10227.","journal-title":"Applied Sciences"},{"key":"10034_CR51","volume-title":"Robot Force Control","author":"B Siciliano","year":"2000","unstructured":"Siciliano, B., & Villani, L. (2000). Robot Force Control (1st ed.). Norwell: Kluwer Academic Publishers.","edition":"1"},{"issue":"7587","key":"10034_CR52","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016). Mastering the game of go with deep neural networks and tree search. Nature, 529(7587), 484.","journal-title":"Nature"},{"key":"10034_CR53","unstructured":"Smith, L., Kew, J. C., Peng, X. B., Ha, S., Tan, J., & Levine, S. (2021). Legged robots that keep on learning: Fine-tuning locomotion policies in the real world. arXiv preprint arXiv:211005457."},{"issue":"3","key":"10034_CR54","doi-asserted-by":"publisher","first-page":"4978","DOI":"10.1109\/LRA.2020.3004787","volume":"5","author":"S Song","year":"2020","unstructured":"Song, S., Zeng, A., Lee, J., & Funkhouser, T. (2020). Grasping in the wild: Learning 6dof closed-loop grasping from low-cost demonstrations. IEEE Robotics and Automation Letters, 5(3), 4978\u20134985.","journal-title":"IEEE Robotics and Automation Letters"},{"issue":"1","key":"10034_CR55","first-page":"9","volume":"3","author":"RS Sutton","year":"1988","unstructured":"Sutton, R. S. (1988). Learning to predict by the methods of temporal differences. Machine learning, 3(1), 9\u201344.","journal-title":"Machine learning"},{"key":"10034_CR56","unstructured":"Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction. MIT press."},{"issue":"13\u201314","key":"10034_CR57","doi-asserted-by":"crossref","first-page":"1455","DOI":"10.1177\/0278364917735594","volume":"36","author":"A ten Pas","year":"2017","unstructured":"ten Pas, A., Gualtieri, M., Saenko, K., & Platt, R. (2017). Grasp pose detection in point clouds. The International Journal of Robotics Research, 36(13\u201314), 1455\u20131473.","journal-title":"The International Journal of Robotics Research"},{"key":"10034_CR58","doi-asserted-by":"crossref","unstructured":"Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., & Abbeel, P. (2017). Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE\/RSJ international conference on intelligent robots and systems (IROS) (pp. 23\u201330). IEEE.","DOI":"10.1109\/IROS.2017.8202133"},{"key":"10034_CR59","doi-asserted-by":"crossref","unstructured":"Todorov, E., Erez, T., & Tassa, Y. (2012). Mujoco: A physics engine for model-based control. In 2012 IEEE\/RSJ international conference on intelligent robots and systems (pp. 5026\u20135033). IEEE.","DOI":"10.1109\/IROS.2012.6386109"},{"issue":"8","key":"10034_CR60","doi-asserted-by":"publisher","first-page":"916","DOI":"10.1080\/0951192X.2015.1130251","volume":"29","author":"P Tsarouchi","year":"2016","unstructured":"Tsarouchi, P., Makris, S., & Chryssolouris, G. (2016). Human-robot interaction review and challenges on task planning and programming. International Journal of Computer Integrated Manufacturing, 29(8), 916\u2013931.","journal-title":"International Journal of Computer Integrated Manufacturing"},{"key":"10034_CR61","first-page":"1","volume-title":"Ai and robotics innovation","author":"V Van Roy","year":"2020","unstructured":"Van Roy, V., Vertesy, D., & Damioli, G. (2020). Ai and robotics innovation (pp. 1\u201335). Human Resources and Population Economics: Handbook of Labor."},{"key":"10034_CR62","unstructured":"Viereck, U., Pas, A., Saenko, K., & Platt, R. (2017). Learning a visuomotor controller for real world robotic grasping using simulated depth images. In Conference on robot learning (pp. 291\u2013300). PMLR."},{"key":"10034_CR63","doi-asserted-by":"crossref","unstructured":"Wei, Q., Wang, L., Liu, Y., & Polycarpou, M. M. (2020). Optimal elevator group control via deep asynchronous actor-critic learning. IEEE Transactions on Neural Networks and Learning Systems.","DOI":"10.1109\/TNNLS.2020.2965208"},{"issue":"11","key":"10034_CR64","doi-asserted-by":"publisher","first-page":"5174","DOI":"10.1109\/TNNLS.2018.2805379","volume":"29","author":"Z Yang","year":"2018","unstructured":"Yang, Z., Merrick, K., Jin, L., & Abbass, H. A. (2018). Hierarchical deep reinforcement learning for continuous action control. IEEE Transactions on Neural Networks and Learning Systems, 29(11), 5174\u20135184.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"}],"container-title":["Autonomous Robots"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10514-022-10034-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10514-022-10034-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10514-022-10034-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,17]],"date-time":"2023-11-17T07:04:04Z","timestamp":1700204644000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10514-022-10034-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,9]]},"references-count":64,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,3]]}},"alternative-id":["10034"],"URL":"https:\/\/doi.org\/10.1007\/s10514-022-10034-z","relation":{},"ISSN":["0929-5593","1573-7527"],"issn-type":[{"value":"0929-5593","type":"print"},{"value":"1573-7527","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,9]]},"assertion":[{"value":"23 November 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 January 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 February 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}