{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T20:19:22Z","timestamp":1777407562757,"version":"3.51.4"},"reference-count":49,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,11,14]],"date-time":"2025-11-14T00:00:00Z","timestamp":1763078400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Robot. AI"],"abstract":"<jats:p>Recent advances in artificial intelligence (AI) have attracted significant attention due to AI\u2019s ability to solve complex problems and the rapid development of learning algorithms and computational power. Among the many AI techniques, transformers stand out for their flexible architectures and high computational capacity. Unlike traditional neural networks, transformers use mechanisms such as self-attention with positional encoding, which enable them to effectively capture long-range dependencies in sequential and spatial data. This paper presents a comparison of various deep Q-learning algorithms and proposes two original techniques that use self-attention into deep Q-learning. The first technique is structured self-attention with deep Q-learning, and the second uses multi-head attention with deep Q-learning. These methods are compared with different types of deep Q-learning and other temporal techniques in uncertain tasks, such as throwing objects to unknown targets. The performance of these algorithms is evaluated in a simplified environment, where the task involves throwing a ball using a robotic arm manipulator. This setup provides a controlled scenario to analyze the algorithms\u2019 efficiency and effectiveness in solving dynamic control problems. Additional constraints are introduced to evaluate performance under more complex conditions, such as a joint lock or the presence of obstacles like a wall near the robot or the target. The output of the algorithm includes the correct joint configurations and trajectories for throwing to unknown target positions. The use of multi-head attention has enhanced the robot\u2019s ability to prioritize and interact with critical environmental features. The paper also includes a comparison of temporal difference algorithms to address constraints on the robot\u2019s joints. These algorithms are capable of finding solutions within the limitations of existing hardware, enabling robots to interact intelligently and autonomously with their environment.<\/jats:p>","DOI":"10.3389\/frobt.2025.1567211","type":"journal-article","created":{"date-parts":[[2025,11,14]],"date-time":"2025-11-14T05:10:41Z","timestamp":1763097041000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Comparative analysis of deep Q-learning algorithms for object throwing using a robot manipulator"],"prefix":"10.3389","volume":"12","author":[{"given":"Mohammad","family":"Al Homsi","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maja","family":"Trumi\u0107","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adriano","family":"Fagiolini","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Giansalvo","family":"Cirrincione","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2025,11,14]]},"reference":[{"key":"B1","article-title":"Analytics","volume-title":"Reinforcement learning: exploration vs exploitation tradeoff","year":"2023"},{"key":"B2","doi-asserted-by":"crossref","first-page":"243","DOI":"10.1007\/978-3-319-32552-1_21","article-title":"Actuators for soft robotics","volume-title":"Handbook of robotics","author":"Albu-Sch\u00e4ffer","year":"2016"},{"key":"B3","first-page":"74","article-title":"Accurate object throwing by an industrial robot manipulator","volume-title":"Proceedings of the Australasian Conference on robotics and automation 2010","author":"August","year":"2010"},{"key":"B4","article-title":"Layer normalization","author":"Ba","year":"2016"},{"key":"B5","doi-asserted-by":"publisher","first-page":"2","DOI":"10.1109\/MRA.2023.3310865","article-title":"Softoss: learning to throw objects with a soft robot","author":"Bianchi","year":"2023","journal-title":"IEEE Robotics and Automation Mag."},{"key":"B6","doi-asserted-by":"publisher","first-page":"127","DOI":"10.1109\/MRA.2022.3177355","article-title":"Dual-arm control for coordinated fast grabbing and tossing of an object: Proposing a new approach","volume":"29","author":"Bombile","year":"2022","journal-title":"IEEE Robotics Automation Mag."},{"key":"B7","doi-asserted-by":"publisher","first-page":"104481","DOI":"10.1016\/j.robot.2023.104481","article-title":"Bimanual dynamic grabbing and tossing of objects onto a moving target","volume":"167","author":"Bombile","year":"2023","journal-title":"Robotics Aut. Syst."},{"key":"B8","article-title":"How growing e-commerce demand is driving growth in mobile robotics","author":"Britt","year":"2020","journal-title":"Robot. Bus. Rev."},{"key":"B9","doi-asserted-by":"publisher","first-page":"292","DOI":"10.1109\/iros.1995.526175","article-title":"Toward a dynamical pick and place","volume":"2","author":"Burridge","year":"1995","journal-title":"Proc. 1995 IEEE\/RSJ Int. Conf. Intelligent Robots Syst. Hum. Robot Interact. Coop. Robots"},{"key":"B10","doi-asserted-by":"publisher","first-page":"425","DOI":"10.1016\/j.measurement.2019.06.039","article-title":"Failure detection in robotic arms using statistical modeling, machine learning and hybrid gradient boosting","volume":"146","author":"Costa","year":"2019","journal-title":"Measurement"},{"key":"B11","first-page":"6718","article-title":"Catch the ball: accurate high-speed motions for mobile manipulators via inverse dynamics learning","volume-title":"2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems, IROS","author":"Dong","year":"2020"},{"key":"B12","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1109\/ROBIO.2006.340302","article-title":"Throwing objects\u2013a bio-inspired approach for the transportation of parts","volume-title":"2006 IEEE International Conference on Robotics and Biomimetics","author":"Frank","year":"2006"},{"key":"B13","volume-title":"Real time Probabilistic models for robot trajectories","author":"Gonzalez","year":"2020"},{"key":"B14","unstructured":"Robotic systems - inverse kinematics\n          \n          \n            \n              Group\n              M. R.\n            \n          \n          \n          2024"},{"key":"B15","doi-asserted-by":"publisher","first-page":"4707","DOI":"10.1109\/tmech.2022.3164247","article-title":"Time-optimal pick-and-throw s-curve trajectories for fast parallel robots","volume":"27","author":"Hassan","year":"2022","journal-title":"IEEE\/ASME Trans. Mechatronics"},{"key":"B16","first-page":"06527","article-title":"Deep recurrent q-learning for partially observable mdps","author":"Hausknecht","year":"2015","journal-title":"Corr."},{"key":"B17","article-title":"AI-Based approach for throwing and Grasping objects from unknown positions by soft robot Upper Body manipulators","volume-title":"Proceedings of the IEEE\/RSJ International Workshop on Intelligent Robots and Systems (IROS)","author":"Homsi","year":"2023"},{"key":"B18","doi-asserted-by":"crossref","DOI":"10.1109\/ECICE47484.2019.8942793","article-title":"Machine learning approach for robot diagnostic system","volume-title":"IEEE Eurasia Conference on IOT","author":"Hu","year":"2019"},{"key":"B19","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48891.2023.10160215","article-title":"Throwing objects into a moving basket while avoiding obstacles","author":"Kasaei","year":"2024","journal-title":"Unspecified"},{"key":"B20","article-title":"Attention layer","year":"2023"},{"key":"B21","first-page":"587","article-title":"Throwing motion generation of a biped human model","volume-title":"2008 2nd IEEE RAS and EMBS International Conference on Biomedical Robotics and Biomechatronics","author":"Kim","year":"2008"},{"key":"B22","volume-title":"Deep reinforcement learning hands-on","author":"Lapan","year":"2018"},{"key":"B23","article-title":"Inverse kinematics tutorial","author":"Learning","year":"2025"},{"key":"B24","doi-asserted-by":"crossref","DOI":"10.1007\/978-981-19-7784-8","volume-title":"Reinforcement learning for sequential decision and optimal control","author":"Li","year":"2023"},{"key":"B25","doi-asserted-by":"publisher","first-page":"333","DOI":"10.3390\/s20020333","article-title":"Ball tracking and trajectory prediction for table-tennis robots","volume":"20","author":"Lin","year":"2020","journal-title":"Sensors"},{"key":"B26","first-page":"1625","article-title":"A solution to adaptive mobile manipulator throwing","volume-title":"2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems, IROS","author":"Liu","year":"2022"},{"key":"B27","doi-asserted-by":"publisher","first-page":"64","DOI":"10.1177\/027836499901800105","article-title":"Dynamic nonprehensile manipulation: Controllability, planning, and experiments","volume":"18","author":"Lynch","year":"1999","journal-title":"Int. J. Robotics Res."},{"key":"B28","doi-asserted-by":"publisher","first-page":"152","DOI":"10.1109\/iros.1993.583093","article-title":"Dynamic manipulation","volume":"1","author":"Mason","year":"1993","journal-title":"Proc. 1993 IEEE\/RSJ Int. Conf. Intelligent Robots Syst."},{"key":"B29","article-title":"As e-commerce booms, robots pick up human slack","author":"Mims","year":"2020","journal-title":"Wall Str. J."},{"key":"B30","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1312.5602","article-title":"Playing atari with deep reinforcement learning","author":"Mnih","year":"2013","journal-title":"arXiv Prepr. arXiv:1312"},{"key":"B31","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"B32","article-title":"Massively parallel methods for deep reinforcement learning","author":"Nair","year":"2015"},{"key":"B33","doi-asserted-by":"crossref","first-page":"1149","DOI":"10.1109\/CASE48305.2020.9216746","article-title":"Robotic pick-and-toss facilitates urban waste sorting","volume-title":"2020 IEEE 16th International Conference on automation Science and engineering (CASE)","author":"Raptopoulos","year":"2020"},{"key":"B34","first-page":"3932","article-title":"A coordinate-free framework for robotic pizza tossing and catching","volume-title":"2016 IEEE International Conference on robotics and automation, ICRA","author":"Satici","year":"2016"},{"key":"B35","first-page":"207","volume-title":"A coordinate-free framework for robotic Pizza Tossing and catching","author":"Satici","year":"2022"},{"key":"B36","doi-asserted-by":"publisher","first-page":"1502","DOI":"10.1109\/tro.2018.2868857","article-title":"Robust ballistic catching: a hybrid system stabilization problem","volume":"34","author":"Schill","year":"2018","journal-title":"IEEE Trans. Robotics"},{"key":"B37","article-title":"Kinematics","volume-title":"Robotics: Modelling, planning, and control","author":"Siciliano","year":"2008"},{"key":"B38","doi-asserted-by":"publisher","DOI":"10.35940\/ijrte.B1004.0782S319","article-title":"Intelligence decision making of fault detection and fault tolerance method for industrial robotic manipulators","author":"Sivasamy","year":"2019","journal-title":"IJRTE"},{"key":"B39","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1512.01693","article-title":"Deep attention recurrent q-network","author":"Sorokin","year":"2015","journal-title":"ArXiv"},{"key":"B40","volume-title":"Reinforcement learning: an introduction","author":"Sutton","year":"2018"},{"key":"B41","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1509.06461","article-title":"Deep reinforcement learning with double q-learning","author":"van Hasselt","year":"2015","journal-title":"arXiv Prepr. arXiv:1509.06461"},{"key":"B42","first-page":"5","article-title":"Deep reinforcement learning with double q-learning","author":"Van Hasselt","year":"2016"},{"key":"B43","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1706.03762","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"B44","volume-title":"Learning from delayed rewards","author":"Watkins","year":"1989"},{"key":"B45","article-title":"Controlling the phantomx pincher robot arm","author":"Wiki","year":"2024"},{"key":"B46","article-title":"Exploration-exploitation dilemma","year":"2023"},{"key":"B47","doi-asserted-by":"publisher","first-page":"1307","DOI":"10.1109\/tro.2020.2988642","article-title":"Tossingbot: learning to throw arbitrary objects with residual physics","volume":"36","author":"Zeng","year":"2020","journal-title":"IEEE Trans. Robotics"},{"key":"B48","doi-asserted-by":"crossref","first-page":"2551","DOI":"10.1109\/ICRA.2012.6225319","article-title":"Sampling-based motion planning with dynamic intermediate state objectives: application to throwing","volume-title":"2012 IEEE International Conference on Robotics and automation","author":"Zhang","year":"2012"},{"key":"B49","doi-asserted-by":"publisher","first-page":"08247v2","DOI":"10.1109\/ACCESS.2020.2972859","article-title":"Machine learning and deep learning algorithms for bearing fault diagnostics \u2013 a comprehensive review","author":"Zhang","year":"2019","journal-title":"Cornell Univ. arXiv:1901"}],"container-title":["Frontiers in Robotics and AI"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frobt.2025.1567211\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,14]],"date-time":"2025-11-14T05:10:48Z","timestamp":1763097048000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frobt.2025.1567211\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,14]]},"references-count":49,"alternative-id":["10.3389\/frobt.2025.1567211"],"URL":"https:\/\/doi.org\/10.3389\/frobt.2025.1567211","relation":{},"ISSN":["2296-9144"],"issn-type":[{"value":"2296-9144","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,14]]},"article-number":"1567211"}}