{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T16:31:39Z","timestamp":1753893099815,"version":"3.41.2"},"reference-count":34,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,6,23]],"date-time":"2025-06-23T00:00:00Z","timestamp":1750636800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Robot. AI"],"abstract":"<jats:p>Long-term use and highly reliable batteries are essential for wearable cyborgs including Hybrid Assistive Limb and wearable vital sensing devices. Consequently, there is ongoing research and development aimed at creating safer next-generation batteries. Researchers, leveraging advanced specialized knowledge and skills, bring products to completion through trial-and-error processes that involve modifying materials, shapes, work protocols, and procedures. When robots can undertake the tedious, repetitive, and attention-demanding tasks currently performed by researchers within facility environments, it will reduce the workload on researchers and ensure reproducibility. In this study, aiming to reduce the workload on researchers and ensure reproducibility in trial-and-error tasks, we proposed and developed a system that collects human motion data, recognizes action sequences, and transfers both physical information (including skeletal coordinates) and task information to a robot. This enables the robot to perform sequential tasks that are traditionally performed by humans. The proposed system employs a non-contact method to acquire three-dimensional skeletal information over time, allowing for quantitative analysis without interfering with sequential tasks. In addition, we developed an action sequence recognition model based on skeletal information and object detection results, which operated independent of background information. This model can adapt to changes in work processes and environments. By translating the human information including the physical and semantic information of a sequential task performed by humans into a robot, the robot can perform the same task. An experiment was conducted to verify this capability using the proposed system. The proposed action sequence recognition method demonstrated high accuracy in recognizing human-performed tasks with an average Edit score of 95.39 and an average F1@10 score of 0.951. In two out of the four trials, the robot adapted to changes in work processes without misrecognizing action sequences and seamlessly executed the sequential task performed by the human. In conclusion, we confirmed the feasibility of using the proposed system.<\/jats:p>","DOI":"10.3389\/frobt.2025.1462833","type":"journal-article","created":{"date-parts":[[2025,6,23]],"date-time":"2025-06-23T04:12:33Z","timestamp":1750651953000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Translating human information into robot tasks: action sequence recognition and robot control based on human motions"],"prefix":"10.3389","volume":"12","author":[{"given":"Taichi","family":"Obinata","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kazutomo","family":"Baba","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Akira","family":"Uehara","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hiroaki","family":"Kawamoto","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yoshiyuki","family":"Sankai","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2025,6,23]]},"reference":[{"key":"B1","first-page":"3782","volume-title":"Translating videos to commands for robotic manipulation with deep recurrent neural networks","author":"Anh","year":"2018"},{"volume-title":"YOLOv4: optimal speed and accuracy of object detection","year":"2020","author":"Bochkovskiy","key":"B2"},{"key":"B3","doi-asserted-by":"publisher","first-page":"237","DOI":"10.1038\/s41586-020-2442-2","article-title":"A mobile robotic chemist","volume":"583","author":"Burger","year":"2020","journal-title":"Nature"},{"year":"2024","key":"B4","article-title":"Annual report on the ageing society [summary] FY2022"},{"key":"B5","doi-asserted-by":"publisher","first-page":"17615","DOI":"10.1039\/c1cp21910c","article-title":"Graphene and carbon nanotube composite electrodes for supercapacitors with ultra-high energy density","volume":"13","author":"Cheng","year":"2011","journal-title":"Phys. Chem. Chem. Phys."},{"key":"B6","doi-asserted-by":"publisher","first-page":"1011","DOI":"10.1109\/tpami.2023.3327284","article-title":"Temporal action segmentation: an analysis of modern techniques","volume":"46","author":"Ding","year":"2024","journal-title":"IEEE Trans. pattern analysis Mach. Intell."},{"key":"B7","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2401.02117","article-title":"Mobile aloha: learning bimanual mobile manipulation with low-cost whole-body teleoperation","author":"Fu","year":"2024","journal-title":"arXiv Prepr. arXiv:2401.02117"},{"volume-title":"Springer handbook of robotics","year":"2016","author":"Khatib","key":"B8"},{"volume-title":"Semi-supervised classification with graph convolutional networks","year":"2017","author":"Kipf","key":"B9"},{"journal-title":"DOF Specifications","article-title":"Kinova Ultra lightweight robotic arm 6","year":"2018","key":"B34"},{"key":"B10","first-page":"780","article-title":"The language of actions: recovering the syntax and semantics of goal-directed human activities","author":"Kuehne","year":"2014"},{"key":"B11","first-page":"156","article-title":"Temporal convolutional networks for action segmentation and detection","author":"Lea","year":"2017"},{"key":"B12","first-page":"1642","article-title":"Learning convolutional action primitives for fine-grained action recognition","author":"Lea","year":"2016"},{"key":"B13","first-page":"287","article-title":"Delving into egocentric actions","author":"Li","year":"2015"},{"key":"B14","first-page":"1113","article-title":"Learning latent plans from play","author":"Lynch","year":"2020"},{"article-title":"Open X-embodiment: robotic learning datasets and RT-X models: open X-embodiment collaboration","year":"2023","author":"Maddukuri","key":"B15"},{"key":"B16","doi-asserted-by":"publisher","first-page":"425","DOI":"10.1093\/geront\/gnr067","article-title":"Japan: super-aging society preparing for the future","volume":"51","author":"Muramatsu","year":"2011","journal-title":"Gerontologist"},{"key":"B17","doi-asserted-by":"publisher","first-page":"926","DOI":"10.20729\/00225500","article-title":"[Study on Action Recognition Including Fine-detailed Hand Motion toward Monitoring System] Tearai sagyou no monitoring ni muketa zenshin to temoto no dousaninshiki syuhou no kenkyuu kaihatu","author":"Obinata","year":"2023","journal-title":"IPSJ J. Inf. Process. Soc. Jpn."},{"key":"B18","first-page":"363","volume-title":"Development of real-time assembly work monitoring system based on 3D skeletal model of arms and fingers, oct 11, 2020","author":"Obinata","year":"2020"},{"key":"B19","first-page":"8748","article-title":"Learning transferable visual models from natural language supervision","author":"Radford","year":"2021"},{"key":"B20","first-page":"21096","article-title":"Assembly101: a large-scale multi-view video dataset for understanding procedural activities","author":"Sener","year":"2022"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.1063\/5.0020370","article-title":"Autonomous materials synthesis by machine learning and robotics","volume":"8","author":"Shimizu","year":"2020","journal-title":"Apl. Mater."},{"key":"B22","first-page":"729","volume-title":"Combining embedded accelerometers with computer vision for recognizing food preparation activities","author":"Stein","year":"2013"},{"key":"B23","doi-asserted-by":"publisher","first-page":"10567","DOI":"10.1109\/lra.2024.3477090","article-title":"GPT-4V(ision) for robotics: multimodal task planning from human demonstration","volume":"9","author":"Wake","year":"2024","journal-title":"IEEE robotics automation Lett."},{"key":"B24","doi-asserted-by":"publisher","first-page":"2171","DOI":"10.1109\/tpami.2023.3330794","article-title":"Temporal action localization in the deep learning era: a survey","volume":"46","author":"Wang","year":"2024","journal-title":"IEEE Trans. pattern analysis Mach. Intell."},{"key":"B25","first-page":"201","article-title":"Mimicplay: long-horizon imitation learning by watching human play","author":"Wang","year":"2023"},{"key":"B26","doi-asserted-by":"publisher","first-page":"1734","DOI":"10.1109\/tnnls.2023.3329525","article-title":"Effective surrogate gradient learning with high-order information bottleneck for spike-based machine intelligence","volume":"36","author":"Yang","year":"","journal-title":"IEEE transaction neural Netw. Learn. Syst."},{"key":"B27","doi-asserted-by":"publisher","first-page":"7852","DOI":"10.1109\/tsmc.2023.3300318","article-title":"SNIB: improving spike-based machine learning using nonlinear information bottleneck","volume":"53","author":"Yang","year":"","journal-title":"IEEE Trans. Syst. man, Cybern. Syst."},{"key":"B28","doi-asserted-by":"publisher","first-page":"128535","DOI":"10.1016\/j.neucom.2024.128535","article-title":"Maximum entropy intrinsic learning for spiking networks towards embodied neuromorphic vision","volume":"610","author":"Yang","year":"2024","journal-title":"Neurocomputing"},{"key":"B29","doi-asserted-by":"publisher","first-page":"4398","DOI":"10.1109\/tnnls.2021.3057070","article-title":"CerebelluMorphic: large-scale neuromorphic model and architecture for supervised motor learning","volume":"33","author":"Yang","year":"2022","journal-title":"IEEE transaction neural Netw. Learn. Syst."},{"key":"B30","doi-asserted-by":"publisher","first-page":"4404","DOI":"10.1109\/tsmc.2023.3248324","article-title":"Watch and act: learning robotic manipulation from visual demonstration","volume":"53","author":"Yang","year":"2023","journal-title":"IEEE Trans. Syst. man, Cybern. Syst."},{"key":"B31","doi-asserted-by":"crossref","DOI":"10.1609\/aaai.v32i1.12328","article-title":"Spatial temporal graph convolutional networks for skeleton-based action recognition","author":"Yan","year":"2018"},{"key":"B32","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2006.10214","article-title":"Mediapipe hands: on-device real-time hand tracking","author":"Zhang","year":"2020","journal-title":"arXiv Prepr. arXiv:2006.10214"},{"key":"B33","doi-asserted-by":"publisher","first-page":"1330","DOI":"10.1109\/34.888718","article-title":"A flexible new technique for camera calibration","volume":"22","author":"Zhang","year":"2000","journal-title":"IEEE Trans. pattern analysis Mach. Intell."}],"container-title":["Frontiers in Robotics and AI"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frobt.2025.1462833\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,23]],"date-time":"2025-06-23T04:12:36Z","timestamp":1750651956000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frobt.2025.1462833\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,23]]},"references-count":34,"alternative-id":["10.3389\/frobt.2025.1462833"],"URL":"https:\/\/doi.org\/10.3389\/frobt.2025.1462833","relation":{},"ISSN":["2296-9144"],"issn-type":[{"type":"electronic","value":"2296-9144"}],"subject":[],"published":{"date-parts":[[2025,6,23]]},"article-number":"1462833"}}