{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T15:47:36Z","timestamp":1753890456006,"version":"3.41.2"},"reference-count":45,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,7,31]],"date-time":"2024-07-31T00:00:00Z","timestamp":1722384000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Comput. Sci."],"abstract":"<jats:p>Utilizing a robot in a new application requires the robot to be programmed at each time. To reduce such programmings efforts, we have been developing \u201cLearning-from-observation (LfO)\u201d that automatically generates robot programs by observing human demonstrations. So far, our previous research has been in the industrial domain. From now on, we want to expand the application field to the household-service domain. One of the main issues with introducing this LfO system into the domain is the cluttered environments, which makes it difficult to discern which movements of the human body parts and their relationships with environment objects are crucial for task execution when observing demonstrations. To overcome this issue, it is necessary for the system to have task common-sense shared with the human demonstrator to focus on the demonstrator's specific movements. Here, task common-sense is defined as the movements humans take almost unconsciously to streamline or optimize the execution of a series of tasks. In this paper, we extract and define three types of task common-sense (semi-conscious movements) that should be focused on when observing demonstrations of household tasks and propose representations to describe them. Specifically, the paper proposes to use labanotation to describe the whole-body movements with respect to the environment, contact-webs to describe the hand-finger movements with respect to the tool for grasping, and physical and semantic constraints to describe the movements of the hand with the tool with respect to the environment. Based on these representations, the paper formulates task models, machine-independent robot programs, that indicate what-to-do and where-to-do. In this design process, the necessary and sufficient set of task models to be prepared in the task-model library are determined on the following criteria: for grasping tasks, according to the classification of contact-webs along the purpose of the grasping, and for manipulation tasks, corresponding to possible transitions between states defined by either physical constraints and semantic constraints. The skill-agent library is also prepared to collect skill-agents corresponding to tasks. The skill-agents in the library are pre-trained using reinforcement learning with the reward functions designed based on the physical and semantic constraints to execute the task when specific parameters are provided. Third, the paper explains the task encoder to obtain task models and task decoder to execute the task models on the robot hardware. The task encoder understands what-to-do from the verbal input and retrieves the corresponding task model in the library. Next, based on the knowledge of each task, the system focuses on specific parts of the demonstration to collect where-to-do parameters for executing the task. The decoder constructs a sequence of skill-agents retrieving from the skill-agnet library corresponding and inserts those parameters obtained from the demonstration into these skill-agents, allowing the robot to perform task sequences with following the Labanotation postures. Finally, this paper presents how the system actually works through several example scenes.<\/jats:p>","DOI":"10.3389\/fcomp.2024.1235239","type":"journal-article","created":{"date-parts":[[2024,7,31]],"date-time":"2024-07-31T04:40:38Z","timestamp":1722400838000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Applying learning-from-observation to household service robots: three task common-sense formulations"],"prefix":"10.3389","volume":"6","author":[{"given":"Katsushi","family":"Ikeuchi","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Takamatsu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kazuhiro","family":"Sasabuchi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Naoki","family":"Wake","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Atsushi","family":"Kanehira","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2024,7,31]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"343","DOI":"10.1007\/s12369-012-0160-0","article-title":"Keyframe-based learning from demonstration: Method and evaluation","volume":"4","author":"Akgun","year":"2012","journal-title":"Int. J. Soc. Robot"},{"key":"B2","doi-asserted-by":"publisher","first-page":"289","DOI":"10.1142\/S0219843608001431","article-title":"Imitation learning of dual-arm manipulation tasks in humanoid robots","volume":"5","author":"Asfour","year":"2008","journal-title":"Int. J. Human. Robot"},{"key":"B3","doi-asserted-by":"crossref","first-page":"1371","DOI":"10.1007\/978-3-540-30301-5_60","article-title":"\u201cRobot programming by demonstration,\u201d","volume-title":"Springer Handbook of Robotics","author":"Billard","year":"2008"},{"key":"B4","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.143","article-title":"\u201cRealtime multi-person 2d pose estimation using part affinity fields,\u201d","author":"Cao","year":"2017","journal-title":"Proceedings of the IEEE conference on computer vision and pattern recognition"},{"key":"B5","doi-asserted-by":"publisher","first-page":"269","DOI":"10.1109\/70.34763","article-title":"On grasp choice, grasp models, and the design of hands for manufacturing tasks","volume":"5","author":"Cutkosky","year":"1989","journal-title":"IEEE Trans. Robot. Autom"},{"key":"B6","doi-asserted-by":"publisher","first-page":"188","DOI":"10.1109\/70.54734","article-title":"And\/or graph representation of assembly plans","volume":"6","author":"De Mello","year":"1990","journal-title":"IEEE Trans. Robot. Autom"},{"key":"B7","doi-asserted-by":"publisher","first-page":"295","DOI":"10.1007\/s13218-010-0060-0","article-title":"Advances in robot programming by demonstration","volume":"24","author":"Dillmann","year":"2010","journal-title":"KI-K\u00fcnstliche Intell"},{"key":"B8","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1109\/THMS.2015.2470657","article-title":"The grasp taxonomy of human grasp types","volume":"46","author":"Feix","year":"2015","journal-title":"IEEE Trans. Hum. Mach. Syst"},{"key":"B9","doi-asserted-by":"publisher","first-page":"381","DOI":"10.1145\/358669.358692","article-title":"Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography","volume":"24","author":"Fischler","year":"1981","journal-title":"Commun. ACM"},{"volume-title":"Labanotation: The System of Analyzing and Recording Movement","year":"1970","author":"Hutchinson-Guest","key":"B10"},{"key":"B11","doi-asserted-by":"publisher","first-page":"1415","DOI":"10.1007\/s11263-018-1123-1","article-title":"Describing upper-body motions based on Labanotation for learning-from-observation robots","volume":"126","author":"Ikeuchi","year":"2018","journal-title":"Int. J. Comput. Vis"},{"key":"B12","article-title":"\u201cAssembly plan from observation,\u201d","volume-title":"Annual Research Review","author":"Ikeuchi","year":"1991"},{"key":"B13","doi-asserted-by":"publisher","first-page":"368","DOI":"10.1109\/70.294211","article-title":"Toward an assembly plan from observation I. task recognition with polyhedral objects","volume":"10","author":"Ikeuchi","year":"1994","journal-title":"IEEE Trans. Robot. Autom"},{"key":"B14","doi-asserted-by":"publisher","first-page":"134","DOI":"10.1177\/02783649231212929","article-title":"Semantic constraints to represent common sense required in household actions for multimodal learning-from-observation robot","volume":"43","author":"Ikeuchi","year":"2024","journal-title":"Int. J. Rob. Res"},{"key":"B15","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1109\/70.554349","article-title":"Toward automatic robot instruction from perception-mapping human grasps to manipulator grasps","volume":"13","author":"Kang","year":"1997","journal-title":"IEEE Trans. Robot. Autom"},{"key":"B16","doi-asserted-by":"publisher","first-page":"799","DOI":"10.1109\/70.338535","article-title":"Learning by watching: extracting reusable task knowledge from visual observation of human performance","volume":"10","author":"Kuniyoshi","year":"1994","journal-title":"IEEE Trans. Robot. Autom"},{"key":"B17","doi-asserted-by":"publisher","first-page":"166","DOI":"10.3390\/robotics12060166","article-title":"Robot learning by demonstration with dynamic parameterization of the orientation: an application to agricultural activities","volume":"12","author":"Lauretti","year":"2023","journal-title":"Robotics"},{"key":"B18","doi-asserted-by":"publisher","first-page":"7661","DOI":"10.1109\/ACCESS.2023.3349093","article-title":"A new dmp scaling method for robot learning by demonstration and application to the agricultural domain","volume":"12","author":"Lauretti","year":"2024","journal-title":"IEEE Access"},{"key":"B19","doi-asserted-by":"publisher","first-page":"821","DOI":"10.1109\/PROC.1983.12681","article-title":"Robot programming","volume":"71","author":"Lozano-Perez","year":"1983","journal-title":"Proc. IEEE"},{"key":"B20","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1177\/027836498400300101","article-title":"Automatic synthesis of fine-motion strategies for robots","volume":"3","author":"Lozano-Perez","year":"1984","journal-title":"Int. J. Rob. Res"},{"key":"B21","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/9780262514620.001.0001","volume-title":"Vision: A Computational Investigation Into the Human Representation and Processing of Visual Information","author":"Marr","year":"2010"},{"key":"B22","doi-asserted-by":"publisher","DOI":"10.1037\/h0043158","article-title":"The magical number seven, plus or minus two: some limits on our capacity for processing information","author":"Miller","year":"1956","journal-title":"Psychol. Rev"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.21236\/ADA200313","author":"Minsky","year":"1988","journal-title":"Society of Mind"},{"key":"B24","doi-asserted-by":"publisher","first-page":"829","DOI":"10.1177\/0278364907079430","article-title":"Learning from observation paradigm: leg task models for enabling a biped humanoid robot to imitate human dances","volume":"26","author":"Nakaoka","year":"2007","journal-title":"Int. J. Rob. Res"},{"key":"B25","doi-asserted-by":"publisher","first-page":"771","DOI":"10.1109\/TRO.2014.2300212","article-title":"Toward a dancing robot with listening capability: keypose-based integration of lower-, middle-, and upper-body motions for varying music tempos","volume":"30","author":"Okamoto","year":"2014","journal-title":"IEEE Trans. Robot"},{"key":"B26","first-page":"1234","article-title":"Keypose and style analysis based on low-dimensional representation","volume":"50","author":"Perera","year":"2009","journal-title":"J. Inf. Proc. Soc. Japan"},{"volume-title":"La psychologie de l'intelligence","year":"2020","author":"Piaget","key":"B27"},{"volume-title":"Phantoms in the Brain.","year":"1998","author":"Ramachandran","key":"B28"},{"key":"B29","article-title":"Exploiting temporal information for 3D pose estimation","author":"Rayat Imtiaz Hossain","year":"2017","journal-title":"arXiv e-prints arXiv-1711"},{"key":"B30","doi-asserted-by":"publisher","DOI":"10.1145\/1283920.1283952","article-title":"\u201cTo dream the possible dream,\u201d","author":"Reddy","year":"2007","journal-title":"ACM Turing Award Lectures"},{"key":"B31","doi-asserted-by":"publisher","DOI":"10.1109\/Humanoids53995.2022.10000167","article-title":"\u201cTask-grasping from a demonstrated human strategy,\u201d","author":"Saito","year":"2022","journal-title":"Proceedings of the International Conference on Humanoids"},{"key":"B32","doi-asserted-by":"publisher","first-page":"54","DOI":"10.1002\/ecjc.20266","article-title":"Multiple model-based reinforcement learning for nonlinear control","volume":"89","author":"Samejima","year":"2006","journal-title":"Electr. Commun. Japan"},{"key":"B33","article-title":"Task-sequencing simulator: Integrated machine learning to execution simulation for robot manipulation","author":"Sasabuchi","year":"2023","journal-title":"arXiv preprint arXiv:2301.01382"},{"key":"B34","doi-asserted-by":"publisher","first-page":"413","DOI":"10.1109\/LRA.2020.3044029","article-title":"Task-oriented motion mapping on robots of various configuration using body role division","volume":"6","author":"Sasabuchi","year":"2021","journal-title":"IEEE Robot. Autom. Lett"},{"key":"B35","doi-asserted-by":"crossref","first-page":"1208","DOI":"10.1109\/IRDS.2002.1043898","article-title":"\u201cTask analysis based on observing hands and objects by vision,\u201d","volume-title":"IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Sato","year":"2002"},{"key":"B36","article-title":"\u201cLearning from demonstration,\u201d","author":"Schaal","year":"1996","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B37","doi-asserted-by":"publisher","first-page":"233","DOI":"10.1016\/S1364-6613(99)01327-3","article-title":"Is imitation learning the route to humanoid robots?","volume":"3","author":"Schaal","year":"1999","journal-title":"Trends Cogn. Sci"},{"key":"B38","doi-asserted-by":"publisher","first-page":"537","DOI":"10.1098\/rstb.2002.1258","article-title":"Computational approaches to motor learning by imitation","volume":"358","author":"Schaal","year":"2003","journal-title":"Philos. Trans. R. Soc. London"},{"key":"B39","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1109\/TRO.2005.855988","article-title":"Representation for knot-tying tasks","volume":"22","author":"Takamatsu","year":"2006","journal-title":"IEEE Trans. Robot"},{"key":"B40","doi-asserted-by":"publisher","first-page":"641","DOI":"10.1177\/0278364907080736","article-title":"Recognizing assembly tasks through human demonstration","volume":"26","author":"Takamatsu","year":"2007","journal-title":"Int. J. Rob. Res"},{"key":"B41","article-title":"Learning-from-observation system considering hardware-level reusability","author":"Takamatsu","year":"2022","journal-title":"arXiv preprint arXiv:2212.09242"},{"key":"B42","doi-asserted-by":"publisher","first-page":"535","DOI":"10.7210\/jrsj.18.535","article-title":"Generation of an assembly-task model analyzing human demonstration","volume":"18","author":"Tsuda","year":"2000","journal-title":"J. Robot. Soc. Japan"},{"key":"B43","article-title":"Interactive learning-from-observation through multimodal human demonstration","author":"Wake","year":"2022","journal-title":"arXiv preprint arXiv:2212.10787"},{"key":"B44","article-title":"\u201cGrasp-type recognition leveraging object affordance,\u201d","author":"Wake","year":"2020","journal-title":"HOBI Workshop, IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN)"},{"key":"B45","doi-asserted-by":"publisher","first-page":"418","DOI":"10.1115\/1.2802490","article-title":"Passive and active closures by constraining mechanisms","volume":"121","author":"Yoshikawa","year":"1999","journal-title":"J. Dyn. Sys. Meas. Control"}],"container-title":["Frontiers in Computer Science"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fcomp.2024.1235239\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,31]],"date-time":"2024-07-31T04:40:47Z","timestamp":1722400847000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fcomp.2024.1235239\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,31]]},"references-count":45,"alternative-id":["10.3389\/fcomp.2024.1235239"],"URL":"https:\/\/doi.org\/10.3389\/fcomp.2024.1235239","relation":{},"ISSN":["2624-9898"],"issn-type":[{"type":"electronic","value":"2624-9898"}],"subject":[],"published":{"date-parts":[[2024,7,31]]},"article-number":"1235239"}}