{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T12:57:17Z","timestamp":1781614637569,"version":"3.54.5"},"reference-count":70,"publisher":"Springer Science and Business Media LLC","issue":"12","license":[{"start":{"date-parts":[[2025,11,5]],"date-time":"2025-11-05T00:00:00Z","timestamp":1762300800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,11,5]],"date-time":"2025-11-05T00:00:00Z","timestamp":1762300800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"TU Wien"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J of Soc Robotics"],"published-print":{"date-parts":[[2025,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The increasing presence of robots in human workspaces underscores the need for intelligent systems that can understand human behaviors and act accordingly for a natural human-robot interaction (HRI). In this work, we propose a method to generate a robot\u2019s behavior for social HRI by integrating both human and robot intentions into the robot\u2019s decision-making process. Our system learns appropriate robot behaviors in social scenarios by observing human-human interactions (HHI). Using a transformer-based model, we first capture the dynamics of each individual and then iteratively adapt both human and robot behavior to achieve a successful interaction. By connecting our model with a human-to-robot motion retargeting framework, our system learns how a robot should behave solely from observing human data. To address the disparity between HHI and HRI, we employ several loss functions that encourage our robot to reproduce the social dynamics observed in humans. As a result, our approach outperforms the state-of-the-art in dyadic human motion forecasting prediction in the largest dataset available and obtains high-quality robot behaviors in human-robot interaction scenarios. Finally, we conduct a thorough evaluation of our performance for HHI, and HRI, and implement and test the system in the real-world TIAGo++ robot.<\/jats:p>","DOI":"10.1007\/s12369-025-01333-3","type":"journal-article","created":{"date-parts":[[2025,11,5]],"date-time":"2025-11-05T13:57:39Z","timestamp":1762351059000},"page":"3211-3230","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Robot Behavior Generation for Social Human-Robot Interaction"],"prefix":"10.1007","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5147-7704","authenticated-orcid":false,"given":"Esteve","family":"Valls Mascaro","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dongheui","family":"Lee","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,11,5]]},"reference":[{"key":"1333_CR1","doi-asserted-by":"publisher","first-page":"140","DOI":"10.1523\/JNEUROSCI.1431-18.2018","volume":"39","author":"MA Thornton","year":"2018","unstructured":"Thornton MA, Weaverdyck ME, Tamir DI (2018) The social brain automatically predicts others\u2019 future mental states. J Neurosci 39:140\u2013148","journal-title":"J. Neurosci"},{"key":"1333_CR2","doi-asserted-by":"crossref","unstructured":"Fragkiadaki K, Levine S, Felsen P, Malik J (2015) Recurrent network models for human dynamics. In IEEE International Conference on Computer Vision, pp 4346\u20134354","DOI":"10.1109\/ICCV.2015.494"},{"key":"1333_CR3","doi-asserted-by":"crossref","unstructured":"Jain A, Zamir AR, Savarese S, Saxena A (2016) Structural-rnn: deep learning on spatio-temporal graphs. In Conference on Computer Vision and Pattern Recognition (CVPR), pp 5308\u20135317","DOI":"10.1109\/CVPR.2016.573"},{"key":"1333_CR4","doi-asserted-by":"crossref","unstructured":"Dang L, Nie Y, Long C, Zhang Q, Li G (2021) Msr-gcn: multi-scale residual graph convolution networks for human motion prediction. In IEEE\/CVF International Conference on Computer Vision (ICCV), pp 11467\u201311476","DOI":"10.1109\/ICCV48922.2021.01127"},{"key":"1333_CR5","doi-asserted-by":"crossref","unstructured":"Mao W, Liu M, Salzmann M, Li H (2019) Learning trajectory dependencies for human motion prediction. In International Conference on Computer Vision (ICCV), pp 9489\u20139497","DOI":"10.1109\/ICCV.2019.00958"},{"key":"1333_CR6","doi-asserted-by":"crossref","unstructured":"Mao W, Liu M, Salzmann M (2020) History repeats itself: human motion prediction via motion attention. In European Conference on Computer Vision (ECCV), pp 474\u2013489","DOI":"10.1007\/978-3-030-58568-6_28"},{"key":"1333_CR7","doi-asserted-by":"crossref","unstructured":"Valls Mascaro E, Ma S, Ahn H, Lee D (2022) Robust human motion forcasting using transformer-based model. In 2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","DOI":"10.1109\/IROS47612.2022.9981877"},{"key":"1333_CR8","unstructured":"Mascaro EV, Ahn H, Lee D (2023). A unified masked autoencoder with patchified skeletons for motion synthesis. arXiv preprint arXiv:2308.07301"},{"key":"1333_CR9","doi-asserted-by":"crossref","unstructured":"Guo W, Du Y, Shen X, Lepetit V, Alameda-Pineda X, Moreno-Noguer F (2023) Back to mlp: a simple baseline for human motion prediction. In IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), pp 4809\u20134819","DOI":"10.1109\/WACV56688.2023.00479"},{"key":"1333_CR10","doi-asserted-by":"crossref","unstructured":"Peng X, Mao S, Wu Z (2023) Trajectory-aware body interaction transformer for multi-person pose forecasting. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 17121\u201317130","DOI":"10.1109\/CVPR52729.2023.01642"},{"key":"1333_CR11","unstructured":"Rahman MRU, Scofano L, De Matteis E, Flaborea A, Sampieri A, Galasso F (2023) Best practices for 2-body pose forecasting. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"1333_CR12","unstructured":"Peng X, Zhou X, Luo Y, Wen H, Wu Z (2023). The MI-Motion dataset and benchmark for 3D multi-person motion prediction. arXiv preprint arXiv:2306.13566, 2023"},{"key":"1333_CR13","unstructured":"Marcard T, Henschel R, Black M, Rosenhahn B, Pons-Moll G (2018) Recovering accurate 3d human pose in the wild using imus and a moving camera. In European Conference on Computer Vision (ECCV)"},{"key":"1333_CR14","unstructured":"Wang J, Xu H, Narasimhan M, Wang X (2021) Multi-person 3d motion prediction with multi-range transformers. In Proceedings of the 35th International Conference on Neural Information Processing Systems. NIP\u2019S 21. Curran Associates Inc, Red Hook, NY, USA"},{"key":"1333_CR15","doi-asserted-by":"crossref","unstructured":"Guo W, Bie X, Alameda-Pineda X, Moreno\u2013Noguer F (2022) Multi-person extreme motion prediction. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR52688.2022.01271"},{"issue":"9","key":"1333_CR16","doi-asserted-by":"publisher","first-page":"3463","DOI":"10.1007\/s11263-024-02042-6","volume":"132","author":"H Liang","year":"2024","unstructured":"Liang H, Zhang W, Li W, Yu J, Xu L (2024) Intergen: diffusion-based multi-human motion generation under complex interactions. Int J Comput Vision 132(9):3463\u20133483","journal-title":"Int J Comput Vision"},{"key":"1333_CR17","doi-asserted-by":"crossref","unstructured":"Kopp T, Baumgartner M, Kinkel S (2021) Success factors for introducing industrial human-robot interaction in practice: an empirically driven framework. Int J Adv Manuf Technol 112","DOI":"10.1007\/s00170-020-06398-0"},{"issue":"2","key":"1333_CR18","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3359616","volume":"9","author":"M Chen","year":"2020","unstructured":"Chen M, Nikolaidis S, Soh H, Hsu D, Srinivasa S (2020) Trust-aware decision making for human-robot collaboration: model learning and planning. ACM Trans Hum Rob Interact (THRI) 9(2):1\u201323","journal-title":"ACM Trans Hum Rob Interact (THRI)"},{"key":"1333_CR19","doi-asserted-by":"crossref","unstructured":"(2020) Recurrent neural network for motion trajectory prediction in human-robot collaborative assembly. CIRP Ann 69(1):9\u201312","DOI":"10.1016\/j.cirp.2020.04.077"},{"key":"1333_CR20","unstructured":"Mascaro EV, Sliwowski D, Lee D (2023) HOI4ABOT: human-object interaction anticipation for human intention reading assistive roBots. In 7th Annual Conference on Robot Learning"},{"key":"1333_CR21","doi-asserted-by":"crossref","unstructured":"Sampieri A, Melendugno GMD, Avogaro A, Cunico F, Setti F, Skenderi G, Cristani M, Galasso F (2022) Pose forecasting in industrial human-robot collaboration. In European Conference on Computer Vision, Springer pp 51\u201369","DOI":"10.1007\/978-3-031-19839-7_4"},{"key":"1333_CR22","doi-asserted-by":"crossref","unstructured":"Kedia K, Bhardwaj A, Dan P, Choudhury S (2024) Interact: transformer models for human intent prediction conditioned on robot actions. In IEEE International Conference on Robotics and Automation","DOI":"10.1109\/ICRA57147.2024.10610681"},{"key":"1333_CR23","doi-asserted-by":"crossref","unstructured":"Valls Mascaro E, Yan Y, Lee D (2024) Robot interaction behavior generation based on social motion forecasting for human-robot interaction. In 2024 IEEE International Conference on Robotics and Automation (ICRA","DOI":"10.1109\/ICRA57147.2024.10610682"},{"key":"1333_CR24","doi-asserted-by":"crossref","unstructured":"Yan Y, Mascaro EV, Lee D (2023). Unsupervised human-to-robot motion retargeting via expressive latent space. arXiv preprint arXiv:2309.05310, 2023","DOI":"10.1109\/Humanoids57100.2023.10375150"},{"issue":"1","key":"1333_CR25","doi-asserted-by":"publisher","first-page":"010804","DOI":"10.1115\/1.4039145","volume":"70","author":"DP Losey","year":"2018","unstructured":"Losey DP, McDonald CG, Battaglia E, O\u2019Malley MK (2018) A review of intent detection, arbitration, and communication aspects of shared control for physical human\u2013robot interaction. Appl Mech Rev 70(1):010804","journal-title":"Appl Mech Rev"},{"key":"1333_CR26","doi-asserted-by":"crossref","unstructured":"Fang J, Wang F, Xue J, Chua T-S (2024) Behavioral intention prediction in driving scenes: a survey. In IEEE Transactions on Intelligent Transportation Systems, pp 1\u201322","DOI":"10.1109\/TITS.2024.3374342"},{"key":"1333_CR27","doi-asserted-by":"crossref","unstructured":"Huang C-M, Mutlu B (2016) Anticipatory robot control for efficient human-robot collaboration. In 2016 11th ACM\/IEEE International Conference on Human-Robot Interaction (HRI), pp 83\u201390","DOI":"10.1109\/HRI.2016.7451737"},{"key":"1333_CR28","doi-asserted-by":"publisher","first-page":"103741","DOI":"10.1016\/j.cviu.2023.103741","volume":"233","author":"Z Ni","year":"2023","unstructured":"Ni Z, Valls Mascar\u00f3 E, Ahn H, Lee D (2023) Human\u2013object interaction prediction in videos through gaze following. Comput Vision Image Underst 233:103741","journal-title":"Comput Vision Image Underst"},{"key":"1333_CR29","doi-asserted-by":"crossref","unstructured":"Schydlo P, Rakovic M, Jamone L, Santos-Victor J (2018) Anticipation in human-robot cooperation: a recurrent neural network approach for multiple action sequences prediction. In 2018 IEEE International Conference on Robotics and Automation (ICRA), IEEE pp 5909\u20135914","DOI":"10.1109\/ICRA.2018.8460924"},{"key":"1333_CR30","doi-asserted-by":"crossref","unstructured":"Zatsarynna O, Gall J (2023) Action anticipation with goal consistency. In 2023 IEEE International Conference on Image Processing (ICIP), pp 1630\u20131634","DOI":"10.1109\/ICIP49359.2023.10222914"},{"key":"1333_CR31","doi-asserted-by":"crossref","unstructured":"Mascaro EV, Ahn H, Lee D (2023) Intention-conditioned long-term human egocentric action anticipation. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), pp 6048\u20136057","DOI":"10.1109\/WACV56688.2023.00599"},{"key":"1333_CR32","doi-asserted-by":"publisher","first-page":"128","DOI":"10.1007\/978-3-540-48113-3_12","volume-title":"Robotics research","author":"Y Nakamura","year":"2007","unstructured":"Nakamura Y, Takano W, Yamane K (2007) Mimetic communication theory for humanoid robots interacting with humans. In: Robotics research. Springer, Berlin, Heidelberg, pp 128\u2013139"},{"issue":"13","key":"1333_CR33","doi-asserted-by":"publisher","first-page":"1684","DOI":"10.1177\/0278364910364164","volume":"29","author":"D Lee","year":"2010","unstructured":"Lee D, Ott C, Nakamura Y (2010) Mimetic communication model with compliant physical contact in human\u2014humanoid interaction. Int J Robot Res 29(13):1684\u20131704","journal-title":"Int J Robot Res"},{"key":"1333_CR34","doi-asserted-by":"crossref","unstructured":"Medina Hern\u00e1ndez J, Lawitzky M, M\u00f6rtl A, Lee D, Hirche S (2011) An experience-driven robotic assistant acquiring human knowledge to improve haptic cooperation 2416\u20132422","DOI":"10.1109\/IROS.2011.6095026"},{"key":"1333_CR35","doi-asserted-by":"crossref","unstructured":"Yang L, Li Y, Huang D (2018) Motion synchronization in human-robot co-transport without force sensing. In 2018 37th Chinese Control Conference (CCC), pp 5369\u20135374","DOI":"10.23919\/ChiCC.2018.8484161"},{"key":"1333_CR36","doi-asserted-by":"crossref","unstructured":"Wang Z, Peer A, Buss M (2009) An hmm approach to realistic haptic human-robot interaction. In World Haptics 2009 - Third Joint EuroHaptics Conference and Symposium on Haptic Interfaces for Virtual Environment and Teleoperator Systems, pp 374\u2013379","DOI":"10.1109\/WHC.2009.4810835"},{"key":"1333_CR37","doi-asserted-by":"crossref","unstructured":"Alahi A, Ramanathan V, Fei-Fei L (2014) Socially-aware large-scale crowd forecasting. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2014.283"},{"key":"1333_CR38","doi-asserted-by":"crossref","unstructured":"Adeli V, Adeli E, Reid I, Niebles JC, Rezatofighi H (2020) Socially and contextually aware human motion and pose forecasting. IEEE Robot Autom Lett 5(4)","DOI":"10.1109\/LRA.2020.3010742"},{"issue":"11","key":"1333_CR39","doi-asserted-by":"publisher","first-page":"999","DOI":"10.1016\/j.visres.2010.02.008","volume":"50","author":"D Baldauf","year":"2010","unstructured":"Baldauf D, Deubel H (2010) Attentional landscapes in reaching and grasping. Vision Res 50(11):999\u20131013","journal-title":"Vision Res"},{"key":"1333_CR40","doi-asserted-by":"crossref","unstructured":"Belardinelli A (2023) Gaze-based intention estimation: principles, methodologies, and applications in hri. ACM Trans Hum Rob Interact","DOI":"10.1145\/3656376"},{"key":"1333_CR41","doi-asserted-by":"crossref","unstructured":"Belardinelli A, Kondapally AR, Ruiken D, Tanneberg D, Watabe T (2022) Intention estimation from gaze and motion features for human-robot shared-control object manipulation. In 2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE pp 9806\u20139813","DOI":"10.1109\/IROS47612.2022.9982249"},{"key":"1333_CR42","unstructured":"Zhao Q, Wang S, Zhang C, Fu C, Do MQ, Agarwal N, Lee K, Sun C (2024) Antgpt: can large language models help long-term action anticipation from videos? ICLR"},{"key":"1333_CR43","doi-asserted-by":"crossref","unstructured":"Martinez J, Black MJ, Romero J (2017) On human motion prediction using recurrent neural networks. In Conference on Computer Vision and Pattern Recognition (CVPR), pp 2891\u20132900","DOI":"10.1109\/CVPR.2017.497"},{"key":"1333_CR44","doi-asserted-by":"crossref","unstructured":"Amirian J, Hayet J-B, Pettr\u00e9 J (2019) Social ways: learning multi-modal distributions of pedestrian trajectories with gans. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","DOI":"10.1109\/CVPRW.2019.00359"},{"key":"1333_CR45","unstructured":"Vendrow E, Kumar S, Adeli E, Rezatofighi H (2022). SoMoFormer: multi-person pose forecasting with transformers. arXiv preprint arXiv:2208.14023, 2022"},{"key":"1333_CR46","unstructured":"Kipf TN, Welling M (2016). Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907"},{"key":"1333_CR47","unstructured":"Sohl-Dickstein J, Weiss E, Maheswaranathan N, Ganguli S (2015) Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, PMLR pp 2256\u20132265"},{"key":"1333_CR48","doi-asserted-by":"crossref","unstructured":"Ahn H, Mascaro EV, Lee D (2023) Can we use diffusion probabilistic models for 3d motion prediction? In 2023 IEEE International Conference on Robotics and Automation (ICRA), IEEE pp 9837\u20139843","DOI":"10.1109\/ICRA48891.2023.10160722"},{"key":"1333_CR49","unstructured":"Tevet G, Raab S, Gordon B, Shafir Y, Cohen-Or D, Bermano AH (2023) Human motion diffusion model. In The Eleventh International Conference on Learning Representations"},{"key":"1333_CR50","doi-asserted-by":"crossref","unstructured":"Xu L, Zhou Y, Yan Y, Jin X, Zhu W, Rao F, Yang X, Zeng W (2024) Regennet: towards human action-reaction synthesis. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR52733.2024.00173"},{"key":"1333_CR51","doi-asserted-by":"crossref","unstructured":"Ghosh A, Dabral R, Golyanik V, Theobalt C, Slusallek P (2023). Remos: reactive 3d motion synthesis for two-person interactions. arXiv preprint arXiv:2311.17057","DOI":"10.1007\/978-3-031-72764-1_24"},{"key":"1333_CR52","unstructured":"Siyao L, Gu T, Yang Z, Lin Z, Liu Z, Ding H, Yang L, Loy CC (2024) Duolando: follower GPT with off-policy reinforcement learning for dance accompaniment. In The Twelfth International Conference on Learning Representations"},{"key":"1333_CR53","doi-asserted-by":"crossref","unstructured":"Gleicher M (1998) Retargetting motion to new characters. In Proceedings of the 25th annual conference on Computer graphics and interactive techniques","DOI":"10.1145\/280814.280820"},{"key":"1333_CR54","doi-asserted-by":"crossref","unstructured":"Koenemann J, Burget F, Bennewitz M (2014) Real-time imitation of human whole-body motions by humanoids","DOI":"10.1109\/ICRA.2014.6907261"},{"key":"1333_CR55","doi-asserted-by":"crossref","unstructured":"Devanne M, Nguyen SM (2017) Multi-level motion analysis for physical exercises assessment in kinaesthetic rehabilitation","DOI":"10.1109\/HUMANOIDS.2017.8246923"},{"key":"1333_CR56","doi-asserted-by":"crossref","unstructured":"Darvish K, Tirupachuri Y, Romualdi G, Rapetti L, Ferigo D, Chavez FJA, Pucci D (2019) Whole-body geometric retargeting for humanoid robots. In 2019 IEEE-RAS 19th International Conference on Humanoid Robots (Humanoids), IEEE pp 679\u2013686","DOI":"10.1109\/Humanoids43949.2019.9035059"},{"key":"1333_CR57","doi-asserted-by":"crossref","unstructured":"Penco L, Clement B, Moduano V, Mingo Hoffman E, Nava G, Pucci D, Tsagarakis N, Mourert J, Ivaldi S (2018) Robust real-time whole-body motion retargeting from human to humanoid 425\u2013432","DOI":"10.1109\/HUMANOIDS.2018.8624943"},{"key":"1333_CR58","doi-asserted-by":"crossref","unstructured":"Delhaisse B, Esteban D, Rozo L, Caldwell D (2017) Transfer learning of shared latent spaces between robots with similar kinematic structure","DOI":"10.1109\/IJCNN.2017.7966379"},{"key":"1333_CR59","doi-asserted-by":"crossref","unstructured":"Villegas R, Yang J, Ceylan D, Lee H (2018) Neural kinematic networks for unsupervised motion retargetting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp 8639\u20138648","DOI":"10.1109\/CVPR.2018.00901"},{"key":"1333_CR60","unstructured":"Lim J, Chang H, Choi J (2019) Pmnet: learning of disentangled pose and movement for unsupervised motion retargeting. In Proceedings of the 30th British Machine Vision Conference (BMVC 2019). British Machine Vision Association, BMVA, Cardiff, UK"},{"issue":"4","key":"1333_CR61","doi-asserted-by":"publisher","first-page":"62","DOI":"10.1145\/3386569.3392462","volume":"39","author":"K Aberman","year":"2020","unstructured":"Aberman K, Li P, Lischinski D, Sorkine-Hornung O, Cohen-Or D, Chen B (2020) Skeleton-aware networks for deep motion retargeting. ACM Trans Graph (TOG) 39(4):62\u20131","journal-title":"ACM Trans Graph (TOG)"},{"key":"1333_CR62","doi-asserted-by":"crossref","unstructured":"Zhang J, Weng J, Kang D, Zhao F, Huang S, Zhe X, Bao L, Shan Y, Wang J, Tu Z (2023) Skinned motion retargeting with residual perception of motion semantics & geometry. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","DOI":"10.1109\/CVPR52729.2023.01332"},{"key":"1333_CR63","doi-asserted-by":"crossref","unstructured":"Choi S, Song MJ, Ahn H, Kim J (2021) Self-supervised motion retargeting with safety guarantee. In 2021 IEEE International Conference on Robotics and Automation (ICRA), IEEE pp 8097\u20138103","DOI":"10.1109\/ICRA48506.2021.9560860"},{"key":"1333_CR64","doi-asserted-by":"crossref","unstructured":"Oreshkin BN, Valkanas A, Harvey FG, M\u00e9nard L-S, Bocquelet F, Coates MJ (2023) Motion in-betweening via deep delta-interpolator. In IEEE Transactions on Visualization and Computer Graphics","DOI":"10.1109\/TVCG.2023.3309107"},{"key":"1333_CR65","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2017) Attention is all you need. Adv Neural Inf Process Syst (NeurIPS) 30"},{"key":"1333_CR66","doi-asserted-by":"crossref","unstructured":"Petrovich M, Black MJ, Varol G (2023) TMR: text-to-motion retrieval using contrastive 3D human motion synthesis. In International Conference on Computer Vision (ICCV)","DOI":"10.1109\/ICCV51070.2023.00870"},{"key":"1333_CR67","doi-asserted-by":"crossref","unstructured":"Guo C, Zou S, Zuo X, Wang S, Ji W, Li X, Cheng L (2022) Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 5152\u20135161","DOI":"10.1109\/CVPR52688.2022.00509"},{"key":"1333_CR68","doi-asserted-by":"crossref","unstructured":"Sofianos T, Sampieri A, Franco L, Galasso F (2021) Space-time-separable graph convolutional network for pose forecasting. In IEEE\/CVF International Conference on Computer Vision, pp 11209\u201311218","DOI":"10.1109\/ICCV48922.2021.01102"},{"key":"1333_CR69","doi-asserted-by":"crossref","unstructured":"Wang C-Y, Yeh I-H, Liao H-YM (2024). Yolov9: learning what you want to learn using programmable gradient information. arXiv preprint arXiv:2402.13616","DOI":"10.1007\/978-3-031-72751-1_1"},{"key":"1333_CR70","doi-asserted-by":"crossref","unstructured":"Li J, Xu C, Chen Z, Bian S, Yang L, Lu C (2021) Hybrik: a hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp 3383\u20133393","DOI":"10.1109\/CVPR46437.2021.00339"}],"container-title":["International Journal of Social Robotics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12369-025-01333-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s12369-025-01333-3","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12369-025-01333-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T12:18:39Z","timestamp":1781612319000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s12369-025-01333-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,5]]},"references-count":70,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2025,12]]}},"alternative-id":["1333"],"URL":"https:\/\/doi.org\/10.1007\/s12369-025-01333-3","relation":{},"ISSN":["1875-4791","1875-4805"],"issn-type":[{"value":"1875-4791","type":"print"},{"value":"1875-4805","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,5]]},"assertion":[{"value":"15 June 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 July 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 October 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 November 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics Approval and Consent to Participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for Publication"}},{"value":"The authors have no conflicts of interest to declare that are relevant to the content of this article.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of Interest"}}]}}