{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T15:47:06Z","timestamp":1785512826495,"version":"3.56.0"},"reference-count":84,"publisher":"SAGE Publications","issue":"4","license":[{"start":{"date-parts":[[2024,1,22]],"date-time":"2024-01-22T00:00:00Z","timestamp":1705881600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"name":"Samsung GRO"},{"name":"ONR","award":["N00014-22-1-2096"],"award-info":[{"award-number":["N00014-22-1-2096"]}]},{"name":"NSF","award":["IIS-2024594"],"award-info":[{"award-number":["IIS-2024594"]}]},{"DOI":"10.13039\/100023581","name":"NSF GRFP","doi-asserted-by":"crossref","award":["DGE2140739"],"award-info":[{"award-number":["DGE2140739"]}],"id":[{"id":"10.13039\/100023581","id-type":"DOI","asserted-by":"crossref"}]},{"name":"GoodAI Research Award"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of Robotics Research"],"published-print":{"date-parts":[[2024,4]]},"abstract":"<jats:p>\n                    To build general robotic agents that can operate in many environments, it is often useful for robots to collect experience in the real world. However, unguided experience collection is often not feasible due to safety, time, and hardware restrictions. We thus propose leveraging the next best thing as real world experience: videos of humans using their hands. To utilize these videos, we develop a method that retargets any 1st person or 3rd person video of human hands and arms into the robot hand and arm trajectories. While retargeting is a difficult problem, our key insight is to rely on only internet human hand video to train it. We use this method to present results in two areas: First, we build a system that enables any human to control a robot hand and arm, simply by demonstrating motions with their own hand. The robot observes the human operator via a\n                    <jats:italic toggle=\"yes\">single RGB camera<\/jats:italic>\n                    and imitates their actions\n                    <jats:italic toggle=\"yes\">in real-time<\/jats:italic>\n                    . This enables the robot to collect real-world experience safely using supervision. See these results at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/robotic-telekinesis.github.io\">https:\/\/robotic-telekinesis.github.io<\/jats:ext-link>\n                    . Second, we retarget in-the-wild human internet video into task-conditioned pseudo-robot trajectories to use as artificial robot experience. This learning algorithm leverages action priors from human hand actions, visual features from the images, and physical priors from dynamical systems to pretrain typical human behavior for a particular robot task. We show that by leveraging internet human hand experience, we need fewer robot demonstrations compared to many other methods. See these results at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/video-dex.github.io\">https:\/\/video-dex.github.io<\/jats:ext-link>\n                  <\/jats:p>","DOI":"10.1177\/02783649241227559","type":"journal-article","created":{"date-parts":[[2024,1,23]],"date-time":"2024-01-23T02:49:00Z","timestamp":1705978140000},"page":"513-532","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":16,"title":["Learning dexterity from human hand motion in internet videos"],"prefix":"10.1177","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-8571-2922","authenticated-orcid":false,"given":"Kenneth","family":"Shaw","sequence":"first","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shikhar","family":"Bahl","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Aravind","family":"Sivakumar","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3104-7560","authenticated-orcid":false,"given":"Aditya","family":"Kannan","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Deepak","family":"Pathak","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2024,1,22]]},"reference":[{"key":"e_1_3_4_2_1","unstructured":"Agarwal A Uppal S Shaw K et al. (2023) Dexterous functional grasping. In: Conference on robot learning. PMLR 3453\u20133467."},{"key":"e_1_3_4_3_1","doi-asserted-by":"crossref","unstructured":"Antotsiou D Garcia-Hernando G Kim TK (2018) Task-oriented hand motion retargeting for dexterous manipulation imitation. In: Proceedings of the European conference on computer vision (ECCV) workshops. Munich Germany 8 September\u201314 September 2018.","DOI":"10.1007\/978-3-030-11024-6_19"},{"key":"e_1_3_4_4_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2203.13251"},{"key":"e_1_3_4_5_1","volume-title":"NeurIPS","author":"Bahl S","year":"2020","unstructured":"Bahl S, Mukadam M, Gupta A, et al. (2020) Neural dynamic policies for end-to-end sensorimotor learning. In: NeurIPS."},{"key":"e_1_3_4_6_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2021.XVII.023"},{"key":"e_1_3_4_7_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2022.XVIII.026"},{"key":"e_1_3_4_8_1","doi-asserted-by":"crossref","unstructured":"Bahl S Mendonca R Chen L et al. (2023) Affordances from human videos as a versatile representation for robotics. In: Proceedings of the IEEE\/CVF Conference on computer vision and pattern recognition Vancouver BC Canada 17 June\u201324 June 2023 13778\u201313790.","DOI":"10.1109\/CVPR52729.2023.01324"},{"key":"e_1_3_4_9_1","unstructured":"Bhat SF Alhashim I Wonka P (2021) Adabins: depth estimation using adaptive bins. In: Proceedings of the IEEE\/CVF Conference on computer vision and pattern recognition Nashville TN USA 20 June\u201325 June 2021 4009\u20134018."},{"key":"e_1_3_4_10_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1604.07316"},{"key":"e_1_3_4_11_1","volume-title":"JAX: Composable Transformations of Python+NumPy Programs","author":"Bradbury J","year":"2018","unstructured":"Bradbury J, Frostig R, Hawkins P, et al. (2018) JAX: Composable Transformations of Python+NumPy Programs. https:\/\/github.com\/google\/jax"},{"key":"e_1_3_4_12_1","volume-title":"Language Models Are Few-Shot Learners","author":"Brown TB","year":"2020","unstructured":"Brown TB, Mann B, Ryder N, et al. (2020) Language Models Are Few-Shot Learners. arXiv."},{"key":"e_1_3_4_13_1","doi-asserted-by":"publisher","DOI":"10.1080\/2151237X.2005.10129202"},{"key":"e_1_3_4_14_1","doi-asserted-by":"publisher","DOI":"10.1080\/2151237X.2005.10129202"},{"key":"e_1_3_4_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2021.3075644"},{"key":"e_1_3_4_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2929257"},{"key":"e_1_3_4_17_1","doi-asserted-by":"crossref","unstructured":"Carpentier J Saurel G Buondonno G et al. (2019) The pinocchio c++ library \u2013 a fast and flexible implementation of rigid body dynamics algorithms and their analytical derivatives. In: IEEE International Symposium on System Integrations (SII). IEEE.","DOI":"10.1109\/SII.2019.8700380"},{"key":"e_1_3_4_18_1","doi-asserted-by":"crossref","unstructured":"Chao YW Yang W Xiang Y et al. (2021) DexYCB: a benchmark for capturing hand grasping of objects. In: IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Nashville TN USA 20 June\u201325 June 2021.","DOI":"10.1109\/CVPR46437.2021.00893"},{"key":"e_1_3_4_19_1","unstructured":"Chen T Kornblith S Norouzi M et al. (2020) A simple framework for contrastive learning of visual representations. In: III HD Singh A (eds) Proceedings of the 37th International conference on machine learning proceedings of machine learning research. PMLR Vol. 119 1597\u20131607. URL https:\/\/proceedings.mlr.press\/v119\/chen20j.html"},{"key":"e_1_3_4_20_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2021.XVII.012"},{"key":"e_1_3_4_21_1","doi-asserted-by":"crossref","unstructured":"Damen D Doughty H Farinella GM et al. (2018) Scaling egocentric vision: the epic-kitchens dataset. In: European conference on computer vision (ECCV). Springer.","DOI":"10.1007\/978-3-030-01225-0_44"},{"key":"e_1_3_4_22_1","doi-asserted-by":"crossref","unstructured":"Das P Xu C Doell RF et al. (2013) A thousand frames in just a few words: lingual description of videos through latent topics and sparse object stitching. In: Proceedings of the IEEE Conference on computer vision and pattern recognition. IEEE 2634\u20132641.","DOI":"10.1109\/CVPR.2013.340"},{"key":"e_1_3_4_23_1","volume-title":"Model-based Inverse Reinforcement Learning from Visual Demonstrations","author":"Das N","year":"2020","unstructured":"Das N, Bechtle S, Davchev T, et al. (2020) Model-based Inverse Reinforcement Learning from Visual Demonstrations. arXiv preprint arXiv:2010.09034."},{"key":"e_1_3_4_24_1","unstructured":"Dasari S Wang J Hong J et al. (2021) Rb2: robotic manipulation benchmarking with a twist. In: NeurIPS datasets and benchmarks track (Round 2). NIPS."},{"key":"e_1_3_4_25_1","doi-asserted-by":"crossref","unstructured":"Deng J Dong W Socher R et al. (2009) Imagenet: a large-scale hierarchical image database. In: 2009 IEEE Conference on computer vision and pattern recognition. IEEE pp. 248\u2013255.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_4_26_1","volume-title":"Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding","author":"Devlin J","year":"2018","unstructured":"Devlin J, Chang MW, Lee K, et al. (2018) Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805."},{"key":"e_1_3_4_27_1","doi-asserted-by":"publisher","DOI":"10.26599\/TST.2018.9010096"},{"key":"e_1_3_4_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/358669.358692"},{"key":"e_1_3_4_29_1","doi-asserted-by":"crossref","unstructured":"Goyal R Ebrahimi Kahou S Michalski V et al. (2017) The \u201dsomething something\u201d video database for learning and evaluating visual common sense. In: Proceedings of the IEEE International conference on computer vision (ICCV). IEEE.","DOI":"10.1109\/ICCV.2017.622"},{"key":"e_1_3_4_30_1","unstructured":"Grauman K Westbury A Byrne E et al. (2022) Ego4d: around the world in 3 000 hours of egocentric video. In: Proceedings of the IEEE\/CVF Conference on computer vision and pattern recognition. IEEE 18995\u201319012."},{"key":"e_1_3_4_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201399"},{"key":"e_1_3_4_32_1","doi-asserted-by":"publisher","unstructured":"Handa A Van Wyk K Yang W et al. (2020) Dexpilot: vision-based teleoperation of dexterous robotic hand-arm system. In: 2020 IEEE International conference on robotics and automation (ICRA). IEEE 9164\u20139170. DOI: 10.1109\/ICRA40945.2020.9197124","DOI":"10.1109\/ICRA40945.2020.9197124"},{"key":"e_1_3_4_33_1","volume-title":"Deep Residual Learning for Image Recognition","author":"He K","year":"2015","unstructured":"He K, Zhang X, Ren S, et al. (2015) Deep Residual Learning for Image Recognition. CoRR abs\/1512.03385. URL https:\/\/arxiv.org\/abs\/1512.03385"},{"key":"e_1_3_4_34_1","doi-asserted-by":"crossref","unstructured":"He K Gkioxari G Dollar P et al. (2017) Mask r-cnn. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). IEEE.","DOI":"10.1109\/ICCV.2017.322"},{"key":"e_1_3_4_35_1","doi-asserted-by":"crossref","unstructured":"He K Chen X Xie S et al. (2022) Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE 16000\u201316009.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_4_36_1","doi-asserted-by":"crossref","unstructured":"Heilbron FC Escorcia V Ghanem B et al. (2015) Activitynet: a large-scale video benchmark for human activity understanding. In: CVPR 961\u2013970.","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_4_37_1","volume-title":"Cmu Graphics Lab Motion Capture Database","author":"Hodgins J","unstructured":"Hodgins J (n.d) Cmu Graphics Lab Motion Capture Database. https:\/\/mocap.cs.cmu.edu\/"},{"key":"e_1_3_4_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.248"},{"key":"e_1_3_4_39_1","volume-title":"Qt-opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation","author":"Kalashnikov D","year":"2018","unstructured":"Kalashnikov D, Irpan A, Pastor P, et al. (2018) Qt-opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation. arXiv preprint arXiv:1806.10293."},{"key":"e_1_3_4_40_1","volume-title":"End-to-End Recovery of Human Shape and Pose","author":"Kanazawa A","year":"2017","unstructured":"Kanazawa A, Black MJ, Jacobs DW, et al. (2017) End-to-End Recovery of Human Shape and Pose. CoRR abs\/171206584. URL https:\/\/arxiv.org\/abs\/1712.06584"},{"key":"e_1_3_4_41_1","volume-title":"Deft: Dexterous Fine-Tuning for Real-World Hand Policies","author":"Kannan A","year":"2023","unstructured":"Kannan A, Shaw K, Bahl S, et al. (2023) Deft: Dexterous Fine-Tuning for Real-World Hand Policies. CoRL."},{"key":"e_1_3_4_42_1","doi-asserted-by":"publisher","unstructured":"Kumar V Todorov E (2015) Mujoco haptix: a virtual reality system for hand manipulation. In: 2015 IEEE-RAS 15th International conference on humanoid robots (Humanoids) Seoul Korea 3 November\u20135 November 2015 657\u2013663. DOI: 10.1109\/HUMANOIDS.2015.7363441","DOI":"10.1109\/HUMANOIDS.2015.7363441"},{"key":"e_1_3_4_43_1","first-page":"1179","article-title":"Conservative q-learning for offline reinforcement learning","volume":"33","author":"Kumar A","year":"2020","unstructured":"Kumar A, Zhou A, Tucker G, et al. (2020) Conservative q-learning for offline reinforcement learning. Advances in Neural Information Processing Systems 33: 1179\u20131191.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_44_1","doi-asserted-by":"crossref","unstructured":"Lee J Ryoo MS (2017) Learning robot activities from first-person human videos using convolutional future regression. In: CVPR Workshops 1\u20132.","DOI":"10.1109\/IROS.2017.8205953"},{"key":"e_1_3_4_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMECH.2016.2634602"},{"key":"e_1_3_4_46_1","volume-title":"End-to-end Training of Deep Visuomotor Policies","author":"Levine S","year":"2016","unstructured":"Levine S, Finn C, Darrell T, et al. (2016) End-to-end Training of Deep Visuomotor Policies. JMLR."},{"key":"e_1_3_4_47_1","doi-asserted-by":"crossref","unstructured":"Li S Ma X Liang H et al. (2019) Vision-based teleoperation of shadow dexterous hand using end-to-end deep neural network. In: 2019 International Conference on Robotics and Automation (ICRA). IEEE 416\u2013422.","DOI":"10.1109\/ICRA.2019.8794277"},{"key":"e_1_3_4_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818013"},{"key":"e_1_3_4_49_1","doi-asserted-by":"publisher","unstructured":"Ma RR Dollar AM (2011) On dexterity and dexterous manipulation. In: 2011 15th International Conference on Advanced Robotics. ICAR Tallinn Estonia 20 June\u201323 June 2011 1\u20137. DOI: 10.1109\/ICAR.2011.6088576","DOI":"10.1109\/ICAR.2011.6088576"},{"key":"e_1_3_4_50_1","volume-title":"Isaac Gym: High Performance Gpu-Based Physics Simulation for Robot Learning","author":"Makoviychuk V","year":"2021","unstructured":"Makoviychuk V, Wawrzyniak L, Guo Y, et al. (2021) Isaac Gym: High Performance Gpu-Based Physics Simulation for Robot Learning. arXiv preprint arXiv:2108.10470."},{"key":"e_1_3_4_51_1","unstructured":"Mandikal P Grauman K (2022) Dexvip: learning dexterous grasping with human hand pose priors from video. In: Conference on Robot Learning Atlanta GA 6 November\u20139 November 2023. PMLR 651\u2013661."},{"key":"e_1_3_4_52_1","doi-asserted-by":"publisher","unstructured":"Mannam P Shaw K Bauer D et al. (2023) Designing anthropomorphic soft hands through interaction. In: 2023 IEEE-RAS 22nd International Conference on Humanoid Robots (Humanoids). IEEE 1\u20138. DOI: 10.1109\/Humanoids57100.2023.10375195","DOI":"10.1109\/Humanoids57100.2023.10375195"},{"key":"e_1_3_4_53_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2023.XIX.012"},{"key":"e_1_3_4_54_1","first-page":"9191","volume-title":"NeurIPS","author":"Nair AV","year":"2018","unstructured":"Nair AV, Pong V, Dalal M, et al. (2018) Visual reinforcement learning with imagined goals. In: NeurIPS. pp: 9191\u20139200."},{"key":"e_1_3_4_55_1","volume-title":"R3m: A Universal Visual Representation for Robot Manipulation","author":"Nair S","year":"2022","unstructured":"Nair S, Rajeswaran A, Kumar V, et al. (2022) R3m: A Universal Visual Representation for Robot Manipulation. arXiv preprint arXiv:2203.12601."},{"key":"e_1_3_4_56_1","volume-title":"The Surprising Effectiveness of Representation Learning for Visual Imitation","author":"Pari J","year":"2021","unstructured":"Pari J, Muhammad N, Arunachalam SP, et al. (2021) The Surprising Effectiveness of Representation Learning for Visual Imitation. arXiv preprint arXiv:2112.01511."},{"key":"e_1_3_4_57_1","doi-asserted-by":"crossref","unstructured":"Pavlakos G Choutas V Ghorbani N et al. (2019) Expressive body capture: 3D hands face and body from a single image. In: Proceedings IEEE Conf. On Computer Vision and Pattern Recognition (CVPR). IEEE 10975\u201310985.","DOI":"10.1109\/CVPR.2019.01123"},{"key":"e_1_3_4_58_1","volume-title":"Learning Agile Robotic Locomotion Skills by Imitating Animals","author":"Peng XB","year":"2020","unstructured":"Peng XB, Coumans E, Zhang T, et al. (2020) Learning Agile Robotic Locomotion Skills by Imitating Animals. arXiv preprint arXiv:2004.00784."},{"key":"e_1_3_4_59_1","volume-title":"The Curious Robot: Learning Visual Representations via Physical Interactions","author":"Pinto L","year":"2016","unstructured":"Pinto L, Gandhi D, Han Y, et al. (2016) The Curious Robot: Learning Visual Representations via Physical Interactions. ECCV."},{"key":"e_1_3_4_60_1","volume-title":"Advances in neural information processing systems","author":"Pomerleau DA","year":"1988","unstructured":"Pomerleau DA (1988) Alvinn: an autonomous land vehicle in a neural network. In: Touretzky D (ed) Advances in neural information processing systems. Morgan-Kaufmann, Vol. 1. https:\/\/proceedings.neurips.cc\/paper\/1988\/file\/812b4ba287f5ee0bc9d43bbf5bbe87fb-Paper.pdf"},{"key":"e_1_3_4_61_1","volume-title":"Dexmv: Imitation Learning for Dexterous Manipulation from Human Videos","author":"Qin Y","year":"2021","unstructured":"Qin Y, Wu YH, Liu S, et al. (2021) Dexmv: Imitation Learning for Dexterous Manipulation from Human Videos. arXiv preprint arXiv:2108.05877."},{"key":"e_1_3_4_62_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2204.12490"},{"key":"e_1_3_4_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3130800.3130883"},{"key":"e_1_3_4_64_1","doi-asserted-by":"crossref","unstructured":"Rong Y Shiratori T Joo H (2021) Frankmocap: a monocular 3d whole-body pose estimation system via regression and integration. In: Proceedings of the IEEE\/CVF International conference on computer vision (ICCV) workshops. IEEE 1749\u20131759.","DOI":"10.1109\/ICCVW54120.2021.00201"},{"key":"e_1_3_4_65_1","volume-title":"Reinforcement Learning with Videos: Combining Offline Observations with Interaction","author":"Schmeckpeper K","year":"2020","unstructured":"Schmeckpeper K, Rybkin O, Daniilidis K, et al. (2020) Reinforcement Learning with Videos: Combining Offline Observations with Interaction. arXiv preprint arXiv:2011.06507."},{"key":"e_1_3_4_66_1","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nberger JL Zheng E Pollefeys M et al. (2016) Pixelwise view selection for unstructured multi-view stereo. In: European conference on computer vision (ECCV). Springer Science.","DOI":"10.1007\/978-3-319-46487-9_31"},{"key":"e_1_3_4_67_1","doi-asserted-by":"crossref","unstructured":"Sermanet P Lynch C Chebotar Y et al. (2018) Time-contrastive networks: self-supervised learning from video. In: ICRA Brisbane QLD Australia. IEEE.","DOI":"10.1109\/ICRA.2018.8462891"},{"key":"e_1_3_4_68_1","doi-asserted-by":"crossref","unstructured":"Shan D Geng J Shu M et al. (2020) Understanding human hands in contact at internet scale. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE 9869\u20139878.","DOI":"10.1109\/CVPR42600.2020.00989"},{"key":"e_1_3_4_69_1","doi-asserted-by":"publisher","DOI":"10.1177\/02783649211046285"},{"key":"e_1_3_4_70_1","volume-title":"Third-Person Visual Imitation Learning via Decoupled Hierarchical Controller","author":"Sharma P","year":"2019","unstructured":"Sharma P, Pathak D, Gupta A (2019) Third-Person Visual Imitation Learning via Decoupled Hierarchical Controller. arXiv preprint arXiv:1911.09676 32."},{"key":"e_1_3_4_71_1","article-title":"Leap Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot Learning","author":"Shaw K","year":"2023","unstructured":"Shaw K, Agarwal A, Pathak D (2023a) Leap Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot Learning. RSS.","journal-title":"RSS"},{"key":"e_1_3_4_72_1","unstructured":"Shaw K Bahl S Pathak D (2023b) Videodex: learning dexterity from internet videos. In: Conference on robot learning. PMLR 654\u2013665."},{"key":"e_1_3_4_73_1","volume-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","author":"Simonyan K","year":"2014","unstructured":"Simonyan K, Zisserman A (2014) Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint arXiv:1409.1556."},{"key":"e_1_3_4_74_1","volume-title":"Robotic Telekinesis: Learning a Robotic Hand Imitator by Watching Humans on Youtube","author":"Sivakumar A","year":"2022","unstructured":"Sivakumar A, Shaw K, Pathak D (2022) Robotic Telekinesis: Learning a Robotic Hand Imitator by Watching Humans on Youtube. arXiv."},{"key":"e_1_3_4_75_1","volume-title":"Avid: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos","author":"Smith L","year":"2020","unstructured":"Smith L, Dhawan N, Zhang M, et al. (2020) Avid: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos. RSS."},{"key":"e_1_3_4_76_1","volume-title":"MuJoCo: A Physics Engine for Model-Based Control","author":"Todorov E","year":"2012","unstructured":"Todorov E, Erez T, Tassa Y (2012) MuJoCo: A Physics Engine for Model-Based Control. IROS."},{"key":"e_1_3_4_77_1","unstructured":"UFactory (n.d) xarm6 by ufactory. https:\/\/www.ufactory.cc\/xarm-collaborative-robot"},{"key":"e_1_3_4_78_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.88573"},{"key":"e_1_3_4_79_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00901"},{"key":"e_1_3_4_80_1","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417788"},{"key":"e_1_3_4_81_1","volume-title":"Masked Visual Pre-training for Motor Control","author":"Xiao T","year":"2022","unstructured":"Xiao T, Radosavovic I, Darrell T, et al. (2022) Masked Visual Pre-training for Motor Control. arXiv preprint arXiv:2203.06173."},{"key":"e_1_3_4_82_1","volume-title":"Visual Imitation Made Easy","author":"Young S","year":"2020","unstructured":"Young S, Gandhi D, Tulsiani S, et al. (2020) Visual Imitation Made Easy. arXiv preprint arXiv:2008.04899."},{"key":"e_1_3_4_83_1","volume-title":"Xirl: Cross-Embodiment Inverse Reinforcement Learning","author":"Zakka K","year":"2021","unstructured":"Zakka K, Zeng A, Florence P, et al. (2021) Xirl: Cross-Embodiment Inverse Reinforcement Learning. arXiv preprint arXiv:2106.03911."},{"key":"e_1_3_4_84_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20077-9_21"},{"key":"e_1_3_4_85_1","doi-asserted-by":"crossref","unstructured":"Zimmermann C Ceylan D Yang J et al. (2019) Freihand: a dataset for markerless capture of hand pose and shape from single RGB images. In: Proceedings of the IEEE\/CVF International conference on computer vision Seoul Korea 27 October\u201328 October 2019 813\u2013822.","DOI":"10.1109\/ICCV.2019.00090"}],"container-title":["The International Journal of Robotics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649241227559","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/02783649241227559","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649241227559","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649241227559","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T09:41:46Z","timestamp":1782294106000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/02783649241227559"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,22]]},"references-count":84,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,4]]}},"alternative-id":["10.1177\/02783649241227559"],"URL":"https:\/\/doi.org\/10.1177\/02783649241227559","relation":{},"ISSN":["0278-3649","1741-3176"],"issn-type":[{"value":"0278-3649","type":"print"},{"value":"1741-3176","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,22]]}}}