{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T14:54:04Z","timestamp":1784645644802,"version":"3.55.0"},"reference-count":99,"publisher":"American Association for the Advancement of Science (AAAS)","issue":"110","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62088101"],"award-info":[{"award-number":["62088101"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62233013"],"award-info":[{"award-number":["62233013"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62293511"],"award-info":[{"award-number":["62293511"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["www.science.org"],"crossmark-restriction":true},"short-container-title":["Sci. Robot."],"published-print":{"date-parts":[[2026,1,28]]},"abstract":"<jats:p>Achieving humanlike dexterity with anthropomorphic multifingered robotic hands requires precise finger coordination. However, dexterous manipulation remains highly challenging because of high-dimensional action-observation spaces, complex hand-object contact dynamics, and frequent occlusions. To address this, we drew inspiration from the human learning paradigm of observation and practice and propose a two-stage learning framework by learning visual-tactile integration representations via self-supervised learning from human demonstrations. We trained a unified multitask policy through reinforcement learning and online imitation learning. This decoupled learning enabled the robot to acquire generalizable manipulation skills using only monocular images and simple binary tactile signals. With the unified policy, we built a multifingered hand manipulation system that performs multiple complicated tasks with low-cost sensing. It achieved an 85% success rate across five complex tasks and 25 objects and further generalized to three unseen tasks that share similar hand-object coordination patterns with the training tasks.<\/jats:p>","DOI":"10.1126\/scirobotics.ady2869","type":"journal-article","created":{"date-parts":[[2026,1,28]],"date-time":"2026-01-28T18:58:46Z","timestamp":1769626726000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark","source":"Crossref","is-referenced-by-count":6,"title":["Visual-tactile pretraining and online multitask learning for humanlike manipulation dexterity"],"prefix":"10.1126","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2285-3402","authenticated-orcid":true,"given":"Qi","family":"Ye","sequence":"first","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-4313-3538","authenticated-orcid":true,"given":"Qingtao","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Siyun","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiaying","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Cui","sequence":"additional","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-7855-8033","authenticated-orcid":true,"given":"Ke","family":"Jin","sequence":"additional","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2107-2048","authenticated-orcid":true,"given":"Huajin","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xuan","family":"Cai","sequence":"additional","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9369-2928","authenticated-orcid":true,"given":"Gaofeng","family":"Li","sequence":"additional","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3155-3145","authenticated-orcid":true,"given":"Jiming","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China."},{"name":"School of Automation, Hangzhou Dianzi University, Hangzhou 310018, China."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"221","reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.adc9244"},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","unstructured":"G. Solak L. Jamone \u201cLearning by demonstration and robust control of dexterous in-hand robotic manipulation skills \u201d in Proceedings of 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2019) pp. 8246\u20138251.","DOI":"10.1109\/IROS40897.2019.8967567"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3145961"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TOH.2023.3300439"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2023.3280028"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/MRA.2024.3433110"},{"key":"e_1_3_2_8_2","unstructured":"I. Mordatch Z. Popovi\u0107 E. Todorov \u201cContact-invariant optimization for hand manipulation \u201d in Proceedings of the ACM SIGGRAPH\/Eurographics Symposium on Computer Animation (ACM 2012) pp. 137\u2013144."},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"V. Kumar Y. Tassa T. Erez E. Todorov \u201cReal-time behaviour synthesis for dynamic hand-manipulation \u201d in Proceedings of 2014 IEEE International Conference on Robotics and Automation (IEEE 2014) pp. 6808\u20136815.","DOI":"10.1109\/ICRA.2014.6907864"},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","unstructured":"G. J. Pollayil G. Grioli M. Bonilla A. Bicchi \u201cPlanning robotic manipulation with tight environment constraints \u201d in Proceedings of 2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2021) pp. 9385\u20139392.","DOI":"10.1109\/IROS51168.2021.9636782"},{"key":"e_1_3_2_11_2","unstructured":"A. Nagabandi K. Konolige S. Levine V. Kumar \u201cDeep dynamics models for learning dexterous manipulation \u201d in Proceedings of 2020 Conference on Robot Learning (PMLR 2020) pp. 1101\u20131112."},{"key":"e_1_3_2_12_2","unstructured":"X. Zhu J. H. Ke Z. Xu Z. Sun B. Bai J. Lv Q. Liu Y. Zeng Q. Ye C. Lu M. Tomizuka L. Shao \u201cDiff-lfd: Contact-aware model-based learning from visual demonstration for robotic manipulation via differentiable physics-based simulation and rendering \u201d in Proceedings of 2023 Conference on Robot Learning (PMLR 2023) pp. 499\u2013512."},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","unstructured":"Y. Jiang M. Yu X. Zhu M. Tomizuka X. Li \u201cContact-implicit model predictive control for dexterous in-hand manipulation: A long-horizon and robust approach \u201d in Proceedings of 2024 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2024) pp. 5260\u20135266.","DOI":"10.1109\/IROS58592.2024.10801751"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abd8803"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abd7710"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","unstructured":"W. Yuan J. A. Stork D. Kragic M. Y. Wang K. Hang \u201cRearrangement with nonprehensile manipulation using deep reinforcement learning \u201d in Proceedings of 2018 IEEE International Conference on Robotics and Automation (IEEE 2018) pp. 270\u2013277.","DOI":"10.1109\/ICRA.2018.8462863"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1080\/01691864.2023.2234428"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"A. Handa A. Allshire V. Makoviychuk A. Petrenko R. Singh J. Liu D. Makoviichuk K. Van Wyk A. Zhurkevich B. Sundaralingam Y. Narang J.-F. Lafleche D. Fox G. State \u201cDextreme: Transfer of agile in-hand manipulation from simulation to reality \u201d in Proceedings of 2023 IEEE International Conference on Robotics and Automation (IEEE 2023) pp. 5977\u20135984.","DOI":"10.1109\/ICRA48891.2023.10160216"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2024.104904"},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","unstructured":"A. Rajeswaran V. Kumar A. Gupta G. Vezzani J. Schulman E. Todorov S. Levine \u201cLearning complex dexterous manipulation with deep reinforcement learning and demonstrations \u201d in Proceedings of Robotics: Science and Systems XIV (RSS 2018); 10.15607\/RSS.2018.XIV.049.","DOI":"10.15607\/RSS.2018.XIV.049"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","unstructured":"Y. Qin Y.-H. Wu S. Liu H. Jiang R. Yang Y. Fu X. Wang \u201cDexmv: Imitation learning for dexterous manipulation from human videos \u201d in Proceedings of the European Conference on Computer Vision (Springer 2022) pp. 570\u2013587.","DOI":"10.1007\/978-3-031-19842-7_33"},{"key":"e_1_3_2_22_2","unstructured":"Y.-H. Wu J. Wang X. Wang \u201cLearning generalizable dexterous manipulation from human grasp affordance \u201d in Proceedings of 2023 Conference on Robot Learning (PMLR 2023) pp. 618\u2013629."},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"Q. Liu Y. Cui Q. Ye Z. Sun H. Li G. Li L. Shao J. Chen \u201cDexrepnet: Learning dexterous robotic grasping network with geometric and spatial hand-object representations \u201d in Proceedings of 2023 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2023) pp. 3153\u20133160.","DOI":"10.1109\/IROS55552.2023.10342334"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"P. Mandikal K. Grauman \u201cLearning dexterous grasping with object-centric visual affordances \u201d in Proceedings of 2021 IEEE International Conference on Robotics and Automation (IEEE 2021) pp. 6169\u20136176.","DOI":"10.1109\/ICRA48506.2021.9561802"},{"key":"e_1_3_2_25_2","unstructured":"P. Mandikal K. Grauman \u201cDexvip: Learning dexterous grasping with human hand pose priors from video \u201d in Proceedings of 2022 Conference on Robot Learning (PMLR 2022) pp. 651\u2013661."},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"Y. Xu W. Wan J. Zhang H. Liu Z. Shan H. Shen R. Wang H. Geng Y. Weng J. Chen T. Liu L. Yi H. Wang \u201cUnidexgrasp: Universal robotic dexterous grasping via learning diverse proposal generation and goal-conditioned policy \u201d in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (IEEE 2023) pp. 4737\u20134746.","DOI":"10.1109\/CVPR52729.2023.00459"},{"key":"e_1_3_2_27_2","doi-asserted-by":"crossref","unstructured":"J. Borja-Diaz O. Mees G. Kalweit L. Hermann J. Boedecker W. Burgard \u201cAffordance learning from play for sample-efficient policy learning \u201d in Proceedings of 2022 International Conference on Robotics and Automation (ACM 2022) pp. 6372\u20136378.","DOI":"10.1109\/ICRA46639.2022.9811889"},{"key":"e_1_3_2_28_2","unstructured":"T. Silver K. Allen J. Tenenbaum L. Kaelbling Residual policy learning. arXiv:1812.06298 [cs.RO] (2018)."},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","unstructured":"K. Li P. Li T. Liu Y. Li S. Huang \u201cManipTrans: Efficient dexterous bimanual manipulation transfer via residual learning \u201d in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (IEEE 2025) pp. 6991\u20137003.","DOI":"10.1109\/CVPR52734.2025.00656"},{"key":"e_1_3_2_30_2","article-title":"ManiDext: Hand-object manipulation synthesis via continuous correspondence embeddings and residual-guided diffusion","author":"Zhang J.","year":"2025","unstructured":"J. Zhang, Y. Zhang, L. An, M. Li, H. Zhang, Z. Hu, Y. Liu, ManiDext: Hand-object manipulation synthesis via continuous correspondence embeddings and residual-guided diffusion. IEEE Trans. Pattern Anal. Mach. Intell. 10.1109\/TPAMI.2025.3588302 (2025).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_31_2","unstructured":"H. Qi B. Yi S. Suresh M. Lambeta Y. Ma R. Calandra J. Malik \u201cGeneral in-hand object rotation with vision and touch \u201d in Proceedings of 2023 Conference on Robot Learning (PMLR 2023) pp. 2549\u20132564."},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Y. Yuan H. Che Y. Qin B. Huang Z.-H. Yin K.-W. Lee Y. Wu S.-C. Lim X. Wang \u201cRobot synesthesia: In-hand manipulation with visuotactile sensing \u201d in Proceedings of 2024 IEEE International Conference on Robotics and Automation (IEEE 2024) pp. 6558\u20136565.","DOI":"10.1109\/ICRA57147.2024.10610532"},{"key":"e_1_3_2_33_2","unstructured":"T.-W. Ke N. Gkanatsios K. Fragkiadaki \u201c3D diffuser actor: Policy diffusion with 3D scene representations \u201d in Proceedings of 2024 Conference on Robot Learning (PMLR 2024) pp. 1949\u20131974."},{"key":"e_1_3_2_34_2","unstructured":"T. Zhang Y. Hu H. Cui H. Zhao Y. Gao \u201cA universal semantic-geometric representation for robotic manipulation \u201d in Proceedings of 2023 Conference on Robot Learning (PMLR 2023) pp. 3342\u20133363."},{"key":"e_1_3_2_35_2","unstructured":"J. Duan W. Yuan W. Pumacay Y. R. Wang K. Ehsani D. Fox R. Krishna \u201cManipulate-anything: Automating real-world robots using vision-language models \u201d in Proceedings of 2024 Conference on Robot Learning (PMLR 2024) pp. 5326\u20135350."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2023.3286071"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3296371"},{"key":"e_1_3_2_38_2","unstructured":"I. Guzey B. Evans S. Chintala L. Pinto \u201cDexterity from touch: Self-supervised pre-training of tactile representations with robotic play \u201d in Proceedings of 2023 Conference on Robot Learning (PMLR 2023) pp. 3142\u20133166."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMECH.2024.3400789"},{"key":"e_1_3_2_40_2","unstructured":"R. Calandra A. Owens M. Upadhyaya W. Yuan J. Lin E. H. Adelson S. Levine \u201cThe feeling of success: Does touch sensing help predict grasp outcomes? \u201d in Proceedings of 2017 Conference on Robot Learning (PMLR 2017) pp. 314\u2013323."},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","unstructured":"C. Sferrazza Y. Seo H. Liu Y. Lee P. Abbeel \u201cThe power of the senses: Generalizable manipulation from vision and touch through masked multimodal learning \u201d in Proceedings of 2024 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2024) pp. 9698\u20139705.","DOI":"10.1109\/IROS58592.2024.10802719"},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","unstructured":"R. Liu X. Liu \u201cMu-mae: Multimodal masked autoencoders-based one-shot learning \u201d in Proceedings of 2024 IEEE 7th International Conference on Multimedia Information Processing and Retrieval (IEEE 2024) pp. 253\u2013259.","DOI":"10.1109\/MIPR62202.2024.00048"},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","unstructured":"T. Lin Y. Zhang Q. Li H. Qi B. Yi S. Levine J. Malik \u201cLearning visuotactile skills with two multifingered hands \u201d in Proceedings of 2025 IEEE International Conference on Robotics and Automation (IEEE 2024) pp. 5637\u20135643.","DOI":"10.1109\/ICRA55743.2025.11128180"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2019.2959445"},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","unstructured":"A. Brohan N. Brown J. Carbajal Y. Chebotar J. Dabis C. Finn K. Gopalakrishnan K. Hausman A. Herzog J. Hsu J. Ibarz B. Ichter A. Irpan T. Jackson S. Jesmonth N. J. Joshi R. Julian D. Kalashnikov Y. Kuang I. Leal K.-H. Lee S. Levine Y. Lu U. Malla D. Manjunath I. Mordatch O. Nachum C. Parada J. Peralta E. Perez K. Pertsch J. Quiambao K. Rao M. Ryoo G. Salazar P. Sanketi K. Sayed J. Singh S. Sontakke A. Stone C. Tan H. Tran V. Vanhoucke S. Vega Q. Vuong F. Xia T. Xiao P. Xu S. Xu T. Yu B. Zitkovich \u201cRt-1: Robotics transformer for real-world control at scale \u201d in Proceedings of Robotics: Science and Systems (RSS 2023).","DOI":"10.15607\/RSS.2023.XIX.025"},{"key":"e_1_3_2_46_2","unstructured":"S. Dasari F. Ebert S. Tian S. Nair B. Bucher K. Schmeckpeper S. Singh S. Levine C. Finn Robonet: \u201cLarge-scale multi-robot learning \u201d in Proceedings of 2019 Conference on Robot Learning (PMLR 2019) pp. 885\u2013897."},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","unstructured":"H. Bharadhwaj J. Vakil M. Sharma A. Gupta S. Tulsiani V. Kumar \u201cRoboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking \u201d in Proceedings of 2024 IEEE International Conference on Robotics and Automation (IEEE 2024) pp. 4788\u20134795.","DOI":"10.1109\/ICRA57147.2024.10611293"},{"key":"e_1_3_2_48_2","unstructured":"E. Jang A. Irpan M. Khansari D. Kappler F. Ebert C. Lynch S. Levine C. Finn \u201cBc-z: Zero-shot task generalization with robotic imitation learning \u201d in Proceedings of 2022 Conference on Robot Learning (PMLR 2022) pp. 991\u20131002."},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","unstructured":"A. Khazatsky K. Pertsch S. Nair A. Balakrishna S. Dasari S. Karamcheti S. Nasiriany M. K. Srirama L. Y. Chen K. Ellis P. D. Fagan J. Hejna M. Itkina M. Lepert Y. J. Ma P. T. Miller J. Wu S. Belkhale S. Dass H. Ha A. Jain A. Lee Y. Lee M. Memmel S. Park I. Radosavovic K. Wang A. Zhan K. Black C. Chi K. B. Hatch S. Lin J. Lu J. Mercat A. Rehman P. R. Sanketi A. Sharma C. Simpson Q. Vuong H. R. Walke B. Wulfe T. Xiao J. H. Yang A. Yavary T. Z. Zhao C. Agia R. Baijal M. G. Castro D. Chen Q. Chen T. Chung J. Drake E. P. Foster J. Gao V. Guizilini D. A. Herrera M. Heo K. Hsu J. Hu M. Z. Irshad D. Jackson C. Le Y. Li K. Lin R. Lin Z. Ma A. Maddukuri S. Mirchandani D. Morton T. Nguyen A. O\u2019Neill R. Scalise D. Seale V. Son S. Tian E. Tran A. E. Wang Y. Wu A. Xie J. Yang P. Yin Y. Zhang O. Bastani G. Berseth J. Bohg K. Goldberg A. Gupta A. Gupta D. Jayaraman J. J. Lim J. Malik R. Mart\u00edn-Mart\u00edn S. Ramamoorthy D. Sadigh S. Song J. Wu M. C. Yip Y. Zhu T. Kollar S. Levine C. Finn \u201cDroid: A large-scale in-the-wild robot manipulation dataset \u201d in Proceedings of Robotics: Science and Systems (RSS 2024).","DOI":"10.15607\/RSS.2024.XX.120"},{"key":"e_1_3_2_50_2","unstructured":"H. R. Walke K. Black T. Z. Zhao Q. Vuong C. Zheng P. Hansen-Estruch A. W. He V. Myers M. J. Kim M. Du A. Lee K. Fang C. Finn S. Levine \u201cBridgedata v2: A dataset for robot learning at scale \u201d in Proceedings of 2023 Conference on Robot Learning (PMLR 2023) pp. 1723\u20131736."},{"key":"e_1_3_2_51_2","unstructured":"Y. Liu Y. Yang Y. Wang X. Wu J. Wang Y. Yao S. Schwertfeger S. Yang W. Wang J. Yu X. He Y. Ma \u201cRealDex: Towards human-like grasping for robotic dexterous hand \u201d in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (ACM 2024) pp. 6859\u20136867."},{"key":"e_1_3_2_52_2","unstructured":"S. Nair A. Rajeswaran V. Kumar C. Finn A. Gupta \u201cR3M: A universal visual representation for robot manipulation \u201d in Proceedings of 2022 Conference on Robot Learning (PMLR 2022) pp. 892\u2013909."},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","unstructured":"S. Karamcheti S. Nair A. Chen T. Kollar \u201cLanguage-driven representation learning for robotics \u201d in Proceedings of Robotics: Science and Systems (RSS 2023).","DOI":"10.15607\/RSS.2023.XIX.032"},{"key":"e_1_3_2_54_2","unstructured":"R. Tian C. Xu M. Tomizuka J. Malik A. Bajcsy \u201cVIP: Towards universal visual reward and representation via value-implicit pre-training \u201d in Proceedings of the Eleventh International Conference on Learning Representations (ICLR 2023)."},{"key":"e_1_3_2_55_2","unstructured":"I. Radosavovic T. Xiao S. James P. Abbeel J. Malik T. Darrell \u201cReal-world robot learning with masked visual pre-training \u201d in Proceedings of 2023 Conference on Robot Learning (PMLR 2023) pp. 416\u2013426."},{"key":"e_1_3_2_56_2","unstructured":"Y. J. Ma W. Liang V. Som V. Kumar A. Zhang O. Bastani D. Jayaraman \u201cLIV: Language-image representations and rewards for robotic control \u201d in Proceedings of International Conference on Machine Learning (PMLR 2023) pp. 23301\u201323320."},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","unstructured":"F. Ceola E. Maiettini L. Rosasco L. Natale \u201cA grasp pose is all you need: Learning multi-fingered grasping with deep reinforcement learning from vision and touch \u201d in Proceedings of 2023 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2023) pp. 2985\u20132992.","DOI":"10.1109\/IROS55552.2023.10341776"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.adi8808"},{"key":"e_1_3_2_59_2","doi-asserted-by":"crossref","unstructured":"J. Hansen F. Hogan D. Rivkin D. Meger M. Jenkin G. Dudek \u201cVisuotactile-rl: Learning multimodal manipulation policies with deep reinforcement learning \u201d in Proceedings of 2022 IEEE International Conference on Robotics and Automation (IEEE 2022) pp. 8298\u20138304.","DOI":"10.1109\/ICRA46639.2022.9812019"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1038\/nn991"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.1097011"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.1106138"},{"key":"e_1_3_2_63_2","doi-asserted-by":"crossref","unstructured":"K. He X. Chen S. Xie Y. Li P. Doll\u00e1r R. Girshick \u201cMasked autoencoders are scalable vision learners \u201d in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (IEEE 2022) pp. 16000\u201316009.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_2_64_2","doi-asserted-by":"crossref","unstructured":"J. Devlin M.-W. Chang K. Lee K. Toutanova \u201cBert: Pre-training of deep bidirectional transformers for language understanding \u201d in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (ACL 2019) pp. 4171\u20134186.","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_2_65_2","doi-asserted-by":"crossref","unstructured":"Q. Liu Q. Ye Z. Sun Y. Cui G. Li J. Chen \u201cMasked visual-tactile pre-training for robot manipulation \u201d in Proceedings of 2024 IEEE International Conference on Robotics and Automation (IEEE 2024) pp. 13859\u201313875.","DOI":"10.1109\/ICRA57147.2024.10610933"},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","unstructured":"M. Li H. Yin K. Tahara A. Billard \u201cLearning object-level impedance control for robust grasping and dexterous manipulation \u201d in Proceedings of 2014 IEEE International Conference on Robotics and Automation (IEEE 2014) pp. 6784\u20136791.","DOI":"10.1109\/ICRA.2014.6907861"},{"key":"e_1_3_2_67_2","doi-asserted-by":"crossref","unstructured":"M. Li Y. Bekiroglu D. Kragic A. Billard \u201cLearning of grasp adaptation through experience and tactile sensing \u201d in Proceedings of 2014 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2014) pp. 3339\u20133346.","DOI":"10.1109\/IROS.2014.6943027"},{"key":"e_1_3_2_68_2","doi-asserted-by":"crossref","unstructured":"A. Bernardino M. Henriques N. Hendrich J. Zhang \u201cPrecision grasp synergies for dexterous robotic hands \u201d in Proceedings of 2013 IEEE International Conference on Robotics and Biomimetics (IEEE 2013) pp. 62\u201367.","DOI":"10.1109\/ROBIO.2013.6739436"},{"key":"e_1_3_2_69_2","unstructured":"T. Yu D. Quillen Z. He R. Julian A. Narayan H. Shively A. Bellathur K. Hausman C. Finn S. Levine \u201cMeta-world: A benchmark and evaluation for multi-task and meta reinforcement learning \u201d in Proceedings of 2020 Conference on Robot Learning (PMLR 2020) pp. 1094\u20131100."},{"key":"e_1_3_2_70_2","unstructured":"D. Kalashnikov J. Varley Y. Chebotar B. Swanson R. Jonschkowski C. Finn S. Levine K. Hausman \u201cMt-opt: Continuous multi-task robotic reinforcement learning at scale \u201d in Proceedings of 2021 Conference on Robot Learning (PMLR 2021) pp. 557\u2013575."},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3339515"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abb2174"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-024-00861-3"},{"key":"e_1_3_2_74_2","doi-asserted-by":"crossref","unstructured":"R. Rahmatizadeh P. Abolghasemi L. B\u00f6l\u00f6ni S. Levine \u201cVision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration \u201d in Proceedings of 2018 IEEE International Conference on Robotics and Automation (IEEE 2018) pp. 3758\u20133765.","DOI":"10.1109\/ICRA.2018.8461076"},{"key":"e_1_3_2_75_2","doi-asserted-by":"crossref","unstructured":"S. Haldar Z. Peng L. Pinto \u201cBAKU: An efficient transformer for multi-task policy learning \u201d in Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS 2024) pp. 141208\u2013141239.","DOI":"10.52202\/079017-4484"},{"key":"e_1_3_2_76_2","unstructured":"S. Ross G. Gordon D. Bagnell \u201cA reduction of imitation learning and structured prediction to no-regret online learning \u201d in Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (PMLR 2011) pp. 627\u2013635."},{"key":"e_1_3_2_77_2","first-page":"4","article-title":"Shadow Hand","volume":"1","author":"Sharma D.","year":"2014","unstructured":"D. Sharma, K. Tokas, A. Puri, K. Sharda, Shadow Hand. J. Adv. Res. Appl. Sci. 1, 4\u20137 (2014).","journal-title":"J. Adv. Res. Appl. Sci."},{"key":"e_1_3_2_78_2","unstructured":"Y. J. Ma W. Liang G. Wang D.-A. Huang O. Bastani D. Jayaraman Y. Zhu L. Fan A. Anandkumar \u201cEureka: Human-level reward design via coding large language models \u201d in Proceedings of the Twelfth International Conference on Learning Representations (ICLR 2024) pp. 26516\u201326560."},{"key":"e_1_3_2_79_2","doi-asserted-by":"crossref","unstructured":"J. Wong V. Makoviychuk A. Anandkumar Y. Zhu \u201cOSCAR: Data-driven operational space control for adaptive and robust robot manipulation \u201d in Proceedings of 2022 IEEE International Conference on Robotics and Automation (IEEE 2022) pp. 10519\u201310526.","DOI":"10.1109\/ICRA46639.2022.9811967"},{"key":"e_1_3_2_80_2","unstructured":"J. Liu C. Li D. Delehelle Z. Li F. Chen Skylark0924\/Rofunc: v0.0.2.5 More examples (Zenodo 2023); https:\/\/doi.org\/10.5281\/zenodo.10016946."},{"key":"e_1_3_2_81_2","unstructured":"J. Schulman F. Wolski P. Dhariwal A. Radford O. Klimov Proximal policy optimization algorithms. arXiv:1707.06347 [cs.LG] (2017)."},{"key":"e_1_3_2_82_2","doi-asserted-by":"crossref","unstructured":"J. Tobin R. Fong A. Ray J. Schneider W. Zaremba P. Abbeel \u201cDomain randomization for transferring deep neural networks from simulation to the real world \u201d in Proceedings of 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2017) pp. 23\u201330.","DOI":"10.1109\/IROS.2017.8202133"},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364919887447"},{"key":"e_1_3_2_84_2","unstructured":"V. Makoviychuk L. Wawrzyniak Y. Guo M. Lu K. Storey M. Macklin D. Hoeller N. Rudin A. Allshire A. Handa G. State \u201cIsaac Gym: High performance GPU-based physics simulation for robot learning \u201d in Proceedings of Annual Conference on Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS 2021)."},{"key":"e_1_3_2_85_2","unstructured":"A. X. Chang T. Funkhouser L. Guibas P. Hanrahan Q. Huang Z. Li S. Savarese M. Savva S. Song H. Su J. Xiao L. Yi F. Yu Shapenet: An information-rich 3D model repository. arXiv:1512.03012 [cs.GR] (2015)."},{"key":"e_1_3_2_86_2","doi-asserted-by":"crossref","unstructured":"R. Wang J. Zhang J. Chen Y. Xu P. Li T. Liu H. Wang \u201cDexgraspnet: A large-scale robotic dexterous grasp dataset for general objects based on simulation \u201d in Proceedings of 2023 IEEE International Conference on Robotics and Automation (IEEE 2023) pp. 11359\u201311366.","DOI":"10.1109\/ICRA48891.2023.10160982"},{"key":"e_1_3_2_87_2","doi-asserted-by":"crossref","unstructured":"F. Xiang Y. Qin K. Mo Y. Xia H. Zhu F. Liu M. Liu H. Jiang Y. Yuan H. Wang L. Yi A. X. Chang L. J. Guibas H. Su \u201cSapien: A simulated part-based interactive environment \u201d in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (IEEE 2020) pp. 11097\u201311107.","DOI":"10.1109\/CVPR42600.2020.01111"},{"key":"e_1_3_2_88_2","doi-asserted-by":"crossref","unstructured":"B. Calli A. Singh A. Walsman S. Srinivasa P. Abbeel A. M. Dollar \u201cThe YCB object and model set: Towards common benchmarks for manipulation research \u201d in Proceedings of 2015 International Conference on Advanced Robotics (IEEE 2015) pp. 510\u2013517.","DOI":"10.1109\/ICAR.2015.7251504"},{"key":"e_1_3_2_89_2","unstructured":"D.-A. Clevert T. Unterthiner S. Hochreiter Fast and accurate deep network learning by exponential linear units (ELUs). arXiv:1511.07289 [cs.LG] (2015)."},{"key":"e_1_3_2_90_2","doi-asserted-by":"crossref","unstructured":"K. Shaw A. Agarwal D. Pathak \u201cLEAP hand: Low-cost efficient and anthropomorphic hand for robot learning \u201d in Proceedings of Robotics: Science and Systems (RSS 2023).","DOI":"10.15607\/RSS.2023.XIX.089"},{"key":"e_1_3_2_91_2","doi-asserted-by":"crossref","unstructured":"K. He X. Zhang S. Ren J. Sun \u201cDeep residual learning for image recognition \u201d in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (IEEE 2016) pp. 770\u2013778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_92_2","doi-asserted-by":"publisher","DOI":"10.1177\/02783649241273668"},{"key":"e_1_3_2_93_2","doi-asserted-by":"crossref","unstructured":"Y. Ze G. Zhang K. Zhang C. Hu M. Wang H. Xu \u201c3D diffusion policy: Generalizable visuomotor policy learning via simple 3D representations \u201d in Proceedings of Robotics: Science and Systems (RSS 2024).","DOI":"10.15607\/RSS.2024.XX.067"},{"key":"e_1_3_2_94_2","doi-asserted-by":"publisher","DOI":"10.3390\/s17122762"},{"key":"e_1_3_2_95_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2020.2977257"},{"key":"e_1_3_2_96_2","unstructured":"F. Yang C. Ma J. Zhang J. Zhu W. Yuan A. Owens \u201cTouch and go: Learning from human-collected vision and touch \u201d in Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS 2022) pp. 8081\u20138103."},{"key":"e_1_3_2_97_2","doi-asserted-by":"crossref","unstructured":"J. Kerr H. Huang A. Wilcox R. Hoque Jeffrey Ichnowski R. Calandra K. Goldber \u201cSelf-supervised visuo-tactile pretraining to locate and follow garment features \u201d in Proceedings of Robotics: Science and Systems (RSS 2023).","DOI":"10.15607\/RSS.2023.XIX.018"},{"key":"e_1_3_2_98_2","doi-asserted-by":"crossref","unstructured":"Y. Dou F. Yang Y. Liu A. Loquercio A. Owens \u201cTactile-augmented radiance fields \u201d in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (IEEE 2024) pp. 26529\u201326539.","DOI":"10.1109\/CVPR52733.2024.02505"},{"key":"e_1_3_2_99_2","doi-asserted-by":"crossref","unstructured":"S. Yu K. Lin A. Xiao J. Duan H. Soh \u201cOctopi: Object property reasoning with large tactile-language models \u201d in Proceedings of Robotics: Science and Systems (RSS 2024).","DOI":"10.15607\/RSS.2024.XX.066"},{"key":"e_1_3_2_100_2","unstructured":"A. Radford J. W. Kim C. Hallacy A. Ramesh G. Goh S. Agarwal G. Sastry A. Askell P. Mishkin J. Clark G. Krueger I. Sutskever \u201cLearning transferable visual models from natural language supervision \u201d in Proceedings of International Conference on Machine Learning (PMLR 2021) pp. 8748\u20138763."}],"container-title":["Science Robotics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.science.org\/doi\/pdf\/10.1126\/scirobotics.ady2869","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,28]],"date-time":"2026-01-28T18:58:55Z","timestamp":1769626735000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.science.org\/doi\/10.1126\/scirobotics.ady2869"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,28]]},"references-count":99,"journal-issue":{"issue":"110","published-print":{"date-parts":[[2026,1,28]]}},"alternative-id":["10.1126\/scirobotics.ady2869"],"URL":"https:\/\/doi.org\/10.1126\/scirobotics.ady2869","relation":{},"ISSN":["2470-9476"],"issn-type":[{"value":"2470-9476","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,28]]},"assertion":[{"value":"2025-04-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-24","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"eady2869"}}