{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T16:35:20Z","timestamp":1778258120625,"version":"3.51.4"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2025,5,22]],"date-time":"2025-05-22T00:00:00Z","timestamp":1747872000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"funder":[{"name":"Prospective Foundation of Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences","award":["T403461"],"award-info":[{"award-number":["T403461"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Comput. Graph. Interact. Tech."],"published-print":{"date-parts":[[2025,5,22]]},"abstract":"<jats:p>\n            Understanding bimanual hand interactions is essential for realistic 3D pose and shape reconstruction. However, existing methods struggle with occlusions, ambiguous appearances, and computational inefficiencies. To address these challenges, we propose Vision Mamba Bimanual Hand Interaction Network (VM-BHINet), introducing state space models (SSMs) into hand reconstruction to enhance interaction modeling while improving computational efficiency. The core component, Vision Mamba Interaction Feature Extraction Block (VM-IFEBlock), combines SSMs with local and global feature operations, enabling deep understanding of hand interactions. Experiments on the InterHand2.6M dataset show that VM-BHINet reduces Mean per-joint position error (MPJPE) and Mean per-vertex position error (MPVPE) by 2-3\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(\\%\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            , significantly surpassing state-of-the-art methods.\n          <\/jats:p>","DOI":"10.1145\/3728308","type":"journal-article","created":{"date-parts":[[2025,5,23]],"date-time":"2025-05-23T03:52:27Z","timestamp":1747972347000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["VM-BHINet:Vision Mamba Bimanual Hand Interaction Network for 3D Interacting Hand Mesh Recovery From a Single RGB Image"],"prefix":"10.1145","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-9568-7614","authenticated-orcid":false,"given":"Han","family":"Bi","sequence":"first","affiliation":[{"name":"University of Chinese Academy of Sciences, Technology and Engineering Center for Space Utilization, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-6461-8165","authenticated-orcid":false,"given":"Ge","family":"Yu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Space Utilization, Technology and Engineering Center for Space Utilization, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0357-681X","authenticated-orcid":false,"given":"Yu","family":"He","sequence":"additional","affiliation":[{"name":"Key Laboratory of Space Utilization, Technology and Engineering Center for Space Utilization, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-4539-0452","authenticated-orcid":false,"given":"Wenzhuo","family":"Liu","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Faculty of Marine Science and Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-3831-4513","authenticated-orcid":false,"given":"Zijie","family":"Zheng","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences, Technology and Engineering Center for Space Utilization, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,5,22]]},"reference":[{"key":"e_1_3_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33783-3_46"},{"key":"e_1_3_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ECNCT63103.2024.10704418"},{"key":"e_1_3_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01107"},{"key":"e_1_3_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02102"},{"key":"e_1_3_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_7_1","first-page":"722","volume-title":"European Conference on Computer Vision","author":"Di Xinhan","year":"2022","unstructured":"Xinhan Di and Pengqian Yu. 2022. LWA-HAND: Lightweight attention hand for interacting hand reconstruction. In European Conference on Computer Vision. Springer, 722\u2013738."},{"key":"e_1_3_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV53792.2021.00011"},{"key":"e_1_3_1_9_1","article-title":"Sifdrivenet: Speed and image fusion for driving behavior classification network","author":"Gong Yan","year":"2023","unstructured":"Yan Gong, Jianli Lu, Wenzhuo Liu, Zhiwei Li, Xinmin Jiang, Xin Gao, and Xingang Wu. 2023. Sifdrivenet: Speed and image fusion for driving behavior classification network. IEEE Transactions on Computational Social Systems (2023).","journal-title":"IEEE Transactions on Computational Social Systems"},{"key":"e_1_3_1_10_1","article-title":"Multi-modal fusion technology based on vehicle information: A survey","author":"Gong Yan","year":"2022","unstructured":"Yan Gong, Jianli Lu, Jiayi Wu, and Wenzhuo Liu. 2022. Multi-modal fusion technology based on vehicle information: A survey. arXiv preprint arXiv: 2211.06080 (2022).","journal-title":"arXiv preprint arXiv: 2211.06080"},{"key":"e_1_3_1_11_1","article-title":"Mamba: Linear-time sequence modeling with selective state spaces","author":"Gu Albert","year":"2023","unstructured":"Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv: 2312.00752 (2023).","journal-title":"arXiv preprint arXiv: 2312.00752"},{"key":"e_1_3_1_12_1","article-title":"Efficiently modeling long sequences with structured state spaces","author":"Gu Albert","year":"2021","unstructured":"Albert Gu, Karan Goel, and Christopher R\u00e9. 2021a.Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv: 2111.00396 (2021).","journal-title":"arXiv preprint arXiv: 2111.00396"},{"key":"e_1_3_1_13_1","first-page":"572","article-title":"Combining recurrent, convolutional, and continuous-time models with linear state space layers","volume":"34","author":"Gu Albert","year":"2021","unstructured":"Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher R\u00e9. 2021b.Combining recurrent, convolutional, and continuous-time models with linear state space layers. Advances in neural information processing systems 34 (2021), 572\u2013585.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01081"},{"key":"e_1_3_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392452"},{"key":"e_1_3_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_17_1","first-page":"1","article-title":"MFE-SSNet: Multi-Modal Fusion-Based End-to-End Steering Angle and Vehicle Speed Prediction Network","author":"Huang Yi","year":"2024","unstructured":"Yi Huang, Wenzhuo Liu, Yaoyu Li, Lei Yang, Hanqi Jiang, Zhiwei Li, and Jun Li. 2024. MFE-SSNet: Multi-Modal Fusion-Based End-to-End Steering Angle and Vehicle Speed Prediction Network. Automotive Innovation (2024), 1\u201314.","journal-title":"Automotive Innovation"},{"key":"e_1_3_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00854"},{"key":"e_1_3_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01113"},{"key":"e_1_3_1_20_1","doi-asserted-by":"crossref","unstructured":"Rudolph\u00a0Emil Kalman. 1960. A new approach to linear filtering and prediction problems. (1960).","DOI":"10.1115\/1.3662552"},{"key":"e_1_3_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01100"},{"key":"e_1_3_1_22_1","article-title":"Adam: A Method for Stochastic Optimization","author":"Kingma Diederik","year":"2014","unstructured":"Diederik Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Optimization. Computer Science (2014).","journal-title":"Computer Science"},{"key":"e_1_3_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00998"},{"key":"e_1_3_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.438"},{"key":"e_1_3_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00278"},{"key":"e_1_3_1_26_1","article-title":"MIPD: A Multi-sensory Interactive Perception Dataset for Embodied Intelligent Driving","author":"Li Zhiwei","year":"2024","unstructured":"Zhiwei Li, Tingzhen Zhang, Meihua Zhou, Dandan Tang, Pengwei Zhang, Wenzhuo Liu, Qiaoning Yang, Tianyu Shen, Kunfeng Wang, and Huaping Liu. 2024. MIPD: A Multi-sensory Interactive Perception Dataset for Embodied Intelligent Driving. arXiv preprint arXiv: 2411.05881 (2024).","journal-title":"arXiv preprint arXiv: 2411.05881"},{"key":"e_1_3_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV48630.2021.00242"},{"key":"e_1_3_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2023.107575"},{"key":"e_1_3_1_29_1","article-title":"FMDNet: Feature-Attention-Embedding-Based Multimodal-Fusion Driving-Behavior-Classification Network","author":"Liu Wenzhuo","year":"2024","unstructured":"Wenzhuo Liu, Jianli Lu, Junbin Liao, Yicheng Qiao, Guoying Zhang, Jiayin Zhu, Bozhang Xu, and Zhiwei Li. 2024b.FMDNet: Feature-Attention-Embedding-Based Multimodal-Fusion Driving-Behavior-Classification Network. IEEE Transactions on Computational Social Systems (2024).","journal-title":"IEEE Transactions on Computational Social Systems"},{"key":"e_1_3_1_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20068-7_22"},{"key":"e_1_3_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00256"},{"key":"e_1_3_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58565-5_33"},{"key":"e_1_3_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3322958"},{"key":"e_1_3_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01169"},{"key":"e_1_3_1_35_1","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2354910"},{"key":"e_1_3_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW60793.2023.00454"},{"key":"e_1_3_1_37_1","unstructured":"Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in pytorch. (2017)."},{"key":"e_1_3_1_38_1","article-title":"Embodied hands: Modeling and capturing hands and bodies together","author":"Romero Javier","year":"2022","unstructured":"Javier Romero, Dimitrios Tzionas, and Michael\u00a0J Black. 2022. Embodied hands: Modeling and capturing hands and bodies together. arXiv preprint arXiv: 2201.02610 (2022).","journal-title":"arXiv preprint arXiv: 2201.02610"},{"key":"e_1_3_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV53792.2021.00053"},{"key":"e_1_3_1_40_1","article-title":"BSSNet: A Real-Time Semantic Segmentation Network for Road Scenes Inspired from AutoEncoder","author":"Shi Xiaoqiang","year":"2023","unstructured":"Xiaoqiang Shi, Zhenyu Yin, Guangjie Han, Wenzhuo Liu, Li Qin, Yuanguo Bi, and Shurui Li. 2023. BSSNet: A Real-Time Semantic Segmentation Network for Road Scenes Inspired from AutoEncoder. IEEE Transactions on Circuits and Systems for Video Technology (2023).","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/AIIoT58432.2024.10574578"},{"key":"e_1_3_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925965"},{"key":"e_1_3_1_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0895-4"},{"key":"e_1_3_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00062"},{"issue":"6","key":"e_1_3_1_45_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3414685.3417852","article-title":"Rgb2hands: real-time tracking of 3d hand interactions from monocular rgb video","volume":"39","author":"Wang Jiayi","year":"2020","unstructured":"Jiayi Wang, Franziska Mueller, Florian Bernard, Suzanne Sorli, Oleksandr Sotnychenko, Neng Qian, Miguel\u00a0A Otaduy, Dan Casas, and Christian Theobalt. 2020. Rgb2hands: real-time tracking of 3d hand interactions from monocular rgb video. ACM Transactions on Graphics (ToG) 39, 6 (2020), 1\u201316.","journal-title":"ACM Transactions on Graphics (ToG)"},{"key":"e_1_3_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01245"},{"key":"e_1_3_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01116"},{"key":"e_1_3_1_48_1","article-title":"Oblique Convolution: A Novel Convolution Idea for Redefining Lane Detection","author":"Zhang Xinyu","year":"2023","unstructured":"Xinyu Zhang, Yan Gong, Jianli Lu, Zhiwei Li, Shixiang Li, Shu Wang, Wenzhuo Liu, Li Wang, and Jun Li. 2023. Oblique Convolution: A Novel Convolution Idea for Redefining Lane Detection. IEEE Transactions on Intelligent Vehicles (2023).","journal-title":"IEEE Transactions on Intelligent Vehicles"}],"container-title":["Proceedings of the ACM on Computer Graphics and Interactive Techniques"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728308","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728308","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:35Z","timestamp":1750295915000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728308"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,22]]},"references-count":47,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,5,22]]}},"alternative-id":["10.1145\/3728308"],"URL":"https:\/\/doi.org\/10.1145\/3728308","relation":{},"ISSN":["2577-6193"],"issn-type":[{"value":"2577-6193","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,22]]},"assertion":[{"value":"2025-05-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}