{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T06:20:40Z","timestamp":1768285240406,"version":"3.49.0"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"7","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,7,31]]},"abstract":"<jats:p>Whole-body mesh reconstruction utilizes neural networks to reconstruct the 3D human body, face, and hands, forming a fundamental task in computer vision. It is used to model human action in many practical applications that prioritize upper body action, particularly the hands. However, accurately estimating the 3D mesh parameters of hands in practical applications remains challenging due to severe self-occlusion and high self-similarity in hand action. To address these challenges, the Accurate Hand Modeling in Whole-Body Mesh Reconstruction (AHM-WBMR) is proposed in this article. It mainly consists of two innovative components: the joint-level features progressive matching and refinement and the kinematic features propagation. In the joint-level features progressive matching and refinement, the 3D deformable cross attention and the 3D deformable transformer decoder are proposed to assist in refining hand joint-level features. Further, in the kinematic features propagation, the kinematic-aware topology network is proposed to distinguish and relate different hand joint-level features using three types of kinematic topology structures. We evaluated AHM-WBMR on the UBody and FreiHAND datasets, both of which contain rich hand movements. Compared with the state-of-the-art methods, we achieved improvements of at least 10.2% and 11.1% in hand-related metrics on the two datasets, respectively.<\/jats:p>","DOI":"10.1145\/3743138","type":"journal-article","created":{"date-parts":[[2025,6,6]],"date-time":"2025-06-06T12:48:40Z","timestamp":1749214120000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Accurate Hand Modeling in Whole-Body Mesh Reconstruction Using Joint-Level Features and Kinematic-Aware Topology"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-1458-7063","authenticated-orcid":false,"given":"Fubin","family":"Guo","sequence":"first","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0896-9444","authenticated-orcid":false,"given":"Qi","family":"Wang","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4264-0180","authenticated-orcid":false,"given":"Qingshan","family":"Wang","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-0279-480X","authenticated-orcid":false,"given":"Sheng","family":"Chen","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,7,18]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3271691"},{"key":"e_1_3_1_3_2","article-title":"Towards robust and expressive whole-body human pose and shape estimation","volume":"36","author":"Pang Hui En","year":"2024","unstructured":"Hui En Pang, Zhongang Cai, Lei Yang, Qingyi Tao, Zhonghua Wu, Tianwei Zhang, and Ziwei Liu. 2024. Towards robust and expressive whole-body human pose and shape estimation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"issue":"7","key":"e_1_3_1_4_2","first-page":"2398","article-title":"Hear sign language: A real-time end-to-end sign language recognition system","volume":"21","author":"Wang Zhibo","year":"2020","unstructured":"Zhibo Wang, Tengda Zhao, Jinxin Ma, Hongkai Chen, Kaixin Liu, Huajie Shao, Qian Wang, and Ju Ren. 2020. Hear sign language: A real-time end-to-end sign language recognition system. IEEE Transactions on Mobile Computing 21, 7 (2020), 2398\u20132410.","journal-title":"IEEE Transactions on Mobile Computing"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/THMS.2022.3146787"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2024.3409357"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475463"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00508"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2024.104050"},{"key":"e_1_3_1_10_2","first-page":"1","volume-title":"Proceedings of the 2021 IEEE Conference on Games (CoG)","author":"Zacharatos Haris","unstructured":"Haris Zacharatos, Christos Gatzoulis, Panayiotis Charalambous, and Yiorgos Chrysanthou. 2021. Emotion recognition from 3D motion capture data using deep CNNs. In Proceedings of the 2021 IEEE Conference on Games (CoG). IEEE, 1\u20135."},{"key":"e_1_3_1_11_2","first-page":"768","volume-title":"Proceedings of the 2022 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct)","author":"Cho Youngwug","year":"2022","unstructured":"Youngwug Cho, Myeongul Jung, and Kwanguk Kim. 2022. Emotion and body movement: A comparative study of automatic emotion recognition using body motions. In Proceedings of the 2022 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct). IEEE, 768\u2013771."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01123"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00868"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00622"},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","first-page":"792","DOI":"10.1109\/3DV53792.2021.00088","volume-title":"Proceedings of the 2021 International Conference on 3D Vision (3DV)","author":"Feng Yao","year":"2021","unstructured":"Yao Feng, Vasileios Choutas, Timo Bolkart, Dimitrios Tzionas, and Michael J. Black. 2021. Collaborative regression of expressive bodies using moderation. In Proceedings of the 2021 International Conference on 3D Vision (3DV). IEEE, 792\u2013804."},{"key":"e_1_3_1_16_2","first-page":"14484","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zanfir Andrei","year":"2021","unstructured":"Andrei Zanfir, Eduard Gabriel Bazavan, Mihai Zanfir, William T. Freeman, Rahul Sukthankar, and Cristian Sminchisescu. 2021. Neural descent for visual 3d human pose and shape. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 14484\u201314493."},{"key":"e_1_3_1_17_2","unstructured":"Jiefeng Li Siyuan Bian Chao Xu Zhicun Chen Lixin Yang and Cewu Lu. 2023. Hybrik-x: Hybrid analytical-neural inverse kinematics for whole-body mesh recovery. arXiv:2304.05690. Retrieved from https:\/\/arxiv.org\/abs\/2304.05690"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00478"},{"key":"e_1_3_1_19_2","article-title":"Smpler-x: Scaling up expressive human pose and shape estimation","volume":"36","author":"Cai Zhongang","year":"2024","unstructured":"Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qingping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, et al. 2024. Smpler-x: Scaling up expressive human pose and shape estimation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems, Vol"},{"key":"e_1_3_1_20_2","first-page":"21159","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lin Jing","year":"2023","unstructured":"Jing Lin, Ailing Zeng, Haoqian Wang, Lei Zhang, and Yu Li. 2023. One-stage 3d whole-body mesh recovery with component aware transformer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 21159\u201321168."},{"key":"e_1_3_1_21_2","unstructured":"Thomas N. Kipf and Max Welling. 2016. Semi-supervised classification with graph convolutional networks. arXiv:1609.02907. Retrieved from https:\/\/arxiv.org\/abs\/1609.02907"},{"key":"e_1_3_1_22_2","unstructured":"Ke Sun Yang Zhao Borui Jiang Tianheng Cheng Bin Xiao Dong Liu Yadong Mu Xinggang Wang Wenyu Liu and Jingdong Wang. 2019. High-resolution representations for labeling pixels and regions. arXiv:1904.04514. Retrieved from https:\/\/arxiv.org\/abs\/1904.04514"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00744"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00339"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00234"},{"issue":"5","key":"e_1_3_1_26_2","first-page":"2610","article-title":"Learning 3d human shape and pose from dense body parts","volume":"44","author":"Zhang Hongwen","year":"2020","unstructured":"Hongwen Zhang, Jie Cao, Guo Lu, Wanli Ouyang, and Zhenan Sun. 2020. Learning 3d human shape and pose from dense body parts. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 5 (2020), 2610\u20132627.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00795"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459936"},{"key":"e_1_3_1_29_2","first-page":"20333","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Lizhen","year":"2022","unstructured":"Lizhen Wang, Zhiyuan Chen, Tao Yu, Chenguang Ma, Liang Li, and Yebin Liu. 2022. Faceverse: A fine-grained and detail-controllable 3d face morphable model from a hybrid dataset. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 20333\u201320342."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00244"},{"key":"e_1_3_1_31_2","first-page":"548","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Moon Gyeongsik","year":"2020","unstructured":"Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Kyoung Mu Lee. 2020. Interhand2. 6m: A dataset and baseline for 3d interacting hand pose estimation from a single RGB image. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 548\u2013564."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00278"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1007\/978-3-030-58607-2_2","volume-title":"Proceedings of the 16th European Conference on Computer Vision\u2013ECCV 2020","author":"Choutas Vasileios","year":"2020","unstructured":"Vasileios Choutas, Georgios Pavlakos, Timo Bolkart, Dimitrios Tzionas, and Michael J. Black. 2020. Monocular expressive body regression through body-driven attention. In Proceedings of the 16th European Conference on Computer Vision\u2013ECCV 2020. Springer, 20\u201340."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW54120.2021.00201"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00257"},{"key":"e_1_3_1_36_2","first-page":"1834","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Sun Qingping","year":"2024","unstructured":"Qingping Sun, Yanjun Wang, Ailing Zeng, Wanqi Yin, Chen Wei, Wenjia Wang, Haiyi Mei, Chi-Sing Leung, Ziwei Liu, Lei Yang, et al. 2024. AiOS: All-in-one-stage expressive human pose and shape estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1834\u20131843."},{"key":"e_1_3_1_37_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems, Vol"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547844"},{"key":"e_1_3_1_39_2","unstructured":"Amin Shabani Amir Abdi Lili Meng and Tristan Sylvain. 2022. Scaleformer: Iterative multi-scale refining transformers for time series forecasting. arXiv:2206.04038. Retrieved from https:\/\/arxiv.org\/abs\/2206.04038"},{"key":"e_1_3_1_40_2","unstructured":"Nikita Kitaev \u0141ukasz Kaiser and Anselm Levskaya. 2020. Reformer: The efficient transformer. arXiv:2001.04451. Retrieved from https:\/\/arxiv.org\/abs\/2001.04451"},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","first-page":"738","DOI":"10.1109\/TIP.2023.3349004","article-title":"TTST: A top-k token selective transformer for remote sensing image super-resolution","volume":"33","author":"Xiao Yi","year":"2024","unstructured":"Yi Xiao, Qiangqiang Yuan, Kui Jiang, Jiang He, Chia-Wen Lin, and Liangpei Zhang. 2024. TTST: A top-k token selective transformer for remote sensing image super-resolution. IEEE Transactions on Image Processing: A Publication of the IEEE Signal Processing Society 33 (2024), 738\u2013752.","journal-title":"IEEE Transactions on Image Processing: A Publication of the IEEE Signal Processing Society"},{"key":"e_1_3_1_42_2","doi-asserted-by":"crossref","unstructured":"Kui Jiang Zhongyuan Wang Chen Chen Zheng Wang Laizhong Cui and Chia-Wen Lin. 2022. Magic ELF: Image deraining meets association learning and transformer. arXiv:2207.10455. Retrieved from https:\/\/arxiv.org\/abs\/2207.10455","DOI":"10.1145\/3503161.3547760"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1656"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01595"},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2024.3376819","article-title":"C 2 former: Calibrated and complementary transformer for RGB-infrared object detection","volume":"62","author":"Yuan Maoxun","year":"2024","unstructured":"Maoxun Yuan and Xingxing Wei. 2024. C 2 former: Calibrated and complementary transformer for RGB-infrared object detection. IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1\u201312.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_46_2","first-page":"9558","volume-title":"Proceedings of the 2023 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Tziafas Georgios","year":"2023","unstructured":"Georgios Tziafas and Hamidreza Kasaei. 2023. Early or late fusion matters: Efficient rgb-d fusion in vision transformers for 3d object recognition. In Proceedings of the 2023 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 9558\u20139565."},{"key":"e_1_3_1_47_2","unstructured":"Xizhou Zhu Weijie Su Lewei Lu Bin Li Xiaogang Wang and Jifeng Dai. 2020. Deformable detr: Deformable transformers for end-to-end object detection. arXiv:2010.04159. https:\/\/arxiv.org\/abs\/2010.04159"},{"key":"e_1_3_1_48_2","first-page":"1","article-title":"A deformable attention network for high-resolution remote sensing images semantic segmentation","volume":"60","author":"Zuo Renxiang","year":"2021","unstructured":"Renxiang Zuo and Guangyun Zhang, Rongting Zhang, and Xiuping Jia. 2021. A deformable attention network for high-resolution remote sensing images semantic segmentation. IEEE Transactions on Geoscience and Remote Sensing 60 (2021), 1\u201314.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2023.3291822","article-title":"Deep blind super-resolution for satellite video","volume":"61","author":"Xiao Yi","year":"2023","unstructured":"Yi Xiao, Qiangqiang Yuan, Qiang Zhang, and Liangpei Zhang. 2023. Deep blind super-resolution for satellite video. IEEE Transactions on Geoscience and Remote Sensing 61 (2023), 1\u201316.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_50_2","first-page":"492","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Wang Jian","year":"2020","unstructured":"Jian Wang, Xiang Long, Yuan Gao, Errui Ding, and Shilei Wen. 2020. Graph-pcnn: Two stage human pose estimation with graph pose refinement. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 492\u2013508."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58529-7_29"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01096"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3199201"},{"key":"e_1_3_1_54_2","article-title":"Fan","author":"Yin Yanfang","year":"2023","unstructured":"Yanfang Yin, Ming Liu, Qigang Zhu, and Shuaishuai Zhang, Naseer Ali Hussien, and Yong Fan. 2023. Multi-branch attention graph convolutional networks for 3D human pose estimation. IEEE Transactions on Instrumentation and Measurement 72 (2023), 1\u201312.","journal-title":"IEEE Transactions on Instrumentation and Measurement 72 (2023), 1\u201312."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00354"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"e_1_3_1_57_2","unstructured":"Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in pytorch. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017)."},{"key":"e_1_3_1_58_2","unstructured":"Javier Romero Dimitrios Tzionas and Michael J. Black. 2022. Embodied hands: Modeling and capturing hands and bodies together. arXiv:2201.02610. Retrieved from https:\/\/arxiv.org\/abs\/2201.02610"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.248"},{"key":"e_1_3_1_60_2","first-page":"196","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Jin Sheng","year":"2020","unstructured":"Sheng Jin, Lumin Xu, Jin Xu, Can Wang, Wentao Liu, Chen Qian, Wanli Ouyang, and Ping Luo. 2020. Whole-body human pose estimation in the wild. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 196\u2013214."},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.471"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00090"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01326"},{"key":"e_1_3_1_64_2","first-page":"42","volume-title":"Proceedings of the 2021 International Conference on 3D Vision (3DV)","author":"Joo Hanbyul","year":"2021","unstructured":"Hanbyul Joo, Natalia Neverova, and Andrea Vedaldi. 2021. Exemplar fine-tuning for 3d human model fitting towards in-the-wild 3d human pose estimation. In Proceedings of the 2021 International Conference on 3D Vision (3DV). IEEE, 42\u201352."},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00256"},{"key":"e_1_3_1_66_2","first-page":"11454","article-title":"Smpler-x: Scaling up expressive human pose and shape estimation","volume":"36","author":"Cai Zhongang","year":"2023","unstructured":"Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qingping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, et al. 2023. Smpler-x: Scaling up expressive human pose and shape estimation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36, 11454\u201311468.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems, Vol"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3652583.3658092"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58571-6_45"},{"key":"e_1_3_1_69_2","first-page":"752","volume-title":"Proceedings of the 16th European Conference on Computer Visio (ECCV \u201920)","author":"Moon Gyeongsik","year":"2020","unstructured":"Gyeongsik Moon and Kyoung Mu Lee. 2020. I2l-meshnet: Image-to-lixel prediction network for accurate 3d human pose and mesh estimation from a single RGB image. In Proceedings of the 16th European Conference on Computer Visio (ECCV \u201920). Springer, 752\u2013768."},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00199"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3743138","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,18]],"date-time":"2025-07-18T23:54:16Z","timestamp":1752882856000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3743138"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,18]]},"references-count":69,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2025,7,31]]}},"alternative-id":["10.1145\/3743138"],"URL":"https:\/\/doi.org\/10.1145\/3743138","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,18]]},"assertion":[{"value":"2024-12-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-25","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-18","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}