{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:48:15Z","timestamp":1750308495274,"version":"3.41.0"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2025,4,8]],"date-time":"2025-04-08T00:00:00Z","timestamp":1744070400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2021ZD0111902"],"award-info":[{"award-number":["2021ZD0111902"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62476179, 62172022 and U21B2038"],"award-info":[{"award-number":["62476179, 62172022 and U21B2038"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Key Laboratory of Multimedia and Intelligent Software Technology"},{"name":"Beijing Artificial Intelligence Institute"},{"name":"Faculty of Information Technology"},{"DOI":"10.13039\/501100003444","name":"Beijing University of Technology","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100003444","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>3D human pose estimation (3DHPE) in images aims at estimating 3D joint positions from images. The existing 3DHPE methods usually define the loss function as the error measured by Euclidean distance between the locations of the predicted joints and the ground truth of joints, which confuses two different kinds of errors with obviously different characteristics and should not be processed equally: the error caused by different pose structures and the others. However, The existing human pose representations are not suitable to distinguish these two kinds of errors. In order to tackle this problem, we propose a novel Multi-Anchor Offset Representation (MAOR) for human pose, which locates the position of each joint using its offsets from a group of selected high-precision joints named Multi-Anchor. Making use of MAOR, the pose error related to the distortion of spatial structure can be measured independently from other errors, which is helpful in promoting the accuracy of pose estimation. We then propose a novel MAOR-based coarse-to-fine diffusion model (MAOR-DiffPose) for pose estimation, which optimizes different types of errors of poses step by step. Firstly, a MAOR-based Denoising Process (MDP) is devised to explicitly optimize spatial structures of 3D poses by using MAOR to describe poses and improves the inductive learning ability of MAOR-DiffPose by extracting view-independent features. Secondly, a Joint Coordinate Denoising Process assisted by MAOR (JCDPaM) is devised to expand the input features meaningfully by combining MAOR with the pose representation based on joint coordinates and optimize the joint coordinates of 3D poses with the assistance of MAOR. MAOR-DiffPose realizes accurate 3DHPE by iterating MDP and JCDPaM modules. Comprehensive experimental results on widely used 3DHPE benchmarks Human3.6M and MPI-INF-3DHP show that the proposed method achieves competitive performance compared with the state-of-the-art methods.<\/jats:p>","DOI":"10.1145\/3716387","type":"journal-article","created":{"date-parts":[[2025,2,7]],"date-time":"2025-02-07T15:50:51Z","timestamp":1738943451000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Multi-Anchor Offset Representation Based Coarse-to-Fine Diffusion Model for Human Pose Estimation"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-6088-0706","authenticated-orcid":false,"given":"Qianxing","family":"Li","sequence":"first","affiliation":[{"name":"Beijing University of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7722-7172","authenticated-orcid":false,"given":"Dehui","family":"Kong","sequence":"additional","affiliation":[{"name":"Beijing University of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5583-8260","authenticated-orcid":false,"given":"Jinghua","family":"Li","sequence":"additional","affiliation":[{"name":"Beijing University of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2011-100X","authenticated-orcid":false,"given":"Dongpan","family":"Chen","sequence":"additional","affiliation":[{"name":"Beijing University of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4164-6647","authenticated-orcid":false,"given":"Baocai","family":"Yin","sequence":"additional","affiliation":[{"name":"Beijing University of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,4,8]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19769-7_10"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10095949"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00236"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-69525-5_14"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3669904"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3057267"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00742"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00235"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3177959"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV56688.2023.00292"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3392302"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01253"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3275914"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3279291"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475219"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20068-7_25"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.248"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2017.100"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00603"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3519305"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01274"},{"key":"e_1_3_1_23_2","first-page":"5130","article-title":"Diffusion-based pose refinement and multi-hypothesis generation for 3D human pose estimation","author":"Kang Hongbo","year":"2024","unstructured":"Hongbo Kang, Yong Wang, Mengyuan Liu, Doudou Wu, Peng Liu, Xinlin Yuan, and Wenming Yang. 2024. Diffusion-based pose refinement and multi-hypothesis generation for 3D human pose estimation. In ICASSP \u201924, 5130\u20135134.","journal-title":"ICASSP \u201924"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i1.25213"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01280"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2205.14217"},{"key":"e_1_3_1_27_2","unstructured":"Yi-Lun Liao and Tess Smidt. 2023. Equiformer: Equivariant graph attention transformer for 3D atomistic graphs. In ICLR \u201823 1\u201310. Retrieved from https:\/\/openreview.net\/forum?id=KwmPfARgOTD"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48506.2021.9561605"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58607-2_19"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-69525-5_6"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01117"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00617"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2017.00064"},{"key":"e_1_3_1_34_2","first-page":"1123","article-title":"KTPFormer: Kinematics and trajectory prior knowledge-enhanced transformer for 3D human pose estimation","author":"Peng Jihua","year":"2024","unstructured":"Jihua Peng, Yanghong Zhou, and P. Y. Mok. 2024. KTPFormer: Kinematics and trajectory prior knowledge-enhanced transformer for 3D human pose estimation. In CVPR \u201924, 1123\u20131132.","journal-title":"CVPR \u201924"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2102.09844"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20065-6_27"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01356"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1503.03585"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.02502"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_33"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00464"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCDS.2022.3185146"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2304.14045"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01101"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2928813"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3136613"},{"key":"e_1_3_1_47_2","first-page":"443","article-title":"A global-part-local approach for 3D human pose estimation from single-view images","author":"Xie Yuhong","year":"2024","unstructured":"Yuhong Xie, Chaoqun Hong, Rongsheng Xie, and Jie Li. 2024. A global-part-local approach for 3D human pose estimation from single-view images. In ICAICE \u201924, 443\u2013448.","journal-title":"ICAICE \u201924"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368066"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01584"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00810"},{"key":"e_1_3_1_51_2","first-page":"1672","article-title":"APP: Adaptive pose pooling for 3D human pose estimation from videos","author":"Zhang Jinyan","year":"2024","unstructured":"Jinyan Zhang, Mengyuan Liu, Hong Liu, Guoquan Wang, and Wenhao Li. 2024. APP: Adaptive pose pooling for 3D human pose estimation from videos. In MM \u201924, 1672\u20131681.","journal-title":"MM \u201924"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01288"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3109517"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1117\/1.JEI.30.4.040502"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-023-01786-x"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01979"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2103.10455"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3051173"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01128"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3716387","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3716387","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T18:43:43Z","timestamp":1750272223000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3716387"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,8]]},"references-count":58,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3716387"],"URL":"https:\/\/doi.org\/10.1145\/3716387","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2025,4,8]]},"assertion":[{"value":"2024-06-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}