{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T16:22:22Z","timestamp":1781108542014,"version":"3.54.1"},"reference-count":49,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2025,7,9]],"date-time":"2025-07-09T00:00:00Z","timestamp":1752019200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Technology Projects of China National Offshore Oil Corporation","award":["KJGG-2024-15-0501"],"award-info":[{"award-number":["KJGG-2024-15-0501"]}]},{"name":"Technology Projects of China National Offshore Oil Corporation","award":["61972353"],"award-info":[{"award-number":["61972353"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["KJGG-2024-15-0501"],"award-info":[{"award-number":["KJGG-2024-15-0501"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61972353"],"award-info":[{"award-number":["61972353"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Human Pose Estimation (HPE) aims to accurately locate the positions of human key points in images or videos. However, the performance of HPE is often significantly reduced in practical application scenarios due to environmental interference. To address this challenge, we propose a ladder side-tuning method for the Vision Transformer (ViT) pre-trained model based on multi-path feature fusion to improve the accuracy of HPE in highly interfering environments. First, we extract the global features, frequency features and multi-scale spatial features through the ViT pre-trained model, the discrete wavelet convolutional network and the atrous spatial pyramid pooling network (ASPP). By comprehensively capturing the information of the human body and the environment, the ability of the model to analyze local details, textures, and spatial information is enhanced. In order to efficiently fuse these features, we devise an adaptive symmetric feature fusion strategy, which dynamically adjusts the intensity of feature fusion according to the similarity among features to achieve the optimal fusion effect. Finally, a multi-graph feature aggregation method is developed. We construct graph structures of different features and deeply explore the subtle differences among the features based on the dual fusion mechanism of points and edges to ensure the information integrity. The experimental results demonstrate that our method achieves 4.3% and 4.2% improvements in the AP metric on the MS COCO dataset and a custom high-interference dataset, respectively, compared with the HRNet. This highlights its superiority for human pose estimation tasks in both general and interfering environments.<\/jats:p>","DOI":"10.3390\/sym17071098","type":"journal-article","created":{"date-parts":[[2025,7,10]],"date-time":"2025-07-10T07:38:27Z","timestamp":1752133107000},"page":"1098","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["A Fine-Tuning Method via Adaptive Symmetric Fusion and Multi-Graph Aggregation for Human Pose Estimation"],"prefix":"10.3390","volume":"17","author":[{"given":"Yinliang","family":"Shi","sequence":"first","affiliation":[{"name":"Digital Intelligence Research Institute, CNOOC Research Institute Ltd., Beijing 100028, China"},{"name":"College of Artificial Intelligence, China University of Petroleum (Beijing), Beijing 102249, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhaonian","family":"Liu","sequence":"additional","affiliation":[{"name":"Digital Intelligence Research Institute, CNOOC Research Institute Ltd., Beijing 100028, China"},{"name":"College of Artificial Intelligence, China University of Petroleum (Beijing), Beijing 102249, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-1348-154X","authenticated-orcid":false,"given":"Bin","family":"Jiang","sequence":"additional","affiliation":[{"name":"Digital Intelligence Research Institute, CNOOC Research Institute Ltd., Beijing 100028, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tianqi","family":"Dai","sequence":"additional","affiliation":[{"name":"Digital Intelligence Research Institute, CNOOC Research Institute Ltd., Beijing 100028, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1801-2507","authenticated-orcid":false,"given":"Yuanfeng","family":"Lian","sequence":"additional","affiliation":[{"name":"College of Artificial Intelligence, China University of Petroleum (Beijing), Beijing 102249, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,7,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Wang, J., Qiu, K., Peng, H., Fu, J., and Zhu, J. (2019, January 21\u201325). Ai coach: Deep human pose estimation and analysis for personalized athletic training assistance. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/3343031.3350910"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Liu, Z., Zhang, H., Chen, Z., Wang, Z., and Ouyang, W. (2020, January 13\u201319). Disentangling and unifying graph convolutions for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00022"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Difini, G.M., Martins, M.G., and Barbosa, J.L.V. (2021, January 5\u201312). Human pose estimation for training assistance: A systematic literature review. Proceedings of the Brazilian Symposium on Multimedia and the Web, Belo Horizonte, Brazil.","DOI":"10.1145\/3470482.3479633"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Gkioxari, G., Arbel\u00e1ez, P., Bourdev, L., and Malik, J. (2013, January 23\u201328). Articulated pose estimation using discriminative armlet classifiers. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.429"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"780","DOI":"10.1109\/34.598236","article-title":"Pfinder: Real-time tracking of the human body","volume":"19","author":"Wren","year":"1997","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201325). Histograms of oriented gradients for human detection. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_7","unstructured":"Ramachandran, P., Parmar, N., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J. (2019). Stand-alone self-attention in vision models. Adv. Neural Inf. Process. Syst., 32\u201344."},{"key":"ref_8","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. Adv. Neural Inf. Process. Syst., 30\u201340."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1023\/B:VISI.0000042934.15159.49","article-title":"Pictorial structures for object recognition","volume":"61","author":"Felzenszwalb","year":"2005","journal-title":"Int. J. Comput. Vis."},{"key":"ref_10","unstructured":"Eichner, M., and Ferrari, V. (2009, January 7\u201310). Better appearance models for pictorial structures. Proceedings of the British Machine Vision Conference, London, UK."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Freifeld, O., Weiss, A., Zuffi, S., and Black, M.J. (2010, January 13\u201318). Contour people: A parameterized model of 2D articulated human shape. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540154"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Yang, Y., and Ramanan, D. (2011, January 20\u201325). Articulated pose estimation with flexible mixtures-of-parts. Proceedings of the CVPR 2011, Providence, RI, USA.","DOI":"10.1109\/CVPR.2011.5995741"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2878","DOI":"10.1109\/TPAMI.2012.261","article-title":"Articulated human detection with flexible mixtures of parts","volume":"35","author":"Yang","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Achilles, F., Ichim, A.-E., Coskun, H., Tombari, F., Noachtar, S., and Navab, N. (2016, January 17\u201321). Patient MoCap: Human pose estimation under blanket occlusion for hospital monitoring applications. Proceedings of the Medical Image Computing and Computer-Assisted Intervention\u2013MICCAI 2016: 19th International Conference, Athens, Greece. Proceedings, Part I 19.","DOI":"10.1007\/978-3-319-46720-7_57"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Toshev, A., and Szegedy, C. (2014, January 23\u201328). Deeppose: Human pose estimation via deep neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.214"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Sun, X., Shang, J., Liang, S., and Wei, Y. (2017, January 22\u201329). Compositional human pose regression. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.284"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Rogez, G., Weinzaepfel, P., and Schmid, C. (2017, January 21\u201326). Lcr-net: Localization-classification-regression for human pose. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.134"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1016\/j.cag.2019.09.002","article-title":"Human pose regression by combining indirect part detection and contextual information","volume":"85","author":"Luvizon","year":"2019","journal-title":"Comput. Graph."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Wei, S.-E., Ramakrishna, V., Kanade, T., and Sheikh, Y. (2016, January 27\u201330). Convolutional pose machines. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.511"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Newell, A., Yang, K., and Deng, J. (2016, January 11\u201314). Stacked hourglass networks for human pose estimation. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Proceedings, Part VIII 14.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Bulat, A., and Tzimiropoulos, G. (2016, January 11\u201314). Human pose estimation via convolutional part heatmap regression. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Proceedings, Part VII 14.","DOI":"10.1007\/978-3-319-46478-7_44"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Sun, K., Xiao, B., Liu, D., and Wang, J. (2019, January 15\u201320). Deep high-resolution representation learning for human pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00584"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Cheng, B., Xiao, B., Wang, J., Shi, H., Huang, T.S., and Zhang, L. (2020, January 13\u201319). Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00543"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Yang, S., Quan, Z., Nie, M., and Yang, W. (2021, January 10\u201317). Transpose: Keypoint localization via transformer. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01159"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Li, Y., Zhang, S., Wang, Z., Yang, S., Yang, W., Xia, S.-T., and Zhou, E. (2021, January 10\u201317). Tokenpose: Learning keypoint tokens for human pose estimation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01112"},{"key":"ref_26","first-page":"7281","article-title":"Hrformer: High-resolution vision transformer for dense predict","volume":"34","author":"Yuan","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"7157","DOI":"10.1109\/TPAMI.2022.3222784","article-title":"Alphapose: Whole-body regional multi-person pose estimation and tracking in real-time","volume":"45","author":"Fang","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Jia, M., Tang, L., Chen, B.-C., Cardie, C., Belongie, S., Hariharan, B., and Lim, S.-N. (2022, January 23\u201327). Visual prompt tuning. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19827-4_41"},{"key":"ref_29","first-page":"1","article-title":"Visual tuning","volume":"56","author":"Yu","year":"2024","journal-title":"ACM Comput. Surv."},{"key":"ref_30","first-page":"34892","article-title":"Visual instruction tuning","volume":"36","author":"Liu","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_31","unstructured":"Zaken, E.B., Ravfogel, S., and Goldberg, Y. (2021). Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. arXiv."},{"key":"ref_32","unstructured":"Bu, Z., Wang, Y.-X., Zha, S., and Karypis, G. (2023). Differentially private bias-term only finetuning of foundation models. arXiv."},{"key":"ref_33","first-page":"3","article-title":"Lora: Low-rank adaptation of large language models","volume":"1","author":"Hu","year":"2022","journal-title":"ICLR"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Valipour, M., Rezagholizadeh, M., Kobyzev, I., and Ghodsi, A. (2022). Dylora: Parameter efficient tuning of pre-trained models using dynamic search-free low-rank adaptation. arXiv.","DOI":"10.18653\/v1\/2023.eacl-main.239"},{"key":"ref_35","unstructured":"Gao, Y., Shi, X., Zhu, Y., Wang, H., Tang, Z., Zhou, X., Li, M., and Metaxas, D.N. (2022). Visual prompt tuning for test-time domain adaptation. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Herzig, R., Abramovich, O., Ben Avraham, E., Arbelle, A., Karlinsky, L., Shamir, A., Darrell, T., and Globerson, A. (2024, January 3\u20138). Promptonomyvit: Multi-task prompt learning improves video transformers using synthetic scene data. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV57701.2024.00666"},{"key":"ref_37","first-page":"6575","article-title":"Feature-proxy transformer for few-shot segmentation","volume":"35","author":"Zhang","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Wang, S., Chang, J., Wang, Z., Li, H., Ouyang, W., and Tian, Q. (2023, January 7\u201314). Fine-grained retrieval prompt tuning. Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA.","DOI":"10.1609\/aaai.v37i2.25363"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Gan, Y., Bai, Y., Lou, Y., Ma, X., Zhang, R., Shi, N., and Luo, L. (2023, January 7\u201314). Decorate the newcomers: Visual domain prompt for continual test time adaptation. Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA.","DOI":"10.1609\/aaai.v37i6.25922"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"4653","DOI":"10.1109\/TCSVT.2023.3327605","article-title":"Pro-tuning: Unified prompt tuning for vision tasks","volume":"34","author":"Nie","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Wang, H., Zhang, T., Yu, M., Sun, J., Ye, W., Wang, C., and Zhang, S. (2020, January 23\u201328). Stacking networks dynamically for image restoration based on the plug-and-play framework. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part XIII 16.","DOI":"10.1007\/978-3-030-58601-0_27"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Ermis, B., Zappella, G., Wistuba, M., Rawal, A., and Archambeau, C. (2022, January 18\u201324). Continual learning with transformers for image classification. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPRW56347.2022.00422"},{"key":"ref_43","first-page":"26462","article-title":"St-adapter: Parameter-efficient image-to-video transfer learning","volume":"35","author":"Pan","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_44","unstructured":"Chen, Z., Duan, Y., Wang, W., He, J., Lu, T., Dai, J., and Qiao, Y. (2022). Vision transformer adapter for dense predictions. arXiv."},{"key":"ref_45","first-page":"16664","article-title":"Adaptformer: Adapting vision transformers for scalable visual recognition","volume":"35","author":"Chen","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Zhang, J.O., Sax, A., Zamir, A., Guibas, L., and Malik, J. (2020, January 23\u201328). Side-tuning: A baseline for network adaptation via additive side networks. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part III 16.","DOI":"10.1007\/978-3-030-58580-8_41"},{"key":"ref_47","first-page":"12991","article-title":"Lst: Ladder side-tuning for parameter and memory efficient transfer learning","volume":"35","author":"Sung","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the Computer vision\u2013ECCV 2014: 13th European conference, Zurich, Switzerland. Proceedings, Part v 13.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"127884","DOI":"10.1016\/j.neucom.2024.127884","article-title":"LMFormer: Lightweight and multi-feature perspective via transformer for human pose estimation","volume":"594","author":"Li","year":"2024","journal-title":"Neurocomputing"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/7\/1098\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:06:52Z","timestamp":1760033212000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/7\/1098"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,9]]},"references-count":49,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2025,7]]}},"alternative-id":["sym17071098"],"URL":"https:\/\/doi.org\/10.3390\/sym17071098","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,9]]}}}