{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T13:35:34Z","timestamp":1774877734430,"version":"3.50.1"},"reference-count":34,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T00:00:00Z","timestamp":1774828800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T00:00:00Z","timestamp":1774828800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/100016808","name":"Natural Science Foundation of Xiamen Municipality","doi-asserted-by":"publisher","award":["3502Z202373023"],"award-info":[{"award-number":["3502Z202373023"]}],"id":[{"id":"10.13039\/100016808","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Category\u2010Agnostic Pose Estimation (CAPE) aims to detect keypoints for objects of any category using only a few labeled samples, making it a challenging yet crucial task for general\u2010purpose visual understanding. Existing methods rely on either visual or textual inputs, but the lack of cross\u2010modal interaction limits generalization. Without a unified input representation, solely using visual features hinders consistent prediction of same\u2010type keypoints, while fixed textual representations fail to capture the diverse characteristics of same\u2010type keypoints, leading to coarse and over\u2010generalized outputs. To address these limitations, we propose two multi\u2010modal frameworks that perform visual\u2010textual integration at both the feature and decision levels. Our feature\u2010level module leverages cross\u2010modal attention to align and enhance keypoint representations, while the decision\u2010level fusion adaptively combines modality\u2010specific predictions through a modality\u2010consistency loss. Experiments on the large\u2010scale MP\u2010100 dataset demonstrate that our method surpasses existing baselines in both accuracy and robustness. Under the challenging 1\u2010shot setting, our model achieves a 0.58% improvement in PCK0.2 over the state\u2010of\u2010the\u2010art CAPE method.<\/jats:p>","DOI":"10.1111\/cgf.70368","type":"journal-article","created":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T12:43:02Z","timestamp":1774874582000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Enhancing Robust Category\u2010Agnostic Pose Estimation through Multi\u2010Modal Feature Alignment"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-6596-6989","authenticated-orcid":false,"given":"Boxuan","family":"Li","sequence":"first","affiliation":[{"name":"Pen\u2010Tung Sah Institute of Micro\u2010Nano Science and Technology Xiamen University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6210-4230","authenticated-orcid":false,"given":"Juan","family":"Liu","sequence":"additional","affiliation":[{"name":"Pen\u2010Tung Sah Institute of Micro\u2010Nano Science and Technology Xiamen University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,3,30]]},"reference":[{"key":"e_1_2_7_2_2","unstructured":"CaoZ. Hidalgo MartinezG. SimonT. WeiS. SheikhY. A.: Openpose: Realtime multi-person 2d pose estimation using part affinity fields.IEEE Transactions on Pattern Analysis and Machine Intelligence(2019). 2"},{"key":"e_1_2_7_3_2","unstructured":"ContributorsM.:Openmmlab pose estimation toolbox and benchmark.https:\/\/github.com\/open-mmlab\/mmpose 2020. 6"},{"key":"e_1_2_7_4_2","doi-asserted-by":"crossref","unstructured":"CaronM. TouvronH. MisraI. J\u00e9gouH. MairalJ. BojanowskiP. JoulinA.: Emerging properties in self-supervised vision transformers. InProceedings of the International Conference on Computer Vision (ICCV)(2021). 6","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"e_1_2_7_5_2","first-page":"1126","article-title":"Model-agnostic meta-learning for fast adaptation of deep networks","volume":"70","author":"Finn C.","year":"2017","journal-title":"Proceedings of the 34th International Conference on Machine Learning (ICML)"},{"key":"e_1_2_7_6_2","unstructured":"GravingJ. M. ChaeD. NaikH. LiL. KogerB. CostelloeB. R. CouzinI. D.: Fast and robust animal pose estimation.bioRxiv(2019) 620245. 5"},{"key":"e_1_2_7_7_2","doi-asserted-by":"crossref","unstructured":"GeY. ZhangR. WangX. TangX. LuoP.: Deepfashion2: A versatile benchmark for detection pose estimation segmentation and re-identification of clothing images. In2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2019) pp.5332\u20135340. doi:10.1109\/CVPR.2019.00548. 5","DOI":"10.1109\/CVPR.2019.00548"},{"key":"e_1_2_7_8_2","unstructured":"HirschornO. AvidanS.:Edge weight prediction for category-agnostic pose estimation 2024. URL:https:\/\/arxiv.org\/abs\/2411.16665 arXiv:2411.16665. 2"},{"key":"e_1_2_7_9_2","unstructured":"HirschornO. AvidanS.:A graph-based approach for category-agnostic pose estimation 2024. URL:https:\/\/github.com\/orhir\/PoseAnything arXiv:2311.17891. 2 3"},{"key":"e_1_2_7_10_2","unstructured":"KimJ. ChungH. KimB.-H.: Capellm: Support-free category-agnostic pose estimation with multimodal large language models.arXiv preprint arXiv:2411.06869(2024). URL:https:\/\/github.com\/Junhojuno\/CapeLLM. 3"},{"key":"e_1_2_7_11_2","doi-asserted-by":"crossref","unstructured":"KhanM. H. McDonaghJ. KhanS. ShahabuddinM. AroraA. KhanF. S. ShaoL. TzimiropoulosG.: Animalweb: A large-scale hierarchical dataset of annotated animal faces. In2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2020) pp.6937\u20136946. doi:10.1109\/CVPR42600.2020.00697. 5","DOI":"10.1109\/CVPR42600.2020.00697"},{"key":"e_1_2_7_12_2","doi-asserted-by":"crossref","unstructured":"K\u00f6stingerM. WohlhartP. RothP. M. BischofH.: Annotated facial landmarks in the wild: A large-scale real-world database for facial landmark localization. In2011 IEEE International Conference on Computer Vision Workshops (ICCV Workshops)(2011) pp.2144\u20132151. doi:10.1109\/ICCVW.2011.6130513. 5","DOI":"10.1109\/ICCVW.2011.6130513"},{"key":"e_1_2_7_13_2","doi-asserted-by":"crossref","unstructured":"LiuZ. LinY. CaoY. HuH. WeiY. ZhangZ. LinS. GuoB.: Swin transformer: Hierarchical vision transformer using shifted windows. In2021 IEEE\/CVF International Conference on Computer Vision (ICCV)(2021) pp.9992\u201310002. doi:10.1109\/ICCV48922.2021.00986. 3","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_2_7_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_2_7_15_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnbeh.2020.581154"},{"key":"e_1_2_7_15_3","doi-asserted-by":"crossref","unstructured":"doi:10.3389\/fnbeh.2020.581154. 5","DOI":"10.3389\/fnbeh.2020.581154"},{"key":"e_1_2_7_16_2","unstructured":"LiZ. ZhangX. ZhangY. LongD. XieP. ZhangM.: Towards general text embeddings with multi-stage contrastive learning.arXiv preprint arXiv:2308.03281(2023). 3"},{"key":"e_1_2_7_17_2","doi-asserted-by":"crossref","unstructured":"MajiD. NagoriS. MathewM. PoddarD.: Yolopose: Enhancing yolo for multi person pose estimation using object key-point similarity loss. In2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)(2022) pp.2636\u20132645. doi:10.1109\/CVPRW56347.2022.00297. 2","DOI":"10.1109\/CVPRW56347.2022.00297"},{"key":"e_1_2_7_18_2","unstructured":"RusanovskyM. HirschornO. AvidanS.: Capex: Category-agnostic pose estimation from textual point explanation. InThe Thirteenth International Conference on Learning Representations(2025). URL:https:\/\/github.com\/matanr\/capex"},{"key":"e_1_2_7_18_3","unstructured":"doi:https:\/\/openreview.net\/forum?id=scKAXgonmq. 2 3"},{"key":"e_1_2_7_19_2","unstructured":"RadfordA. KimJ. W. HallacyC. RameshA. GohG. AgarwalS. SastryG. AskellA. MishkinP. ClarkJ.:Learning transferable visual models from natural language supervision. 3"},{"key":"e_1_2_7_20_2","doi-asserted-by":"crossref","unstructured":"ReddyN. D. VoM. NarasimhanS. G.: Carfusion: Combining point tracking and part detection for dynamic 3d reconstruction of vehicles. In2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2018) pp.1906\u20131915. doi:10.1109\/CVPR.2018.00204. 5","DOI":"10.1109\/CVPR.2018.00204"},{"key":"e_1_2_7_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2016.01.002"},{"key":"e_1_2_7_21_3","doi-asserted-by":"crossref","unstructured":"doi:https:\/\/doi.org\/10.1016\/j.imavis.2016.01.002. 5","DOI":"10.1016\/j.imavis.2016.01.002"},{"key":"e_1_2_7_22_2","doi-asserted-by":"crossref","unstructured":"ShiM. HuangZ. MaX. HuX. CaoZ.: Matching is not enough: A two-stage framework for category-agnostic pose estimation. In2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2023) pp.7308\u20137317. URL:https:\/\/github.com\/flyinglynx\/CapeFormer","DOI":"10.1109\/CVPR52729.2023.00706"},{"key":"e_1_2_7_22_3","doi-asserted-by":"crossref","unstructured":"doi:10.1109\/CVPR52729.2023.00706. 2 3","DOI":"10.1109\/CVPR52729.2023.00706"},{"key":"e_1_2_7_23_2","first-page":"4077","article-title":"Prototypical networks for few-shot learning","volume":"30","author":"Snell J.","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_7_24_2","doi-asserted-by":"crossref","unstructured":"ToshevA. SzegedyC.: Deeppose: Human pose estimation via deep neural networks. In2014 IEEE Conference on Computer Vision and Pattern Recognition(2014) pp.1653\u20131660. doi:10.1109\/CVPR.2014.214. 2","DOI":"10.1109\/CVPR.2014.214"},{"key":"e_1_2_7_25_2","unstructured":"WahC. BransonS. WelinderP. PeronaP. BelongieS.:The Caltech-UCSD Birds-200-2011 Dataset. Tech. Rep. CNS-TR-2011-001 California Institute of Technology 2011. 5"},{"key":"e_1_2_7_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2018.2879980"},{"key":"e_1_2_7_27_2","doi-asserted-by":"crossref","unstructured":"XuL. JinS. ZengW. LiuW. QianC. OuyangW. LuoP. WangX.: Pose for Everything: Towards Category-Agnostic Pose Estimation.Computer Vision \u2013 ECCV 2022 13671(2022) 398\u2013416. Code available athttps:\/\/github.com\/luminxu\/Pose-for-Everything. URL:https:\/\/github.com\/luminxu\/Pose-for-Everything. 2 3","DOI":"10.1007\/978-3-031-20068-7_23"},{"key":"e_1_2_7_28_2","unstructured":"YuH. XuY. ZhangJ. ZhaoW. GuanZ. TaoD.: Ap-10k: A benchmark for animal pose estimation in the wild. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)(2021). 5"},{"key":"e_1_2_7_29_2","first-page":"17301","article-title":"Apt-36k: A large-scale benchmark for animal pose estimation and tracking","volume":"35","author":"Yang Y.","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_7_30_2","unstructured":"ZhangY. HuangX. MaJ. LiZ. LuoZ. XieY. QinY. LuoT. LiY. LiuS. et al.: Recognize anything: A strong image tagging model.arXiv preprint arXiv:2306.03514(2023). 1"},{"key":"e_1_2_7_31_2","unstructured":"ZouX. YangJ. ZhangH. LiF. LiL. WangJ. WangL. GaoJ. LeeY. J.: Segment everything everywhere all at once. InAdvances in Neural Information Processing Systems 36 (NeurIPS 2023)(2023). 37th Conference on Neural Information Processing Systems (NeurIPS) New Orleans LA Dec 10\u201316 2023. 1"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70368","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70368","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70368","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T12:43:08Z","timestamp":1774874588000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70368"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,30]]},"references-count":34,"alternative-id":["10.1111\/cgf.70368"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70368","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,30]]},"assertion":[{"value":"2026-03-30","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70368"}}