{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,19]],"date-time":"2025-12-19T05:15:06Z","timestamp":1766121306981,"version":"3.48.0"},"reference-count":66,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,12,1]],"date-time":"2025-12-01T00:00:00Z","timestamp":1764547200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,12,19]],"date-time":"2025-12-19T00:00:00Z","timestamp":1766102400000},"content-version":"vor","delay-in-days":18,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100014718","name":"Innovative Research Group Project of the National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62401447"],"award-info":[{"award-number":["62401447"]}],"id":[{"id":"10.13039\/100014718","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100015401","name":"Key Research and Development Projects of Shaanxi Province","doi-asserted-by":"publisher","award":["2024GX-YBXM-051"],"award-info":[{"award-number":["2024GX-YBXM-051"]}],"id":[{"id":"10.13039\/501100015401","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100015401","name":"Key Research and Development Projects of Shaanxi Province","doi-asserted-by":"publisher","award":["2024CY2-GJHX08"],"award-info":[{"award-number":["2024CY2-GJHX08"]}],"id":[{"id":"10.13039\/501100015401","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"published-print":{"date-parts":[[2025,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The aim of human-object interaction (HOI) detection is to identify the triplets consisting of a human, a verb, and an object. Although existing methods leverage vision-language models (e.g., CLIP) to transfer textual information for unseen compositions, they often fail to capture the fine-grained visual cues that are essential for complex interactions, such as spatial configurations and object affordances. In this paper, we introduce visual guidance as an alternative approach to achieving the desired outcome. We define a new visual-guided HOI detection task for the first time, aiming at detecting unseen HOI categories using a small number of guidance examples. To support this new task, we have constructed a new benchmark dataset, which contains one base set and four novel sets, taking into account the peculiarities of HOI. Then, we propose a VG-HOI model with progressive guidance, query reconstruction, and a conditional uncoupling decoder to supplement common HOI knowledge and task-specific cues to improve the generalization capability of our model. Besides, we explore a new guidance sampling strategy \u2014 disentangled guidance \u2014 for real-world scenarios. Our in-depth analysis of the experimental results shows that the proposed model can improve the ability to generalize when detecting visual-guided HOI.<\/jats:p>","DOI":"10.1007\/s44267-025-00102-0","type":"journal-article","created":{"date-parts":[[2025,12,19]],"date-time":"2025-12-19T03:39:02Z","timestamp":1766115542000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Visual-guided human-object interaction detection"],"prefix":"10.1007","volume":"3","author":[{"given":"Fang","family":"Nan","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ni","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nian","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haochen","family":"Han","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9266-4685","authenticated-orcid":false,"given":"Binglu","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,12,19]]},"reference":[{"key":"102_CR1","first-page":"8359","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"G. Gkioxari","year":"2018","unstructured":"Gkioxari, G., Girshick, R., Doll\u00e1r, P., & He, K. (2018). Detecting and recognizing human-object interactions. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 8359\u20138367). Piscataway: IEEE."},{"key":"102_CR2","first-page":"381","volume-title":"Proceedings of the IEEE winter conference on applications of computer vision","author":"Y.-W. Chao","year":"2018","unstructured":"Chao, Y.-W., Liu, Y., Liu, X., Zeng, H., & Deng, J. (2018). Learning to detect human-object interactions. In Proceedings of the IEEE winter conference on applications of computer vision (pp. 381\u2013389). Piscataway: IEEE."},{"key":"102_CR3","first-page":"20104","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"F. Zhang","year":"2022","unstructured":"Zhang, F., Campbell, D., & Gould, S. (2022). Efficient two-stage detection of human-object interactions with a novel unary-pairwise transformer. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 20104\u201320112). Piscataway: IEEE."},{"key":"102_CR4","doi-asserted-by":"publisher","first-page":"3853","DOI":"10.1109\/TCSVT.2021.3119892","volume":"32","author":"D. Yang","year":"2021","unstructured":"Yang, D., Zou, Y., Zhang, C., Cao, M., & Chen, J. (2021). RR-Net: relation reasoning for end-to-end human-object interaction detection. IEEE Transactions on Circuits and Systems for Video Technology, 32, 3853\u20133865.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"102_CR5","first-page":"482","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Liao","year":"2020","unstructured":"Liao, Y., Liu, S., Wang, F., Chen, Y., Qian, C., & Feng, J. (2020). PPDM: parallel point detection and matching for real-time human-object interaction detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 482\u2013490). Piscataway: IEEE."},{"key":"102_CR6","first-page":"213","volume-title":"Proceedings of the 16th European conference on computer vision","author":"N. Carion","year":"2020","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-end object detection with transformers. In A. Vedaldi, H. Bischof, T. Brox, & J. Frahm (Eds.), Proceedings of the 16th European conference on computer vision (pp. 213\u2013229). Cham: Springer."},{"key":"102_CR7","first-page":"10410","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"M. Tamura","year":"2021","unstructured":"Tamura, M., Ohashi, H., & Yoshinaga, T. (2021). QPIC: query-based pairwise human-object interaction detection with image-wide contextual information. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 10410\u201310419). Piscataway: IEEE."},{"key":"102_CR8","first-page":"11825","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"C. Zou","year":"2021","unstructured":"Zou, C., Wang, B., Hu, Y., Liu, J., Wu, Q., Zhao, Y., Li, B., Zhang, C., Zhang, C., Wei, Y., et al. (2021). End-to-end human object interaction detection with HOI transformer. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11825\u201311834). Piscataway: IEEE."},{"key":"102_CR9","first-page":"9004","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"M. Chen","year":"2021","unstructured":"Chen, M., Liao, Y., Liu, S., Chen, Z., Wang, F., & Qian, C. (2021). Reformulating HOI detection as adaptive set prediction. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 9004\u20139013). Piscataway: IEEE."},{"key":"102_CR10","first-page":"74","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"B. Kim","year":"2021","unstructured":"Kim, B., Lee, J., Kang, J., Kim, E.-S., & Kim, H. (2021). HOTR: end-to-end human-object interaction detection with transformers. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 74\u201383). Piscataway: IEEE."},{"key":"102_CR11","first-page":"1","volume-title":"Proceedings of the 18th international conference on machine vision and applications","author":"J. Chen","year":"2023","unstructured":"Chen, J., & Yanai, K. (2023). QAHOI: query-based anchors for human-object interaction detection. In Proceedings of the 18th international conference on machine vision and applications. (pp. 1\u20135). Piscataway: IEEE."},{"key":"102_CR12","doi-asserted-by":"publisher","first-page":"1827","DOI":"10.1109\/TCSVT.2022.3216663","volume":"33","author":"Y. Cheng","year":"2022","unstructured":"Cheng, Y., Wang, Z., Zhan, W., & Duan, H. (2022). Multi-scale human-object interaction detector. IEEE Transactions on Circuits and Systems for Video Technology, 33, 1827\u20131838.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"102_CR13","first-page":"87","volume-title":"Proceedings of the 17th European conference on computer vision","author":"D. Tu","year":"2022","unstructured":"Tu, D., Min, X., Duan, H., Guo, G., Zhai, G., & Shen, W. (2022). Iwin: human-object interaction detection via transformer with irregular windows. In S. Avidan, G. J. Brostow, M. Ciss\u00e9, G. M. Farinella, & T. Hassner (Eds.), Proceedings of the 17th European conference on computer vision (pp. 87\u2013103). Cham: Springer."},{"key":"102_CR14","first-page":"1019","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"J. Park","year":"2022","unstructured":"Park, J., Lee, S., Heo, H., Choi, H., & Kim, H. (2022). Consistency learning via decoding path augmentation for transformers in human object interaction detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 1019\u20131028). Piscataway: IEEE."},{"key":"102_CR15","first-page":"2925","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"S. Kim","year":"2023","unstructured":"Kim, S., Jung, D., & Cho, M. (2023). Relational context learning for human-object interaction detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 2925\u20132934). Piscataway: IEEE."},{"key":"102_CR16","doi-asserted-by":"publisher","first-page":"5603","DOI":"10.1109\/TCSVT.2024.3358952","volume":"34","author":"Y. Wang","year":"2024","unstructured":"Wang, Y., Liu, Q., & Lei, Y. (2024). TED-Net: dispersal attention for perceiving interaction region in indirectly-contact HOI detection. IEEE Transactions on Circuits and Systems for Video Technology, 34, 5603\u20135615.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"102_CR17","doi-asserted-by":"publisher","first-page":"9760","DOI":"10.1109\/TCSVT.2024.3402247","volume":"34","author":"W. Ren","year":"2024","unstructured":"Ren, W., Luo, J., Jiang, W., Qu, L., Han, Z., Tian, J., & Liu, H. (2024). Learning self- and cross-triplet context clues for human-object interaction detection. IEEE Transactions on Circuits and Systems for Video Technology, 34, 9760\u20139773.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"102_CR18","first-page":"3931","volume-title":"Proceedings of the 39th AAAI conference on artificial intelligence","author":"M. Jia","year":"2025","unstructured":"Jia, M., Zhao, L., Li, G., & Zheng, Y. (2025). ContextHOI: spatial context learning for human-object interaction detection. In T. Walsh, J. Shah, & Z. Kolter (Eds.), Proceedings of the 39th AAAI conference on artificial intelligence (pp. 3931\u20133939). Palo Alto: AAAI Press."},{"key":"102_CR19","first-page":"20126","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Y. Hu","year":"2025","unstructured":"Hu, Y., Ding, C., Sun, C., Huang, S., & Xu, X. (2025). Bilateral collaboration with large vision-language models for open vocabulary human-object interaction detection. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 20126\u201320136). Piscataway: IEEE."},{"key":"102_CR20","first-page":"8748","volume-title":"Proceedings of the international conference on machine learning","author":"A. Radford","year":"2025","unstructured":"Radford, A., Kim, J., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2025). Learning transferable visual models from natural language supervision. In M. Meila, & T. Zhang (Eds.), Proceedings of the international conference on machine learning (pp. 8748\u20138763). Retrieved November 25, 2025, from https:\/\/proceedings.mlr.press\/v139\/radford21a\/radford21a.pdf."},{"key":"102_CR21","first-page":"1130","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"X. Wang","year":"2023","unstructured":"Wang, X., Zhang, X., Cao, Y., Wang, W., Shen, C., & Huang, T. (2023). SegGPT: towards segmenting everything in context. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 1130\u20131140). Piscataway: IEEE."},{"key":"102_CR22","first-page":"6830","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"X. Wang","year":"2023","unstructured":"Wang, X., Wang, W., Cao, Y., Shen, C., & Huang, T. (2023). Images speak in images: a generalist painter for in-context visual learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 6830\u20136839). Piscataway: IEEE."},{"key":"102_CR23","first-page":"91","volume-title":"Proceedings of the 29th international conference on neural information processing systems","author":"S. Ren","year":"2015","unstructured":"Ren, S., He, K., Girshick, R., & Sun, J. (2015). Faster R-CNN: towards real-time object detection with region proposal networks. In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, & R. Garnett (Eds.), Proceedings of the 29th international conference on neural information processing systems (pp. 91\u201399). Red Hook: Curran Associates."},{"key":"102_CR24","first-page":"17209","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"A. Zhang","year":"2021","unstructured":"Zhang, A., Liao, Y., Liu, S., Lu, M., Wang, Y., Gao, C., & Li, X. (2021). Mining the benefits of two-stage and one-stage HOI detection. In M. Ranzato, A. Beygelzimer, Y. N. Dauphin, P. Liang, & J. W. Vaughan (Eds.), Proceedings of the 35th international conference on neural information processing systems (pp. 17209\u201317220). Red Hook: Curran Associates."},{"key":"102_CR25","first-page":"19568","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"D. Zhou","year":"2022","unstructured":"Zhou, D., Liu, Z., Wang, J., Wang, L., Hu, T., Ding, E., & Wang, J. (2022). Human-object interaction detection via disentangled transformer. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 19568\u201319577). Piscataway: IEEE."},{"key":"102_CR26","first-page":"19548","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Zhang","year":"2022","unstructured":"Zhang, Y., Pan, Y., Yao, T., Huang, R., Mei, T., & Chen, C.-W. (2022). Exploring structure-aware transformer over interaction proposals for human-object interaction detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 19548\u201319557). Piscataway: IEEE."},{"key":"102_CR27","first-page":"19538","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"L. Dong","year":"2022","unstructured":"Dong, L., Li, Z., Xu, K., Zhang, Z., Yan, L., Zhong, S., & Zou, X. (2022). Category-aware transformer network for better human-object interaction detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 19538\u201319547). Piscataway: IEEE."},{"key":"102_CR28","first-page":"5353","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"A. Iftekhar","year":"2022","unstructured":"Iftekhar, A., Chen, H., Kundu, K., Li, X., Tighe, J., & Modolo, D. (2022). What to look at and where: semantic and spatial refined transformer for detecting human-object interactions. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 5353\u20135363). Piscataway: IEEE."},{"key":"102_CR29","first-page":"19558","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"X. Qu","year":"2022","unstructured":"Qu, X., Ding, C., Li, X., Zhong, X., & Tao, D. (2022). Distillation using oracle queries for transformer-based human-object interaction detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 19558\u201319567). Piscataway: IEEE."},{"key":"102_CR30","first-page":"20123","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Liao","year":"2022","unstructured":"Liao, Y., Zhang, A., Lu, M., Wang, Y., Li, X., & Liu, S. (2022). GEN-VLKT: simplify association and enhance interaction understanding for HOI detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 20123\u201320132). Piscataway: IEEE."},{"key":"102_CR31","first-page":"19578","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"B. Kim","year":"2022","unstructured":"Kim, B., Mun, J., On, K.-W., Shin, M., Lee, J., & Kim, E.-S. (2022). MSTR: multi-scale transformer for end-to-end human-object interaction detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 19578\u201319587). Piscataway: IEEE."},{"key":"102_CR32","first-page":"939","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"S. Wang","year":"2022","unstructured":"Wang, S., Duan, Y., Ding, H., Tan, Y.-P., Yap, K.-H., & Yuan, J. (2022). Learning transferable human-object interaction detector with natural language supervision. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 939\u2013948). Piscataway: IEEE."},{"key":"102_CR33","first-page":"23507","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"S. Ning","year":"2023","unstructured":"Ning, S., Qiu, L., Liu, Y., & He, X. (2023). HOICLIP: efficient knowledge transfer for HOI detection with vision-language models. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 23507\u201323517). Piscataway: IEEE."},{"key":"102_CR34","first-page":"28212","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"J. Luo","year":"2024","unstructured":"Luo, J., Ren, W., Jiang, W., Chen, X., Wang, Q., Han, Z., & Liu, H. (2024). Discovering syntactic interaction clues for human-object interaction detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 28212\u201328222). Piscataway: IEEE."},{"key":"102_CR35","first-page":"23492","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Y. Cao","year":"2023","unstructured":"Cao, Y., Tang, Q., Yang, F., Su, X., You, S., Lu, X., & Xu, C. (2023). Re-mine, learn and reason: exploring the cross-modal semantic correlations for language-guided HOI detection. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 23492\u201323503). Piscataway: IEEE."},{"key":"102_CR36","first-page":"6480","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"T. Lei","year":"2023","unstructured":"Lei, T., Caba, F., Chen, Q., Jin, H., Peng, Y., & Liu, Y. (2023). Efficient adaptive human-object interaction detection with concept-guided memory. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 6480\u20136490). Piscataway: IEEE."},{"key":"102_CR37","first-page":"45895","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"Y. Mao","year":"2023","unstructured":"Mao, Y., Deng, J., Zhou, W., Li, L., Fang, Y., & Li, H. (2023). CLIP4HOI: towards adapting clip for practical zero-shot HOI detection. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 45895\u201345906). Red Hook: Curran Associates."},{"key":"102_CR38","first-page":"27970","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"G. Wang","year":"2024","unstructured":"Wang, G., Guo, Y., Xu, Z., & Kankanhalli, M. (2024). Bilateral adaptation for human-object interaction detection with occlusion-robustness. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 27970\u201327980). Piscataway: IEEE."},{"key":"102_CR39","first-page":"37416","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"H. Yuan","year":"2022","unstructured":"Yuan, H., Jiang, J., Albanie, S., Feng, T., Huang, Z., Ni, D., & Tang, M. (2022). RLIP: relational language-image pre-training for human-object interaction detection. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Proceedings of the 36th international conference on neural information processing systems (pp. 37416\u201337431). Red Hook: Curran Associates."},{"key":"102_CR40","first-page":"21649","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"H. Yuan","year":"2023","unstructured":"Yuan, H., Zhang, S., Wang, X., Albanie, S., Pan, Y., Feng, T., Jiang, J., Ni, D., Zhang, Y., & Zhao, D. (2023). RLIPv2: fast scaling of relational language-image pre-training. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 21649\u201321661). Piscataway: IEEE."},{"key":"102_CR41","first-page":"19392","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"S. Zheng","year":"2023","unstructured":"Zheng, S., Xu, B., & Jin, Q. (2023). Open-category human-object interaction pre-training via language modeling framework. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 19392\u201319402). Piscataway: IEEE."},{"key":"102_CR42","first-page":"8420","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"B. Kang","year":"2019","unstructured":"Kang, B., Liu, Z., Wang, X., Yu, F., Feng, J., & Darrell, T. (2019). Few-shot object detection via feature reweighting. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 8420\u20138429). Piscataway: IEEE."},{"key":"102_CR43","first-page":"1844","volume-title":"Proceedings of the 37th AAAI conference on artificial intelligence","author":"X. Lu","year":"2023","unstructured":"Lu, X., Diao, W., Mao, Y., Li, J., Wang, P., Sun, X., & Fu, K. (2023). Breaking immutable: information-coupled prototype elaboration for few-shot object detection. In B. Williams, Y. Chen, & J. Neville (Eds.), Proceedings of the 37th AAAI conference on artificial intelligence (pp. 1844\u20131852). Palo Alto: AAAI Press."},{"key":"102_CR44","first-page":"12832","volume":"45","author":"G. Zhang","year":"2022","unstructured":"Zhang, G., Luo, Z., Cui, K., Lu, S., & Xing, E. (2022). Meta-DETR: image-level few-shot detection with inter-class correlation exploitation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45, 12832\u201312843.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"102_CR45","first-page":"11793","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"A. Bulat","year":"2023","unstructured":"Bulat, A., Guerrero, R., Martinez, B., & Tzimiropoulos, G. (2023). FS-DETR: few-shot DEtection TRansformer with prompting and without re-training. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 11793\u201311802). Piscataway: IEEE."},{"key":"102_CR46","doi-asserted-by":"publisher","first-page":"9221","DOI":"10.1109\/TPAMI.2024.3421340","volume":"46","author":"Y. Han","year":"2024","unstructured":"Han, Y., Zhang, J., Xue, Z., Xu, C., Shen, X., Wang, Y., Wang, C., Liu, Y., & Li, X. (2024). Reference twice: a simple and unified baseline for few-shot instance segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46, 9221\u20139238.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"102_CR47","doi-asserted-by":"publisher","first-page":"1050","DOI":"10.1109\/TPAMI.2020.3013717","volume":"44","author":"Z. Tian","year":"2020","unstructured":"Tian, Z., Zhao, H., Shu, M., Yang, Z., Li, R., & Jia, J. (2020). Prior guided feature enrichment network for few-shot segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44, 1050\u20131065.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"102_CR48","first-page":"37484","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"Y. Sun","year":"2022","unstructured":"Sun, Y., Chen, Q., He, X., Wang, J., Feng, H., Han, J., Ding, E., Cheng, J., Li, Z., & Wang, J. (2022). Singular value fine-tuning: few-shot segmentation requires few-parameters fine-tuning. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Proceedings of the 36th international conference on neural information processing systems (pp. 37484\u201337496). Red Hook: Curran Associates."},{"key":"102_CR49","unstructured":"Zhang, R., Jiang, Z., Guo, Z., Yan, S., Pan, J., Ma, X., Dong, H., Gao, P., & Li, H. (2023). Personalize segment anything model with one shot (pp. 13754\u201313783). arXiv preprint. arXiv:2305.03048."},{"key":"102_CR50","first-page":"4015","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"A. Kirillov","year":"2023","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A., Lo, W.-Y., et al. (2023). Segment anything. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 4015\u20134026). Piscataway: IEEE."},{"key":"102_CR51","first-page":"23565","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Sun","year":"2024","unstructured":"Sun, Y., Chen, J., Zhang, S., Zhang, X., Chen, Q., Zhang, G., Ding, E., Wang, J., & Li, Z. (2024). VRP-SAM: SAM with visual reference prompt. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 23565\u201323574). Piscataway: IEEE."},{"key":"102_CR52","first-page":"38020","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"Y. Liu","year":"2022","unstructured":"Liu, Y., Liu, N., Yao, X., & Han, J. (2022). Intermediate prototype mining transformer for few-shot semantic segmentation. In Proceedings of the 36th international conference on neural information processing systems (pp. 38020\u201338031). Red Hook: Curran Associates."},{"key":"102_CR53","first-page":"11573","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Liu","year":"2022","unstructured":"Liu, Y., Liu, N., Cao, Q., Yao, X., Han, J., & Shao, L. (2022). Learning non-target knowledge for few-shot semantic segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11573\u201311582). Piscataway: IEEE."},{"key":"102_CR54","doi-asserted-by":"crossref","unstructured":"Shaban, A., Bansal, S., Liu, Z., Essa, I., & Boots, B. (2017). One-shot learning for semantic segmentation. arXiv preprint. arXiv:1709.03410.","DOI":"10.5244\/C.31.167"},{"key":"102_CR55","first-page":"11085","volume-title":"Proceedings of the 34th AAAI conference on artificial intelligence","author":"Z. Ji","year":"2020","unstructured":"Ji, Z., Liu, X., Pang, Y., & Li, X. (2020). SGAP-Net: semantic-guided attentive prototypes network for few-shot human-object interaction recognition. In Proceedings of the 34th AAAI conference on artificial intelligence (pp. 11085\u201311092). Palo Alto: AAAI Press."},{"key":"102_CR56","doi-asserted-by":"publisher","first-page":"1648","DOI":"10.1109\/TIP.2020.3046861","volume":"30","author":"Z. Ji","year":"2020","unstructured":"Ji, Z., Liu, X., Pang, Y., Ouyang, W., & Li, X. (2020). Few-shot human-object interaction recognition with semantic-guided attentive prototypes network. IEEE Transactions on Image Processing, 30, 1648\u20131661.","journal-title":"IEEE Transactions on Image Processing"},{"key":"102_CR57","first-page":"10460","volume-title":"Proceedings of the 34th AAAI conference on artificial intelligence","author":"A. Bansal","year":"2020","unstructured":"Bansal, A., Rambhatla, S., Shrivastava, A., & Chellappa, R. (2020). Detecting human-object interactions via functional generalization. In Proceedings of the 34th AAAI conference on artificial intelligence (pp. 10460\u201310469). Palo Alto: AAAI Press."},{"key":"102_CR58","doi-asserted-by":"publisher","first-page":"4235","DOI":"10.1145\/3394171.3413600","volume-title":"Proceedings of the 28th ACM international conference on multimedia","author":"Y. Liu","year":"2020","unstructured":"Liu, Y., Yuan, J., & Chen, C. (2020). ConsNet: learning consistency graph for zero-shot human-object interaction detection. In C. Chen, R. Cucchiara, X.-S. Hua, G.-J. Qi, E. Ricci, Z. Zhang, & R. Zimmermann (Eds.), Proceedings of the 28th ACM international conference on multimedia (pp. 4235\u20134243). New York: ACM."},{"key":"102_CR59","first-page":"584","volume-title":"Proceedings of the 16th European conference on computer vision","author":"Z. Hou","year":"2020","unstructured":"Hou, Z., Peng, X., Qiao, Y., & Tao, D. (2020). Visual compositional learning for human-object interaction detection. In C. W. Chen, R. Cucchiara, X. Hua, G. Qi, E. Ricci, Z. Zhang, & R. Zimmermann (Eds.), Proceedings of the 16th European conference on computer vision (pp. 584\u2013600). Cham: Springer."},{"key":"102_CR60","first-page":"2961","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"K. He","year":"2017","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., & Girshick, R. (2017). Mask R-CNN. In Proceedings of the IEEE international conference on computer vision (pp. 2961\u20132969). Piscataway: IEEE."},{"key":"102_CR61","first-page":"21981","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"C. Doersch","year":"2020","unstructured":"Doersch, C., Gupta, A., & Zisserman, A. (2020). CrossTransformers: spatially-aware few-shot transfer. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, & H. Lin (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 21981\u201321993). Red Hook: Curran Associates."},{"key":"102_CR62","first-page":"2567","volume-title":"Proceedings of the 36th AAAI conference on artificial intelligence","author":"Y. Wang","year":"2022","unstructured":"Wang, Y., Zhang, X., Yang, T., & Sun, J. (2022). Anchor DETR: query design for transformer-based detector. In Proceedings of the 36th AAAI conference on artificial intelligence (pp. 2567\u20132575). Palo Alto: AAAI Press."},{"key":"102_CR63","first-page":"3651","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"D. Meng","year":"2021","unstructured":"Meng, D., Chen, X., Fan, Z., Zeng, G., Li, H., Yuan, Y., Sun, L., & Wang, J. (2021). Conditional DETR for fast training convergence. In Proceedings of the IEEE international conference on computer vision (pp. 3651\u20133660). Piscataway: IEEE."},{"key":"102_CR64","unstructured":"Liu, S., Li, F., Zhang, H., Yang, X., Qi, X., Su, H., Zhu, J., & Zhang, L. (2022). DAB-DETR: dynamic anchor boxes are better queries for DETR. arXiv preprint. arXiv:2201.12329."},{"key":"102_CR65","first-page":"1017","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Y.-W. Chao","year":"2015","unstructured":"Chao, Y.-W., Wang, Z., He, Y., Wang, J., & Deng, J. (2015). HICO: a benchmark for recognizing human-object interactions in images. In Proceedings of the IEEE international conference on computer vision (pp. 1017\u20131025). Piscataway: IEEE."},{"key":"102_CR66","first-page":"770","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"K. He","year":"2016","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770\u2013778). Piscataway: IEEE."}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-025-00102-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-025-00102-0","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-025-00102-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,19]],"date-time":"2025-12-19T05:01:36Z","timestamp":1766120496000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-025-00102-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12]]},"references-count":66,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,12]]}},"alternative-id":["102"],"URL":"https:\/\/doi.org\/10.1007\/s44267-025-00102-0","relation":{},"ISSN":["2097-3330","2731-9008"],"issn-type":[{"type":"print","value":"2097-3330"},{"type":"electronic","value":"2731-9008"}],"subject":[],"published":{"date-parts":[[2025,12]]},"assertion":[{"value":"6 August 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 November 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 December 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 December 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no competing interests to declare relevant to this article\u2019s content.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"30"}}