{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,7]],"date-time":"2026-05-07T16:43:27Z","timestamp":1778172207143,"version":"3.51.4"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,12,23]],"date-time":"2024-12-23T00:00:00Z","timestamp":1734912000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Project of China","award":["2021ZD0110700"],"award-info":[{"award-number":["2021ZD0110700"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U19B2043, 61976185"],"award-info":[{"award-number":["U19B2043, 61976185"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["226-2022-00051"],"award-info":[{"award-number":["226-2022-00051"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"name":"HKUST Special Support for Young Faculty","award":["F0927"],"award-info":[{"award-number":["F0927"]}]},{"name":"HKUST Sports Science and Technology Research","award":["SSTRG24EG04"],"award-info":[{"award-number":["SSTRG24EG04"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,1,31]]},"abstract":"<jats:p>\n            Today's scene graph generation (SGG) models typically require abundant manual annotations to learn new predicate types. Therefore, it is difficult to apply them to real-world applications with massive uncommon predicate categories whose annotations are hard to collect. In this article, we focus on\n            <jats:italic>Few-Shot SGG (FSSGG)<\/jats:italic>\n            , which encourages SGG models to be able to quickly transfer previous knowledge and recognize unseen predicates well with only a few examples. However, current methods for FSSGG are hindered by the high intra-class variance of predicate categories in SGG: On one hand, each predicate category commonly has multiple semantic meanings under different contexts. On the other hand, the visual appearance of relation triplets with the same predicate differs greatly under different subject\u2013object compositions. Such great variance of inputs makes it hard to learn generalizable representation for each predicate category with current few-shot learning (FSL) methods. However, we found that this intra-class variance of predicates is highly related to the composed subjects and objects. To model the intra-class variance of predicates with subject\u2013object context, we propose a novel\n            <jats:italic>Decomposed Prototype Learning (DPL)<\/jats:italic>\n            model for FSSGG. Specifically, we first construct a decomposable prototype space to capture diverse semantics and visual patterns of subjects and objects for predicates by decomposing them into multiple prototypes. Afterwards, we integrate these prototypes with different weights to generate query-adaptive predicate representation with more reliable semantics for each query sample. We conduct extensive experiments and compare with various baseline methods to show the effectiveness of our method.\n          <\/jats:p>","DOI":"10.1145\/3700877","type":"journal-article","created":{"date-parts":[[2024,10,21]],"date-time":"2024-10-21T15:55:21Z","timestamp":1729526121000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Decomposed Prototype Learning for Few-Shot Scene Graph Generation"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8204-1302","authenticated-orcid":false,"given":"Xingchen","family":"Li","sequence":"first","affiliation":[{"name":"Zhejiang University, Hangzhou, China and China Mobile (Zhejiang) Innovation Research Institute, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0303-134X","authenticated-orcid":false,"given":"Jun","family":"Xiao","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9227-007X","authenticated-orcid":false,"given":"Guikun","family":"Chen","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9136-0965","authenticated-orcid":false,"given":"Yinfu","family":"Feng","sequence":"additional","affiliation":[{"name":"Alibaba Group, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0512-880X","authenticated-orcid":false,"given":"Yi","family":"Yang","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5755-9145","authenticated-orcid":false,"given":"An-An","family":"Liu","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6148-9709","authenticated-orcid":false,"given":"Long","family":"Chen","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,12,23]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"NeurIPS","author":"Alayrac Jean-Baptiste","year":"2022","unstructured":"Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: A visual language model for few-shot learning. In NeurIPS."},{"key":"e_1_3_2_3_2","volume-title":"NeurIPS","author":"Chen Guikun","year":"2024","unstructured":"Guikun Chen, Jin Li, and Wenguan Wang. 2024. Scene Graph Generation with Role-Playing Large Language Models. In NeurIPS."},{"key":"e_1_3_2_4_2","volume-title":"Addressing predicate overlap in scene graph generation with semantic granularity controller. In ICME","author":"Chen Guikun","year":"2023","unstructured":"Guikun Chen, Lin Li, Yawei Luo, and Jun Xiao. 2023. Addressing predicate overlap in scene graph generation with semantic granularity controller. In ICME."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01750"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00471"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01882"},{"key":"e_1_3_2_8_2","first-page":"0","volume-title":"ICCV Workshops","author":"Dornadula Apoorva","year":"2019","unstructured":"Apoorva Dornadula, Austin Narcomey, Ranjay Krishna, Michael Bernstein, and Fei-Fei Li. 2019. Visual relationships as functions: Enabling few-shot scene graph prediction. In ICCV Workshops, 0\u20130."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01369"},{"key":"e_1_3_2_10_2","first-page":"1126","volume-title":"ICML","author":"Finn Chelsea","year":"2017","unstructured":"Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In ICML, 1126\u20131135."},{"key":"e_1_3_2_11_2","volume-title":"ICLR","author":"Gao Kaifeng","year":"2023","unstructured":"Kaifeng Gao, Long Chen, Hanwang Zhang, Jun Xiao, and Qianru Sun. 2023. Compositional prompt tuning with motion cues for open-vocabulary video relation detection. In ICLR."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00459"},{"key":"e_1_3_2_13_2","unstructured":"Yuxian Gu Xu Han Zhiyuan Liu and Minlie Huang. 2021. PPT: Pre-trained prompt tuning for few-shot learning. arXiv:2109.04332. Retrieved from https:\/\/arxiv.org\/abs\/2109.04332"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414025"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19815-1_4"},{"key":"e_1_3_2_16_2","first-page":"1091","volume-title":"IEEE TCSVT","author":"Jiang Wen","year":"2020","unstructured":"Wen Jiang, Kai Huang, Jie Geng, and Xinyang Deng. 2020. Multi-scale metric learning for few-shot learning. IEEE TCSVT 31, 3 (2020), 1091\u20131102."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00969"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00889"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01091"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00823"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01982"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01830"},{"key":"e_1_3_2_24_2","article-title":"Zero-shot visual relation detection via composite visual cues from large language models","volume":"36","author":"Li Lin","year":"2023","unstructured":"Lin Li, Jun Xiao, Guikun Chen, Jian Shao, Yueting Zhuang, and Long Chen. 2023. Zero-shot visual relation detection via composite visual cues from large language models. In NeurIPS, Vol. 36.","journal-title":"NeurIPS"},{"issue":"1","key":"e_1_3_2_25_2","first-page":"195","article-title":"Label semantic knowledge distillation for unbiased scene graph generation","volume":"34","author":"Li Lin","year":"2023","unstructured":"Lin Li, Jun Xiao, Hanrong Shi, Wenxiao Wang, Jian Shao, An-An Liu, Yi Yang, and Long Chen. 2023. Label semantic knowledge distillation for unbiased scene graph generation. IEEE TCSVT 34, 1 (2023), 195\u2013206.","journal-title":"IEEE TCSVT"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3387349"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01888"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00743"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018642"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548164"},{"key":"e_1_3_2_31_2","volume-title":"BMVC","author":"Li Xingchen","year":"2022","unstructured":"Xingchen Li, Long Chen, Jian Shao, Shaoning Xiao, Songyang Zhang, and Jun Xiao. 2022. Rethinking the evaluation of unbiased scene graph generation. In BMVC."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.353"},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","unstructured":"Yiming Li Xiaoshan Yang Xuhui Huang Zhe Ma and Changsheng Xu. 2022. Zero-shot predicate prediction for scene graph parsing. IEEE TMM 25 (2022) 3140\u20133153.","DOI":"10.1109\/TMM.2022.3155928"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_51"},{"key":"e_1_3_2_35_2","unstructured":"Xinyu Lyu Lianli Gao Junlin Xie Pengpeng Zeng Yulu Tian Jie Shao and Heng Tao Shen. 2023. Generalized unbiased scene graph generation. arXiv:2308.04802. Retrieved from https:\/\/arxiv.org\/abs\/2308.04802"},{"issue":"11","key":"e_1_3_2_36_2","first-page":"13921","article-title":"Adaptive fine-grained predicates learning for scene graph generation","volume":"45","author":"Lyu Xinyu","year":"2023","unstructured":"Xinyu Lyu, Lianli Gao, Pengpeng Zeng, Heng Tao Shen, and Jingkuan Song. 2023. Adaptive fine-grained predicates learning for scene graph generation. IEEE TPAMI 45, 11 (2023), 13921\u201313940.","journal-title":"IEEE TPAMI"},{"issue":"9","key":"e_1_3_2_37_2","first-page":"4616","article-title":"Understanding and mitigating overfitting in prompt tuning for vision-language models","volume":"33","author":"Ma Chengcheng","year":"2023","unstructured":"Chengcheng Ma, Yang Liu, Jiankang Deng, Lingxi Xie, Weiming Dong, and Changsheng Xu. 2023. Understanding and mitigating overfitting in prompt tuning for vision-language models. IEEE TCSVT 33, 9 (2023), 4616\u20134629.","journal-title":"IEEE TCSVT"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20062-5_39"},{"key":"e_1_3_2_39_2","first-page":"8748","volume-title":"ICML","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In ICML. PMLR, 8748\u20138763."},{"key":"e_1_3_2_40_2","first-page":"1842","volume-title":"ICML","author":"Santoro Adam","year":"2016","unstructured":"Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap. 2016. Meta-learning with memory-augmented neural networks. In ICML. PMLR, 1842\u20131850."},{"key":"e_1_3_2_41_2","unstructured":"Hanrong Shi Lin Li Jun Xiao Yueting Zhuang and Long Chen. 2024. From easy to hard: Learning curricular shape-aware features for robust panoptic scene graph generation. IJCV (2024) 1\u201320. Retrieved from https:\/\/link.springer.com\/article\/10.1007\/s11263-024-02190-9"},{"key":"e_1_3_2_42_2","article-title":"Prototypical networks for few-shot learning","volume":"30","author":"Snell Jake","year":"2017","unstructured":"Jake Snell, Kevin Swersky, and Richard Zemel. 2017. Prototypical networks for few-shot learning. In NeurIPS, Vol. 30.","journal-title":"NeurIPS"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00131"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00377"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00678"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01883"},{"key":"e_1_3_2_47_2","first-page":"200","article-title":"Multimodal few-shot learning with frozen language models","volume":"34","author":"Tsimpoukelli Maria","year":"2021","unstructured":"Maria Tsimpoukelli, Jacob L. Menick, Serkan Cabi, S. M. Eslami, Oriol Vinyals, and Felix Hill. 2021. Multimodal few-shot learning with frozen language models. In NeurIPS, Vol. 34, 200\u2013212.","journal-title":"NeurIPS"},{"key":"e_1_3_2_48_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS, Vol. 30.","journal-title":"NeurIPS"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00929"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6904"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_41"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02212"},{"key":"e_1_3_2_53_2","unstructured":"Tianyu Yu Yangning Li Jiaoyan Chen Yinghui Li Hai-Tao Zheng Xi Chen Qingbin Liu Wenqiang Liu Dongxiao Huang Bei Wu et al. 2023. Knowledge-augmented few-shot visual relation detection. arXiv:2303.05342. Retrieved from https:\/\/arxiv.org\/abs\/2303.05342"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00611"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00553"},{"issue":"1","key":"e_1_3_2_56_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3514041","article-title":"Boosting scene graph generation with visual relation saliency","volume":"19","author":"Zhang Yong","year":"2023","unstructured":"Yong Zhang, Yingwei Pan, Ting Yao, Rui Huang, Tao Mei, and Chang-Wen Chen. 2023. Boosting scene graph generation with visual relation saliency. ACM TOMM 19, 1 (2023), 1\u201317.","journal-title":"ACM TOMM"},{"issue":"3","key":"e_1_3_2_57_2","first-page":"1743","article-title":"Dual-branch hybrid learning network for unbiased scene graph generation","volume":"34","author":"Zheng Chaofan","year":"2023","unstructured":"Chaofan Zheng, Lianli Gao, Xinyu Lyu, Pengpeng Zeng, Abdulmotaleb El Saddik, and Heng Tao Shen. 2023. Dual-branch hybrid learning network for unbiased scene graph generation. IEEE TCSVT 34, 3 (2023), 1743\u20131756.","journal-title":"IEEE TCSVT"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02182"},{"issue":"5","key":"e_1_3_2_59_2","first-page":"2102","article-title":"Quaternion-valued correlation learning for few-shot semantic segmentation","volume":"33","author":"Zheng Zewen","year":"2022","unstructured":"Zewen Zheng, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Hongrui Liu, and Wing-Kuen Ling. 2022. Quaternion-valued correlation learning for few-shot semantic segmentation. IEEE TCSVT 33, 5 (2022), 2102\u20132115.","journal-title":"IEEE TCSVT"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00184"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01631"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-022-01653-1"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3673231"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3700877","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3700877","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:10:23Z","timestamp":1750295423000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3700877"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,23]]},"references-count":62,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1,31]]}},"alternative-id":["10.1145\/3700877"],"URL":"https:\/\/doi.org\/10.1145\/3700877","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,23]]},"assertion":[{"value":"2023-12-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-23","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}