{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T01:22:27Z","timestamp":1773105747285,"version":"3.50.1"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,11,10]],"date-time":"2023-11-10T00:00:00Z","timestamp":1699574400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Integrated Program of National Natural Science Foundation of China","award":["91948303"],"award-info":[{"award-number":["91948303"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>Classifying and accurately locating a visual category with few annotated training samples in computer vision has motivated the few-shot object detection technique, which exploits transfering the source-domain detection model to the target domain. Under this paradigm, however, such transferred source-domain detection model usually encounters difficulty in the classification of the target domain because of the low data diversity of novel training samples. To combat this, we present a simple yet effective few-shot detector, Transferable RCNN. To transfer general knowledge learned from data-abundant base classes to data-scarce novel classes, we propose a weight transfer strategy to promote model transferability and an attention-based feature enhancement mechanism to learn more robust object proposal feature representations. Further, we ensure strong discrimination by optimizing the contrastive objectives of feature maps via a supervised spatial contrastive loss. Meanwhile, we introduce an angle-guided additive margin classifier to augment instance-level inter-class difference and intra-class compactness, which is beneficial for improving the discriminative power of the few-shot classification head under a few supervisions. Our proposed framework outperforms the current works in various settings of PASCAL VOC and MSCOCO datasets; this demonstrates the effectiveness and generalization ability.<\/jats:p>","DOI":"10.1145\/3608478","type":"journal-article","created":{"date-parts":[[2023,7,12]],"date-time":"2023-07-12T11:43:38Z","timestamp":1689162218000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Boosting Few-shot Object Detection with Discriminative Representation and Class Margin"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8822-1558","authenticated-orcid":false,"given":"Yanyan","family":"Shi","sequence":"first","affiliation":[{"name":"College of Computer, National University of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8398-5612","authenticated-orcid":false,"given":"Shaowu","family":"Yang","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6997-0406","authenticated-orcid":false,"given":"Wenjing","family":"Yang","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8112-371X","authenticated-orcid":false,"given":"Dianxi","family":"Shi","sequence":"additional","affiliation":[{"name":"National Innovation Institute of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-0099-4463","authenticated-orcid":false,"given":"Xuehui","family":"Li","sequence":"additional","affiliation":[{"name":"National Innovation Institute of Defense Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,11,10]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3519022"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11716"},{"key":"e_1_3_1_4_2","article-title":"A closer look at few-shot classification","author":"Chen Wei-Yu","year":"2019","unstructured":"Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, and Jia-Bin Huang. 2019. A closer look at few-shot classification. arXiv Preprint arXiv:1904.04232 (2019).","journal-title":"arXiv Preprint arXiv:1904.04232"},{"key":"e_1_3_1_5_2","article-title":"Leveraging bottom-up and top-down attention for few-shot object detection","author":"Chen Xianyu","year":"2020","unstructured":"Xianyu Chen, Ming Jiang, and Qi Zhao. 2020. Leveraging bottom-up and top-down attention for few-shot object detection. arXiv Preprint arXiv:2007.12104 (2020).","journal-title":"arXiv Preprint arXiv:2007.12104"},{"issue":"2","key":"e_1_3_1_6_2","first-page":"3","article-title":"A new meta-baseline for few-shot learning","volume":"1","author":"Chen Yinbo","year":"2020","unstructured":"Yinbo Chen, Xiaolong Wang, Zhuang Liu, Huijuan Xu, Trevor Darrell, et\u00a0al. 2020. A new meta-baseline for few-shot learning. arXiv Preprint arXiv:2003.04390 1, 2 (2020), 3.","journal-title":"arXiv Preprint arXiv:2003.04390"},{"key":"e_1_3_1_7_2","first-page":"168","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201918)","author":"Cheng Hao","year":"2018","unstructured":"Hao Cheng, Dongze Lian, Shenghua Gao, and Yanlin Geng. 2018. Evaluating capability of deep neural networks for image classification via information plane. In Proceedings of the European Conference on Computer Vision (ECCV \u201918). 168\u2013182."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00407"},{"key":"e_1_3_1_9_2","first-page":"4527","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fan Zhibo","year":"2021","unstructured":"Zhibo Fan, Yuchen Ma, Zeming Li, and Jian Sun. 2021. Generalized few-shot object detection without forgetting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 4527\u20134536."},{"key":"e_1_3_1_10_2","first-page":"1126","volume-title":"International Conference on Machine Learning","author":"Finn Chelsea","year":"2017","unstructured":"Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning. PMLR, 1126\u20131135."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00494"},{"key":"e_1_3_1_12_2","first-page":"14","volume-title":"Proceedings of the 24th International Conference on Neural Information Processing (ICONIP \u201917), , Part III 24","author":"Han Guangxing","year":"2017","unstructured":"Guangxing Han, Xuan Zhang, and Chongrong Li. 2017. Revisiting faster R-CNN: A deeper look at region proposal network. In Proceedings of the 24th International Conference on Neural Information Processing (ICONIP \u201917), , Part III 24. Springer, 14\u201324."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00092"},{"issue":"2","key":"e_1_3_1_14_2","first-page":"1","article-title":"Multi-peak graph-based multi-instance learning for weakly supervised object detection","volume":"17","author":"Ji Ruyi","year":"2021","unstructured":"Ruyi Ji, Zeyu Liu, Libo Zhang, Jianwei Liu, Xin Zuo, Yanjun Wu, Chen Zhao, Haofeng Wang, and Lin Yang. 2021. Multi-peak graph-based multi-instance learning for weakly supervised object detection. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 17, 2s (2021), 1\u201321.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)"},{"key":"e_1_3_1_15_2","first-page":"8420","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Kang Bingyi","year":"2019","unstructured":"Bingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu, Jiashi Feng, and Trevor Darrell. 2019. Few-shot object detection via feature reweighting. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 8420\u20138429."},{"key":"e_1_3_1_16_2","first-page":"5197","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Karlinsky Leonid","year":"2019","unstructured":"Leonid Karlinsky, Joseph Shtok, Sivan Harary, Eli Schwartz, Amit Aides, Rogerio Feris, Raja Giryes, and Alex M. Bronstein. 2019. Repmet: Representative-based metric learning for classification and few-shot object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5197\u20135206."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00019"},{"key":"e_1_3_1_18_2","first-page":"18661","article-title":"Supervised contrastive learning","volume":"33","author":"Khosla Prannay","year":"2020","unstructured":"Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. Advances in Neural Information Processing Systems 33 (2020), 18661\u201318673.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.aab3050"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01259"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01348"},{"key":"e_1_3_1_22_2","first-page":"15395","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Yiting","year":"2021","unstructured":"Yiting Li, Haiyue Zhu, Yu Cheng, Wenxin Wang, Chek Sing Teo, Cheng Xiang, Prahlad Vadakkepat, and Tong Heng Lee. 2021. Few-shot object detection via classification refinement and distractor retreatment. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 15395\u201315403."},{"key":"e_1_3_1_23_2","article-title":"Self-EMD: Self-supervised object detection without Imagenet","author":"Liu Songtao","year":"2020","unstructured":"Songtao Liu, Zeming Li, and Jian Sun. 2020. Self-EMD: Self-supervised object detection without Imagenet. arXiv preprint arXiv:2011.13677 (2020).","journal-title":"arXiv preprint arXiv:2011.13677"},{"key":"e_1_3_1_24_2","first-page":"21","volume-title":"Proceedings of the 14th European Conference on Computer Vision (ECCV \u201916), Part I 14","author":"Liu Wei","year":"2016","unstructured":"Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg. 2016. SSD: Single shot multibox detector. In Proceedings of the 14th European Conference on Computer Vision (ECCV \u201916), Part I 14. Springer, 21\u201337."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3381086"},{"issue":"3","key":"e_1_3_1_26_2","first-page":"4","article-title":"Reptile: A scalable metalearning algorithm","volume":"2","author":"Nichol Alex","year":"2018","unstructured":"Alex Nichol and John Schulman. 2018. Reptile: A scalable metalearning algorithm. arXiv Preprint arXiv:1803.02999 2, 3 (2018), 4.","journal-title":"arXiv Preprint arXiv:1803.02999"},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","first-page":"671","DOI":"10.1007\/978-3-030-86486-6_41","volume-title":"Proceedings of the European Conference on Machine Learning and Knowledge Discovery in Databases. Research Track (ECML PKDD \u201921), Part I 21","author":"Ouali Yassine","year":"2021","unstructured":"Yassine Ouali, C\u00e9line Hudelot, and Myriam Tami. 2021. Spatial contrastive learning for few-shot classification. In Proceedings of the European Conference on Machine Learning and Knowledge Discovery in Databases. Research Track (ECML PKDD \u201921), Part I 21. Springer, 671\u2013686."},{"key":"e_1_3_1_28_2","volume-title":"Proceedings of the IEEE\/CVF International Conference on Learning Representations","author":"Ravi Sachin","year":"2016","unstructured":"Sachin Ravi and Hugo Larochelle. 2016. Optimization as a model for few-shot learning. In Proceedings of the IEEE\/CVF International Conference on Learning Representations."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_3_1_30_2","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems 28 (2015).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1146\/annurev-bioeng-071516-044442"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.136"},{"key":"e_1_3_1_33_2","article-title":"Prototypical networks for few-shot learning","volume":"30","author":"Snell Jake","year":"2017","unstructured":"Jake Snell, Kevin Swersky, and Richard Zemel. 2017. Prototypical networks for few-shot learning. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00727"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00131"},{"key":"e_1_3_1_36_2","first-page":"266","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XIV 16","author":"Tian Yonglong","year":"2020","unstructured":"Yonglong Tian, Yue Wang, Dilip Krishnan, Joshua B. Tenenbaum, and Phillip Isola. 2020. Rethinking few-shot image classification: A good embedding is all you need? In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XIV 16. Springer, 266\u2013282."},{"key":"e_1_3_1_37_2","article-title":"Meta-learning: A survey","author":"Vanschoren Joaquin","year":"2018","unstructured":"Joaquin Vanschoren. 2018. Meta-learning: A survey. arXiv Preprint arXiv:1810.03548 (2018).","journal-title":"arXiv Preprint arXiv:1810.03548"},{"key":"e_1_3_1_38_2","article-title":"Matching networks for one shot learning","volume":"29","author":"Vinyals Oriol","year":"2016","unstructured":"Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et\u00a0al. 2016. Matching networks for one shot learning. Advances in Neural Information Processing Systems 29 (2016).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00552"},{"key":"e_1_3_1_40_2","article-title":"Frustratingly simple few-shot object detection","author":"Wang Xin","year":"2020","unstructured":"Xin Wang, Thomas E. Huang, Trevor Darrell, Joseph E. Gonzalez, and Fisher Yu. 2020. Frustratingly simple few-shot object detection. arXiv Preprint arXiv:2003.06957 (2020).","journal-title":"arXiv Preprint arXiv:2003.06957"},{"key":"e_1_3_1_41_2","first-page":"1831","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Xin","year":"2019","unstructured":"Xin Wang, Fisher Yu, Ruth Wang, Trevor Darrell, and Joseph E. Gonzalez. 2019. Tafe-net: Task-aware feature embeddings for low shot learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1831\u20131840."},{"key":"e_1_3_1_42_2","first-page":"3024","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Xinlong","year":"2021","unstructured":"Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li. 2021. Dense contrastive learning for self-supervised visual pre-training. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3024\u20133033."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.01002"},{"key":"e_1_3_1_44_2","first-page":"456","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XVI 16","author":"Wu Jiaxi","year":"2020","unstructured":"Jiaxi Wu, Songtao Liu, Di Huang, and Yunhong Wang. 2020. Multi-scale positive sample refinement for few-shot object detection. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XVI 16. Springer, 456\u2013472."},{"issue":"3","key":"e_1_3_1_45_2","first-page":"3090","article-title":"Few-shot object detection and viewpoint estimation for objects in the wild","volume":"45","author":"Xiao Yang","year":"2022","unstructured":"Yang Xiao, Vincent Lepetit, and Renaud Marlet. 2022. Few-shot object detection and viewpoint estimation for objects in the wild. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 3 (2022), 3090\u20133106.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00967"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3472393"},{"key":"e_1_3_1_48_2","first-page":"3521","article-title":"Restoring negative information in few-shot object detection","volume":"33","author":"Yang Yukuan","year":"2020","unstructured":"Yukuan Yang, Fangyun Wei, Miaojing Shi, and Guoqi Li. 2020. Restoring negative information in few-shot object detection. Advances in Neural Information Processing Systems 33 (2020), 3521\u20133532.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_49_2","first-page":"818","volume-title":"Proceedings of the 13th European Conference on Computer Vision (ECCV \u201914),Part I 13","author":"Zeiler Matthew D.","year":"2014","unstructured":"Matthew D. Zeiler and Rob Fergus. 2014. Visualizing and understanding convolutional networks. In Proceedings of the 13th European Conference on Computer Vision (ECCV \u201914),Part I 13. Springer, 818\u2013833."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3608478","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3608478","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:29:45Z","timestamp":1750285785000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3608478"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,10]]},"references-count":48,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3608478"],"URL":"https:\/\/doi.org\/10.1145\/3608478","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,10]]},"assertion":[{"value":"2022-03-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-29","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}