{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,21]],"date-time":"2026-06-21T05:42:40Z","timestamp":1782020560254,"version":"3.54.5"},"reference-count":132,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100000923","name":"Australian Research Council","doi-asserted-by":"crossref","award":["ARC DP210101682, DP210102674"],"award-info":[{"award-number":["ARC DP210101682, DP210102674"]}],"id":[{"id":"10.13039\/501100000923","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Defence Science and Technology Group (DSTG) for the project \u201cLow Observer Detection of Small Objects in Maritime Scenes\u201d"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>Transformers have rapidly gained popularity in computer vision, especially in the field of object detection. Upon examining the outcomes of state-of-the-art object detection methods, we noticed that transformers consistently outperformed well-established CNN-based detectors in almost every video or image dataset. Small objects have been identified as one of the most challenging object types in detection frameworks due to their low visibility. This article aims to explore the performance benefits offered by such extensive networks and identify potential reasons for their Small Object Detection (SOD) superiority. We aim to investigate potential strategies that could further enhance transformers\u2019 performance in SOD. This survey presents a taxonomy of over 60 research studies on developed transformers for the task of SOD, spanning the years 2020 to 2023. These studies encompass a variety of detection applications, including SOD in generic images, aerial images, medical images, active millimeter images, underwater images, and videos. We also compile and present a list of 12 large-scale datasets suitable for SOD that were overlooked in previous studies and compare the performance of the reviewed studies using popular metrics such as mean Average Precision (mAP), Frames Per Second (FPS), and number of parameters. Researchers can keep track of newer studies on our web page, which is available at: https:\/\/github.com\/arekavandi\/Transformer-SOD.<\/jats:p>","DOI":"10.1145\/3758090","type":"journal-article","created":{"date-parts":[[2025,8,11]],"date-time":"2025-08-11T11:26:16Z","timestamp":1754911576000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":26,"title":["Transformers in Small Object Detection: A Benchmark and Survey of State-of-the-Art"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9542-759X","authenticated-orcid":false,"given":"Aref","family":"Miri Rekavandi","sequence":"first","affiliation":[{"name":"The University of Western Australia","place":["Perth, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5154-2385","authenticated-orcid":false,"given":"Shima","family":"Rashidi","sequence":"additional","affiliation":[{"name":"RMIT University","place":["Melbourne, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7250-7407","authenticated-orcid":false,"given":"Farid","family":"Boussaid","sequence":"additional","affiliation":[{"name":"The University of Western Australia","place":["Perth, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-5181-3042","authenticated-orcid":false,"given":"Stephen","family":"Hoefs","sequence":"additional","affiliation":[{"name":"Defence Science and Technology Group","place":["Perth, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3760-6722","authenticated-orcid":false,"given":"Emre","family":"Akbas","sequence":"additional","affiliation":[{"name":"Middle East Technical University","place":["Ankara, Turkey"]},{"name":"Helmholtz Center Munich German Research Center for Environmental Health","place":["Ankara, Turkey"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6603-3257","authenticated-orcid":false,"given":"Mohammed","family":"Bennamoun","sequence":"additional","affiliation":[{"name":"School of Computer Science and Software Engineering, The University of Western Australia","place":["Perth, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,10]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"2020. Udacity self-driving car driving data 2017 transformer. Retrieved Sep 2023 from https:\/\/github.com\/udacity\/self-driving-car\/tree\/master\/annotations"},{"key":"e_1_3_2_3_2","unstructured":"Josh Beal Eric Kim Eric Tzeng Dong Huk Park Andrew Zhai and Dmitry Kislyuk. 2020. Toward transformer-based object detection. Retrieved from https:\/\/arxiv.org\/abs\/2012.09958"},{"key":"e_1_3_2_4_2","unstructured":"Alexey Bochkovskiy et\u00a0al. 2020. Yolov4: Optimal speed and accuracy of object detection. Retrieved from https:\/\/arxiv.org\/abs\/2004.10934"},{"key":"e_1_3_2_5_2","first-page":"4777","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cai Likun","year":"2022","unstructured":"Likun Cai, Zhi Zhang, Yi Zhu, Li Zhang, Mu Li, and Xiangyang Xue. 2022. Bigdetection: A large-scale benchmark for improved object detector pre-training. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 4777\u20134787."},{"issue":"5","key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"1483","DOI":"10.1109\/TPAMI.2019.2956516","article-title":"Cascade R-CNN: High quality object detection and instance segmentation","volume":"43","author":"Cai Zhaowei","year":"2021","unstructured":"Zhaowei Cai and Nuno Vasconcelos. 2021. Cascade R-CNN: High quality object detection and instance segmentation. TPAMI 43, 5 (2021), 1483\u20131498.","journal-title":"TPAMI"},{"key":"e_1_3_2_7_2","first-page":"185","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"36","author":"Cao Xipeng","year":"2022","unstructured":"Xipeng Cao, Peng Yuan, Bailan Feng, and Kun Niu. 2022. CF-DETR: Coarse-to-fine transformers for end-to-end object detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 185\u2013193."},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1007\/978-3-030-58452-8_13","volume-title":"ECCV 2020: Proceedings of the 16th European Conference on Computer Vision, Part I 16","author":"Carion Nicolas","year":"2020","unstructured":"Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In ECCV 2020: Proceedings of the 16th European Conference on Computer Vision, Part I 16. Springer, 213\u2013229."},{"key":"e_1_3_2_9_2","first-page":"1","volume-title":"ICASSP 2023-Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Chen Dong","year":"2023","unstructured":"Dong Chen, Duoqian Miao, and Xuerong Zhao. 2023. Hyneter: Hybrid network transformer for object detection. In ICASSP 2023-Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1\u20135."},{"issue":"4","key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"1076","DOI":"10.3390\/rs15041076","article-title":"HTDet: A hybrid transformer-based approach for underwater small object detection","volume":"15","author":"Chen Gangqi","year":"2023","unstructured":"Gangqi Chen, Zhaoyong Mao, Kai Wang, and Junge Shen. 2023. HTDet: A hybrid transformer-based approach for underwater small object detection. Remote Sensing 15, 4 (2023), 1076.","journal-title":"Remote Sensing"},{"issue":"2","key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"371","DOI":"10.3390\/rs15020371","article-title":"MDCT: Multi-kernel dilated convolution and transformer for one-stage object detection of remote sensing images","volume":"15","author":"Chen Juanjuan","year":"2023","unstructured":"Juanjuan Chen, Hansheng Hong, Bin Song, Jie Guo, Chen Chen, and Junjie Xu. 2023. MDCT: Multi-kernel dilated convolution and transformer for one-stage object detection of remote sensing images. Remote Sensing 15, 2 (2023), 371.","journal-title":"Remote Sensing"},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","first-page":"70","DOI":"10.1007\/978-3-031-20080-9_5","volume-title":"ECCV 2022: Proceedings of the 17th European Conference on Computer Vision, Part X","author":"Chen Peixian","year":"2022","unstructured":"Peixian Chen, Mengdan Zhang, Yunhang Shen, Kekai Sheng, Yuting Gao, Xing Sun, Ke Li, and Chunhua Shen. 2022. Efficient decoder-free object detection with transformers. In ECCV 2022: Proceedings of the 17th European Conference on Computer Vision, Part X. Springer, 70\u201386."},{"key":"e_1_3_2_13_2","first-page":"6633","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Chen Qiang","year":"2023","unstructured":"Qiang Chen, Xiaokang Chen, Jian Wang, Shan Zhang, Kun Yao, Haocheng Feng, Junyu Han, Errui Ding, Gang Zeng, and Jingdong Wang. 2023. Group DETR: Fast DETR training with group-wise one-to-many assignment. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 6633\u20136642."},{"key":"e_1_3_2_14_2","unstructured":"Qiang Chen Jian Wang Chuchu Han Shan Zhang Zexian Li Xiaokang Chen Jiahui Chen Xiaodi Wang Shuming Han Gang Zhang Haocheng Feng Kun Yao Junyu Han Errui Ding and Jingdong Wang. 2022. Group DETR v2: Strong object detector with encoder-decoder pretraining. Retrieved from https:\/\/arxiv.org\/abs\/2211.03594"},{"key":"e_1_3_2_15_2","unstructured":"Xiaokang Chen Fangyun Wei Gang Zeng and Jingdong Wang. 2022. Conditional DETR v2: Efficient detection transformer with box queries. Retrieved from https:\/\/arxiv.org\/abs\/2207.08914"},{"key":"e_1_3_2_16_2","first-page":"8126","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Xin","year":"2021","unstructured":"Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. 2021. Transformer tracking. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8126\u20138135."},{"key":"e_1_3_2_17_2","first-page":"5621","article-title":"Reppoints v2: Verification meets regression for object detection","volume":"33","author":"Chen Yihong","year":"2020","unstructured":"Yihong Chen, Zheng Zhang, Yue Cao, Liwei Wang, Stephen Lin, and Han Hu. 2020. Reppoints v2: Verification meets regression for object detection. Advances in Neural Information Processing Systems 33 (2020), 5621\u20135631.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_18_2","article-title":"Towards large-scale small object detection: Survey and benchmarks","author":"Cheng Gong","year":"2023","unstructured":"Gong Cheng, Xiang Yuan, Xiwen Yao, Kebing Yan, Qinghua Zeng, Xingxing Xie, and Junwei Han. 2023. Towards large-scale small object detection: Survey and benchmarks. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 11 (2023), 13467\u201313488.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"12","key":"e_1_3_2_19_2","doi-asserted-by":"crossref","first-page":"7405","DOI":"10.1109\/TGRS.2016.2601622","article-title":"Learning rotation-invariant convolutional neural networks for object detection in VHR optical remote sensing images","volume":"54","author":"Cheng Gong","year":"2016","unstructured":"Gong Cheng, Peicheng Zhou, and Junwei Han. 2016. Learning rotation-invariant convolutional neural networks for object detection in VHR optical remote sensing images. IEEE Transactions on Geoscience and Remote Sensing 54, 12 (2016), 7405\u20137415.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_2_20_2","first-page":"13564","article-title":"Relationnet++: Bridging visual representations for object detection via transformer decoder","volume":"33","author":"Chi Cheng","year":"2020","unstructured":"Cheng Chi, Fangyun Wei, and Han Hu. 2020. Relationnet++: Bridging visual representations for object detection via transformer decoder. Advances in Neural Information Processing Systems 33 (2020), 13564\u201313574.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"9","key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"2995","DOI":"10.3390\/s18092995","article-title":"Face detection in nighttime images using visible-light camera sensors with two-step faster region-based convolutional neural network","volume":"18","author":"Cho Se Woon","year":"2018","unstructured":"Se Woon Cho, Na Rae Baek, Min Cheol Kim, Ja Hyung Koo, Jong Hyun Kim, and Kang Ryoung Park. 2018. Face detection in nighttime images using visible-light camera sensors with two-step faster region-based convolutional neural network. Sensors 18, 9 (2018), 2995.","journal-title":"Sensors"},{"key":"e_1_3_2_22_2","first-page":"1","volume-title":"Proceedings of the 2021 17th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS)","author":"Coluccia Angelo","year":"2021","unstructured":"Angelo Coluccia, Alessio Fascista, Arne Schumann, Lars Sommer, Anastasios Dimou, Dimitrios Zarpalas, Fatih Cagatay Akyon, Ogulcan Eryuksel, Kamil Anil Ozfuttu, Sinan Onur Altinuc, et\u00a0al. 2021. Drone-vs-bird detection challenge at IEEE AVSS2021. In Proceedings of the 2021 17th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS). IEEE, 1\u20138."},{"key":"e_1_3_2_23_2","first-page":"6365","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cui Yiming","year":"2023","unstructured":"Yiming Cui. 2023. Feature aggregated queries for transformer-based video object detectors. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6365\u20136376."},{"key":"e_1_3_2_24_2","article-title":"R-FCN: Object detection via region-based fully convolutional networks","author":"Sun Jifeng Dai, Yi Li, Kaiming He, and Jian","year":"2016","unstructured":"Jifeng Dai, Yi Li, Kaiming He, and Jian Sun. 2016. R-FCN: Object detection via region-based fully convolutional networks. In Proceedings of the 30th International Conference on Neural Information Processing Systems (NeurIPS).","journal-title":"Proceedings of the 30th International Conference on Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_2_25_2","article-title":"Ao2-detr: Arbitrary-oriented object detection transformer","author":"Dai Linhui","year":"2023","unstructured":"Linhui Dai, Hong Liu, Hao Tang, Zhiwei Wu, and Pinhao Song. 2023. Ao2-detr: Arbitrary-oriented object detection transformer. IEEE Transactions on Circuits and Systems for Video Technology 33, 5 (2023), 2342\u20132356.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_26_2","first-page":"2988","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Dai Xiyang","year":"2021","unstructured":"Xiyang Dai, Yinpeng Chen, Jianwei Yang, Pengchuan Zhang, Lu Yuan, and Lei Zhang. 2021. Dynamic DETR: End-to-end object detection with dynamic attention. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 2988\u20132997."},{"key":"e_1_3_2_27_2","doi-asserted-by":"crossref","first-page":"119560","DOI":"10.1016\/j.eswa.2023.119560","article-title":"Sw-YoloX: An anchor-free detector based transformer for sea surface object detection","author":"Ding Jiangang","year":"2023","unstructured":"Jiangang Ding, Wei Li, Lili Pei, Ming Yang, Chao Ye, and Bo Yuan. 2023. Sw-YoloX: An anchor-free detector based transformer for sea surface object detection. Expert Systems with Applications 217 (2023), 119560.","journal-title":"Expert Systems with Applications"},{"key":"e_1_3_2_28_2","first-page":"2849","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ding Jian","year":"2019","unstructured":"Jian Ding, Nan Xue, Yang Long, Gui-Song Xia, and Qikai Lu. 2019. Learning RoI transformer for oriented object detection in aerial images. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2849\u20132858."},{"issue":"1","key":"e_1_3_2_29_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s11554-023-01280-0","article-title":"DeoT: An end-to-end encoder-only Transformer object detector","volume":"20","author":"Ding Tonghe","year":"2023","unstructured":"Tonghe Ding, Kaili Feng, Yanjun Wei, Yu Han, and Tianping Li. 2023. DeoT: An end-to-end encoder-only Transformer object detector. Journal of Real-Time Image Processing 20, 1 (2023), 1.","journal-title":"Journal of Real-Time Image Processing"},{"key":"e_1_3_2_30_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly Jakob Uszkoreit and Neil Houlsby. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"issue":"5","key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"3509","DOI":"10.1109\/TPAMI.2023.3342120","article-title":"CenterNet++ for object detection","volume":"46","author":"Duan Kaiwen","year":"2023","unstructured":"Kaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian. 2023. CenterNet++ for object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 5 (2023), 3509\u20133521.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","first-page":"103620","DOI":"10.1016\/j.jvcir.2022.103620","article-title":"Improving small objects detection using transformer","volume":"89","author":"Dubey Shikha","year":"2022","unstructured":"Shikha Dubey, Farrukh Olimov, Muhammad Aasim Rafique, and Moongu Jeon. 2022. Improving small objects detection using transformer. Journal of Visual Communication and Image Representation 89 (2022), 103620.","journal-title":"Journal of Visual Communication and Image Representation"},{"key":"e_1_3_2_33_2","first-page":"26183","article-title":"You only look at one sequence: Rethinking transformer in vision through object detection","volume":"34","author":"Fang Yuxin","year":"2021","unstructured":"Yuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu, Jianwei Niu, and Wenyu Liu. 2021. You only look at one sequence: Rethinking transformer in vision through object detection. Advances in Neural Information Processing Systems 34 (2021), 26183\u201326197.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","first-page":"65886","DOI":"10.1109\/ACCESS.2022.3184031","article-title":"Video sparse transformer with attention-guided memory for video object detection","volume":"10","author":"Fujitake Masato","year":"2022","unstructured":"Masato Fujitake and Akihiro Sugimoto. 2022. Video sparse transformer with attention-guided memory for video object detection. IEEE Access 10 (2022), 65886\u201365900.","journal-title":"IEEE Access"},{"key":"e_1_3_2_35_2","unstructured":"Yuan Gao Hui Shen Donghong Zhong Jian Wang Zeyu Liu Ti Bai Xiang Long and Shilei Wen. 2019. A solution for densely annotated large scale object detection task. (2019)."},{"key":"e_1_3_2_36_2","first-page":"1440","volume-title":"Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV)","author":"Girshick Ross","year":"2015","unstructured":"Ross Girshick. 2015. Fast R-CNN. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV). 1440\u20131448."},{"key":"e_1_3_2_37_2","first-page":"5227","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Goldman Eran","year":"2019","unstructured":"Eran Goldman, Roei Herzig, Aviv Eisenschtat, Jacob Goldberger, and Tal Hassner. 2019. Precise detection in densely packed scenes. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5227\u20135236."},{"issue":"12","key":"e_1_3_2_38_2","doi-asserted-by":"crossref","first-page":"2861","DOI":"10.3390\/rs14122861","article-title":"Swin-transformer-enabled YOLOv5 with attention mechanism for small object detection on satellite images","volume":"14","author":"Gong Hang","year":"2022","unstructured":"Hang Gong, Tingkui Mu, Qiuxia Li, Haishan Dai, Chunlai Li, Zhiping He, Wenjing Wang, Feng Han, Abudusalamu Tuniyazi, Haoyang Li, et\u00a0al. 2022. Swin-transformer-enabled YOLOv5 with attention mechanism for small object detection on satellite images. Remote Sensing 14, 12 (2022), 2861.","journal-title":"Remote Sensing"},{"key":"e_1_3_2_39_2","first-page":"2786","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Han Jiaming","year":"2021","unstructured":"Jiaming Han, Jian Ding, Nan Xue, and Gui-Song Xia. 2021. Redet: A rotation-equivariant detector for aerial object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2786\u20132795."},{"key":"e_1_3_2_40_2","volume-title":"Proceedings of the 33rd British Machine Vision Conference","author":"Hashmi Khurram Azeem","year":"2022","unstructured":"Khurram Azeem Hashmi, Didier Stricker, and Muhammamd Zeshan Afzal. 2022. Spatio-temporal learnable proposals for end-to-end video object detection. In Proceedings of the 33rd British Machine Vision Conference."},{"issue":"9","key":"e_1_3_2_41_2","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","year":"2015","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Spatial pyramid pooling in deep convolutional networks for visual recognition. TPAMI 37, 9 (2015), 1904\u20131916.","journal-title":"TPAMI"},{"key":"e_1_3_2_42_2","first-page":"2961","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV)","author":"Girshick Kaiming He, Georgia Gkioxari, Piotr Doll\u00e1r, and Ross","year":"2017","unstructured":"Kaiming He, Georgia Gkioxari, Piotr Doll\u00e1r, and Ross Girshick. 2017. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2961\u20132969."},{"key":"e_1_3_2_43_2","first-page":"9377","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He Liqiang","year":"2022","unstructured":"Liqiang He and Sinisa Todorovic. 2022. DESTR: Object detection with split transformer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9377\u20139386."},{"key":"e_1_3_2_44_2","first-page":"1507","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"He Lu","year":"2021","unstructured":"Lu He, Qianyu Zhou, Xiangtai Li, Li Niu, Guangliang Cheng, Xiao Li, Wenxuan Liu, Yunhai Tong, Lizhuang Ma, and Liqing Zhang. 2021. End-to-end video object detection with spatial-temporal transformers. In Proceedings of the 29th ACM International Conference on Multimedia. 1507\u20131516."},{"key":"e_1_3_2_45_2","first-page":"31","article-title":"The human brain in numbers: A linearly scaled-up primate brain","author":"Herculano-Houzel Suzana","year":"2009","unstructured":"Suzana Herculano-Houzel. 2009. The human brain in numbers: A linearly scaled-up primate brain. Frontiers in Human Neuroscience 3 (2009), 31.","journal-title":"Frontiers in Human Neuroscience"},{"key":"e_1_3_2_46_2","first-page":"4678","volume-title":"Proceedings of the 2022 26th International Conference on Pattern Recognition (ICPR)","author":"Isaac-Medina Brian KS","year":"2022","unstructured":"Brian KS Isaac-Medina, Chris G. Willcocks, and Toby P. Breckon. 2022. Multi-view vision transformers for object detection. In Proceedings of the 2022 26th International Conference on Pattern Recognition (ICPR). IEEE, 4678\u20134684."},{"key":"e_1_3_2_47_2","unstructured":"Glenn Jocher Alex Stoken Jirka Borovec NanoCode012 ChristopherSTAN Liu Changyu Laughing Adam Hogan lorenzomammana tkianai et\u00a0al. 2020. yolov5. Code repository. Retrieved Sep. 2023 from https:\/\/github.com\/ultralytics\/yolov5. (2020)."},{"key":"e_1_3_2_48_2","unstructured":"Sehoon Kim Coleman Hooper Thanakul Wattanawong Minwoo Kang Ruohan Yan Hasan Genc Grace Dinh Qijing Huang Kurt Keutzer Michael W. Mahoney Yakun Sophia Shao and Amir Gholami. 2023. Full stack optimization of transformer inference: A survey. Retrieved from https:\/\/arxiv.org\/abs\/2302.14017"},{"key":"e_1_3_2_49_2","first-page":"734","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Law Hei","year":"2018","unstructured":"Hei Law and Jia Deng. 2018. Cornernet: Detecting objects as paired keypoints. In Proceedings of the European Conference on Computer Vision (ECCV). 734\u2013750."},{"key":"e_1_3_2_50_2","unstructured":"Chuyi Li Lulu Li Hongliang Jiang Kaiheng Weng Yifei Geng Liang Li Zaidan Ke Qingyuan Li Meng Cheng Weiqiang Nie Yiduo Li Bo Zhang Yufei Liang Linyuan Zhou Xiaoming Xu Xiangxiang Chu Xiaoming Wei and Xiaolin Wei. 2022. YOLOv6: A single-stage object detection framework for industrial applications. arXiv:2209.02976. Retrieved from https:\/\/arxiv.org\/abs\/2209.02976"},{"key":"e_1_3_2_51_2","first-page":"13619","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Feng","year":"2022","unstructured":"Feng Li, Hao Zhang, Shilong Liu, Jian Guo, Lionel M. Ni, and Lei Zhang. 2022. DN-DETR: Accelerate DETR training by introducing query denoising. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 13619\u201313627."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3615862"},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","first-page":"296","DOI":"10.1016\/j.isprsjprs.2019.11.023","article-title":"Object detection in optical remote sensing images: A survey and a new benchmark","volume":"159","author":"Li Ke","year":"2020","unstructured":"Ke Li, Gang Wan, Gong Cheng, Liqiu Meng, and Junwei Han. 2020. Object detection in optical remote sensing images: A survey and a new benchmark. ISPRS Journal of Photogrammetry and Remote Sensing 159 (2020), 296\u2013307.","journal-title":"ISPRS Journal of Photogrammetry and Remote Sensing"},{"issue":"4","key":"e_1_3_2_54_2","doi-asserted-by":"crossref","first-page":"984","DOI":"10.3390\/rs14040984","article-title":"Transformer with transfer CNN for remote-sensing-image object detection","volume":"14","author":"Li Qingyun","year":"2022","unstructured":"Qingyun Li, Yushi Chen, and Ying Zeng. 2022. Transformer with transfer CNN for remote-sensing-image object detection. Remote Sensing 14, 4 (2022), 984.","journal-title":"Remote Sensing"},{"issue":"18","key":"e_1_3_2_55_2","doi-asserted-by":"crossref","first-page":"6939","DOI":"10.3390\/s22186939","article-title":"Ghostformer: A GhostNet-based two-stage transformer for small object detection","volume":"22","author":"Li Sijia","year":"2022","unstructured":"Sijia Li, Furkat Sultonov, Jamshid Tursunboev, Jun-Hyun Park, Sangseok Yun, and Jae-Mo Kang. 2022. Ghostformer: A GhostNet-based two-stage transformer for small object detection. Sensors 22, 18 (2022), 6939.","journal-title":"Sensors"},{"key":"e_1_3_2_56_2","first-page":"6054","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li Yanghao","year":"2019","unstructured":"Yanghao Li, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang. 2019. Scale-aware trident networks for object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 6054\u20136063."},{"key":"e_1_3_2_57_2","first-page":"280","volume-title":"ECCV 2022: Proceedings of the 17th European Conference on Computer Vision, Part IX","author":"Li Yanghao","year":"2022","unstructured":"Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. 2022. Exploring plain vision transformer backbones for object detection. In ECCV 2022: Proceedings of the 17th European Conference on Computer Vision, Part IX. Springer, 280\u2013296."},{"key":"e_1_3_2_58_2","doi-asserted-by":"crossref","first-page":"6893","DOI":"10.1109\/TIP.2022.3216771","article-title":"Cbnet: A composite backbone network architecture for object detection","volume":"31","author":"Liang Tingting","year":"2022","unstructured":"Tingting Liang, Xiaojie Chu, Yudong Liu, Yongtao Wang, Zhi Tang, Wei Chu, Jingdong Chen, and Haibin Ling. 2022. Cbnet: A composite backbone network architecture for object detection. IEEE Transactions on Image Processing 31 (2022), 6893\u20136906.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_2_59_2","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1016\/j.aiopen.2022.10.001","article-title":"A survey of transformers","volume":"3","author":"Lin Tianyang","year":"2022","unstructured":"Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu. 2022. A survey of transformers. AI Open 3 (2022), 111\u2013132.","journal-title":"AI Open"},{"key":"e_1_3_2_60_2","first-page":"2117","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","year":"2017","unstructured":"Tsung-Yi Lin, Piotr Doll\u00e1r, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2117\u20132125."},{"key":"e_1_3_2_61_2","first-page":"2980","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV)","year":"2017","unstructured":"Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Doll\u00e1r. 2017. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2980\u20132988."},{"key":"e_1_3_2_62_2","first-page":"740","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision. Springer, 740\u2013755."},{"key":"e_1_3_2_63_2","first-page":"2588","volume-title":"ICASSP 2020-Proceedings of the 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Lin Wei-Hong","year":"2020","unstructured":"Wei-Hong Lin, Jia-Xing Zhong, Shan Liu, Thomas Li, and Ge Li. 2020. RoIMix: Proposal-fusion among multiple images for underwater object detection. In ICASSP 2020-Proceedings of the 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2588\u20132592."},{"key":"e_1_3_2_64_2","first-page":"954","volume-title":"Proceedings of the 2021 IEEE International Conference on Unmanned Systems (ICUS)","author":"Liu Chang","year":"2021","unstructured":"Chang Liu, Sheng Xu, and Baochang Zhang. 2021. Aerial small object tracking with transformers. In Proceedings of the 2021 IEEE International Conference on Unmanned Systems (ICUS). IEEE, 954\u2013959."},{"key":"e_1_3_2_65_2","volume-title":"ICLR","author":"Liu Shilong","year":"2023","unstructured":"Shilong Liu, Feng Li, Hao Zhang, Xiao Yang, Xianbiao Qi, Hang Su, Jun Zhu, and Lei Zhang. 2023. DAB-DETR: Dynamic anchor boxes are better queries for DETR. In ICLR."},{"issue":"12","key":"e_1_3_2_66_2","doi-asserted-by":"crossref","first-page":"9909","DOI":"10.1109\/TIE.2019.2893843","article-title":"Concealed object detection for activate millimeter wave image","volume":"66","author":"Liu Ting","year":"2019","unstructured":"Ting Liu, Yao Zhao, Yunchao Wei, Yufeng Zhao, and Shikui Wei. 2019. Concealed object detection for activate millimeter wave image. IEEE Transactions on Industrial Electronics 66, 12 (2019), 9909\u20139917.","journal-title":"IEEE Transactions on Industrial Electronics"},{"key":"e_1_3_2_67_2","first-page":"21","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","year":"2016","unstructured":"Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg. 2016. SSD: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, 21\u201337."},{"key":"e_1_3_2_68_2","doi-asserted-by":"crossref","first-page":"114602","DOI":"10.1016\/j.eswa.2021.114602","article-title":"A survey and performance evaluation of deep learning methods for small object detection","volume":"172","author":"Liu Yang","year":"2021","unstructured":"Yang Liu, Peng Sun, Nickolas Wergeles, and Yi Shang. 2021. A survey and performance evaluation of deep learning methods for small object detection. Expert Systems with Applications 172 (2021), 114602.","journal-title":"Expert Systems with Applications"},{"key":"e_1_3_2_69_2","doi-asserted-by":"crossref","first-page":"57120","DOI":"10.1109\/ACCESS.2019.2913882","article-title":"MR-CNN: A multi-scale region-based convolutional neural network for small traffic sign recognition","volume":"7","author":"Liu Zhigang","year":"2019","unstructured":"Zhigang Liu, Juan Du, Feng Tian, and Jiazheng Wen. 2019. MR-CNN: A multi-scale region-based convolutional neural network for small traffic sign recognition. IEEE Access 7 (2019), 57120\u201357128.","journal-title":"IEEE Access"},{"key":"e_1_3_2_70_2","first-page":"10012","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Liu Ze","year":"2021","unstructured":"Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 10012\u201310022."},{"key":"e_1_3_2_71_2","article-title":"A CNN-transformer hybrid model based on CSWin transformer for UAV image object detection","author":"Lu Wanjie","year":"2023","unstructured":"Wanjie Lu, Chaozhen Lan, Chaoyang Niu, Wei Liu, Liang Lyu, Qunshan Shi, and Shiju Wang. 2023. A CNN-transformer hybrid model based on CSWin transformer for UAV image object detection. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 16 (2023), 1211\u20131231.","journal-title":"IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing"},{"key":"e_1_3_2_72_2","unstructured":"Teli Ma Mingyuan Mao Honghui Zheng Peng Gao Xiaodi Wang Shumin Han Errui Ding Baochang Zhang and David Doermann. 2021. Oriented object detection with transformer. Retrieved from https:\/\/arxiv.org\/abs\/2106.03146"},{"key":"e_1_3_2_73_2","volume-title":"Proceedings of the 17th European Conference on Computer Vision (ECCV)","author":"Maaz Muhammad","year":"2022","unstructured":"Muhammad Maaz, Hanoona Rasheed, Salman Khan, Fahad Shahbaz Khan, Rao Muhammad Anwer, and Ming-Hsuan Yang. 2022. Class-agnostic Object detection with multi-modal transformer. In Proceedings of the 17th European Conference on Computer Vision (ECCV). Springer."},{"key":"e_1_3_2_74_2","first-page":"8844","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Meinhardt Tim","year":"2022","unstructured":"Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer. 2022. Trackformer: Multi-object tracking with transformers. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8844\u20138854."},{"key":"e_1_3_2_75_2","first-page":"3651","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Meng Depu","year":"2021","unstructured":"Depu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng, Houqiang Li, Yuhui Yuan, Lei Sun, and Jingdong Wang. 2021. Conditional DETR for fast training convergence. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 3651\u20133660."},{"key":"e_1_3_2_76_2","doi-asserted-by":"crossref","first-page":"445","DOI":"10.1007\/978-3-319-46448-0_27","volume-title":"ECCV 2016: Proceedings of the 14th European Conference on Computer Vision, Part I 14","author":"Mueller Matthias","year":"2016","unstructured":"Matthias Mueller, Neil Smith, and Bernard Ghanem. 2016. A benchmark and simulator for UAV tracking. In ECCV 2016: Proceedings of the 14th European Conference on Computer Vision, Part I 14. Springer, 445\u2013461."},{"issue":"10","key":"e_1_3_2_77_2","doi-asserted-by":"crossref","first-page":"3388","DOI":"10.1109\/TPAMI.2020.2981890","article-title":"Imbalance problems in object detection: A review","volume":"43","author":"Oksuz Kemal","year":"2020","unstructured":"Kemal Oksuz, Baris Can Cam, Sinan Kalkan, and Emre Akbas. 2020. Imbalance problems in object detection: A review. IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 10 (2020), 3388\u20133415.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_78_2","unstructured":"Jeffrey Ouyang-Zhang Jang Hyun Cho Xingyi Zhou and Philipp Kr\u00e4henb\u00fchl. 2022. NMS strikes back. Retrieved from https:\/\/arxiv.org\/abs\/2212.06137"},{"key":"e_1_3_2_79_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Unel F. Ozge","year":"2019","unstructured":"F. Ozge Unel, Burak O. Ozkalayci, and Cevahir Cigla. 2019. The power of tiling for small object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops."},{"key":"e_1_3_2_80_2","first-page":"821","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","year":"2019","unstructured":"Jiangmiao Pang, Kai Chen, Jianping Shi, Huajun Feng, Wanli Ouyang, and Dahua Lin. 2019. Libra R-CNN: Towards balanced learning for object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 821\u2013830."},{"key":"e_1_3_2_81_2","article-title":"Conformer: Local features coupling global representations for recognition and detection","author":"Peng Zhiliang","year":"2023","unstructured":"Zhiliang Peng, Zonghao Guo, Wei Huang, Yaowei Wang, Lingxi Xie, Jianbin Jiao, Qi Tian, and Qixiang Ye. 2023. Conformer: Local features coupling global representations for recognition and detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 8 (2023), 9454\u20139468.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_82_2","first-page":"367","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Peng Zhiliang","year":"2021","unstructured":"Zhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie, Yaowei Wang, Jianbin Jiao, and Qixiang Ye. 2021. Conformer: Local features coupling global representations for visual recognition. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 367\u2013376."},{"key":"e_1_3_2_83_2","first-page":"9288","article-title":"Optimal visual search based on a model of target detectability in natural images","volume":"33","author":"Rashidi Shima","year":"2020","unstructured":"Shima Rashidi, Krista Ehinger, Andrew Turpin, and Lars Kulik. 2020. Optimal visual search based on a model of target detectability in natural images. Advances in Neural Information Processing Systems 33 (2020), 9288\u20139299.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_84_2","article-title":"IT-RUDA: Information theory assisted robust unsupervised domain adaptation","author":"Rashidi Shima","year":"2025","unstructured":"Shima Rashidi, Ruwan Tennakoon, Aref Miri Rekavandi, Papangkorn Jessadatavornwong, Amanda Freis, Garret Huff, Mark Easton, Adrian Mouritz, Reza Hoseinnezhad, and Alireza Bab-Hadiashar. 2025. IT-RUDA: Information theory assisted robust unsupervised domain adaptation. ACM Transactions on Intelligent Systems and Technology (2025).","journal-title":"ACM Transactions on Intelligent Systems and Technology"},{"key":"e_1_3_2_85_2","first-page":"779","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Redmon Joseph","year":"2016","unstructured":"Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2016. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 779\u2013788."},{"key":"e_1_3_2_86_2","first-page":"7263","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Redmon Joseph","year":"2017","unstructured":"Joseph Redmon and Ali Farhadi. 2017. YOLO9000: Better, faster, stronger. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 7263\u20137271."},{"key":"e_1_3_2_87_2","unstructured":"Joseph Redmon and Ali Farhadi. 2018. Yolov3: An incremental improvement. Retrieved from https:\/\/arxiv.org\/abs\/1804.02767"},{"key":"e_1_3_2_88_2","doi-asserted-by":"crossref","first-page":"5017","DOI":"10.1109\/TIP.2021.3077139","article-title":"Robust subspace detectors based on  \\(\\alpha\\) -divergence with application to detection in imaging","volume":"30","author":"Rekavandi Aref Miri","year":"2021","unstructured":"Aref Miri Rekavandi, Abd-Krim Seghouane, and Robin J. Evans. 2021. Robust subspace detectors based on \\(\\alpha\\) -divergence with application to detection in imaging. IEEE Transactions on Image Processing 30 (2021), 5017\u20135031.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_2_89_2","article-title":"A guide to image-and video-based small object detection using deep learning: Case study of maritime surveillance","author":"Rekavandi Aref Miri","year":"2025","unstructured":"Aref Miri Rekavandi, Lian Xu, Farid Boussaid, Abd-Krim Seghouane, Stephen Hoefs, and Mohammed Bennamoun. 2025. A guide to image-and video-based small object detection using deep learning: Case study of maritime surveillance. IEEE Transactions on Intelligent Transportation Systems 26, 3 (2025), 2851\u20132879.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"e_1_3_2_90_2","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the 29th International Conference on Neural Information Processing Systems (NeurIPS) ."},{"key":"e_1_3_2_91_2","doi-asserted-by":"crossref","first-page":"93453","DOI":"10.1109\/ACCESS.2022.3203399","article-title":"DAFA: Diversity-aware feature aggregation for attention-based video object detection","volume":"10","author":"Roh Si-Dong","year":"2022","unstructured":"Si-Dong Roh and Ki-Seok Chung. 2022. DAFA: Diversity-aware feature aggregation for attention-based video object detection. IEEE Access 10 (2022), 93453\u201393463.","journal-title":"IEEE Access"},{"key":"e_1_3_2_92_2","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"Imagenet large scale visual recognition challenge","volume":"115","author":"Russakovsky Olga","year":"2015","unstructured":"Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et\u00a0al. 2015. Imagenet large scale visual recognition challenge. International Journal of Computer Vision 115 (2015), 211\u2013252.","journal-title":"International Journal of Computer Vision"},{"key":"e_1_3_2_93_2","first-page":"1032","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Shen Ruoyue","year":"2023","unstructured":"Ruoyue Shen, Nakamasa Inoue, and Koichi Shinoda. 2023. Text-guided object detector for multi-modal video question answering. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 1032\u20131042."},{"key":"e_1_3_2_94_2","article-title":"Object detection in medical images based on hierarchical transformer and mask mechanism","volume":"2022","author":"Shou Yuntao","year":"2022","unstructured":"Yuntao Shou, Tao Meng, Wei Ai, Canhao Xie, Haiyan Liu, and Yina Wang. 2022. Object detection in medical images based on hierarchical transformer and mask mechanism. Computational Intelligence and Neuroscience 2022 (2022), 1\u201312.","journal-title":"Computational Intelligence and Neuroscience"},{"key":"e_1_3_2_95_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Song Hwanjun","year":"2022","unstructured":"Hwanjun Song, Deqing Sun, Sanghyuk Chun, Varun Jampani, Dongyoon Han, Byeongho Heo, Wonjae Kim, and Ming-Hsuan Yang. 2022. ViDT: An efficient and effective fully transformer-based object detector. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_96_2","first-page":"2325","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Stewart Russell","year":"2016","unstructured":"Russell Stewart, Mykhaylo Andriluka, and Andrew Y. Ng. 2016. End-to-end people detection in crowded scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2325\u20132333."},{"issue":"9","key":"e_1_3_2_97_2","doi-asserted-by":"crossref","first-page":"6148","DOI":"10.1109\/TCSVT.2022.3161815","article-title":"Multi-source aggregation transformer for concealed object detection in millimeter-wave images","volume":"32","author":"Sun Peng","year":"2022","unstructured":"Peng Sun, Ting Liu, Xiaotong Chen, Shiyin Zhang, Yao Zhao, and Shikui Wei. 2022. Multi-source aggregation transformer for concealed object detection in millimeter-wave images. IEEE Transactions on Circuits and Systems for Video Technology 32, 9 (2022), 6148\u20136159.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_98_2","first-page":"3611","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Sun Zhiqing","year":"2021","unstructured":"Zhiqing Sun, Shengcao Cao, Yiming Yang, and Kris M. Kitani. 2021. Rethinking transformer-based set prediction for object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 3611\u20133620."},{"key":"e_1_3_2_99_2","doi-asserted-by":"crossref","unstructured":"Yudi Tang Bing Wang Wangli He and Feng Qian. 2023. Pointdet++: An object detection framework based on human local features with transformer encoder. Neural Comput & Applic 35 (2023) 10097\u201310108.","DOI":"10.1007\/s00521-022-06938-7"},{"key":"e_1_3_2_100_2","first-page":"9627","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Tian Zhi","year":"2019","unstructured":"Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. 2019. Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 9627\u20139636."},{"key":"e_1_3_2_101_2","first-page":"10347","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Touvron Hugo","year":"2021","unstructured":"Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv\u00e9 J\u00e9gou. 2021. Training data-efficient image transformers & distillation through attention. In Proceedings of the International Conference on Machine Learning. PMLR, 10347\u201310357."},{"key":"e_1_3_2_102_2","first-page":"2768","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Varga Leon Amadeus","year":"2021","unstructured":"Leon Amadeus Varga and Andreas Zell. 2021. Tackling the background bias in sparse object detection via cropped windows. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 2768\u20132777."},{"key":"e_1_3_2_103_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017), 1\u201311.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_104_2","first-page":"7464","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Chien-Yao","year":"2023","unstructured":"Chien-Yao Wang, Alexey Bochkovskiy, and Hong-Yuan Mark Liao. 2023. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 7464\u20137475."},{"key":"e_1_3_2_105_2","doi-asserted-by":"crossref","first-page":"732","DOI":"10.1007\/978-3-031-20074-8_42","volume-title":"ECCV 2022: Proceedings of the 17th European Conference on Computer Vision, Part VIII","author":"Wang Han","year":"2022","unstructured":"Han Wang, Jun Tang, Xiaodong Liu, Shanyan Guan, Rong Xie, and Li Song. 2022. PTSEFormer: Progressive temporal-spatial enhanced transformer towards video object detection. In ECCV 2022: Proceedings of the 17th European Conference on Computer Vision, Part VIII. Springer, 732\u2013747."},{"key":"e_1_3_2_106_2","unstructured":"Jinwang Wang Chang Xu Wen Yang and Lei Yu. 2021. A normalized Gaussian Wasserstein distance for tiny object detection. Retrieved from https:\/\/arxiv.org\/abs\/2110.13389"},{"key":"e_1_3_2_107_2","first-page":"6450","volume-title":"IGARSS 2023-Proceedings of the 2023 IEEE International Geoscience and Remote Sensing Symposium","author":"Wang Liya","year":"2023","unstructured":"Liya Wang and Alex Tien. 2023. Aerial image object detection with vision transformer detector (ViTDet). In IGARSS 2023-Proceedings of the 2023 IEEE International Geoscience and Remote Sensing Symposium. IEEE, 6450\u20136453."},{"key":"e_1_3_2_108_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Wang Wen","year":"2022","unstructured":"Wen Wang, Yang Cao, Jing Zhang, and Dacheng Tao. 2022. Fp-detr: Detection transformer advanced by fully pre-training. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_109_2","article-title":"Anchor DETR: Query design for transformer-based object detection","author":"Wang Yingming","year":"2022","unstructured":"Yingming Wang, Xiangyu Zhang, Tong Yang, and Jian Sun. 2022. Anchor DETR: Query design for transformer-based object detection. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) .","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI)"},{"key":"e_1_3_2_110_2","first-page":"9217","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Wu Haiping","year":"2019","unstructured":"Haiping Wu, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang. 2019. Sequence level semantics aggregation for video object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 9217\u20139225."},{"key":"e_1_3_2_111_2","first-page":"2012","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"Wu Jialian","year":"2020","unstructured":"Jialian Wu, Chunluan Zhou, Qian Zhang, Ming Yang, and Junsong Yuan. 2020. Self-mimic learning for small-scale pedestrian detection. In Proceedings of the 28th ACM International Conference on Multimedia. 2012\u20132020."},{"key":"e_1_3_2_112_2","first-page":"3974","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Xia Gui-Song","year":"2018","unstructured":"Gui-Song Xia, Xiang Bai, Jian Ding, Zhen Zhu, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liangpei Zhang. 2018. DOTA: A large-scale dataset for object detection in aerial images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3974\u20133983."},{"issue":"15","key":"e_1_3_2_113_2","doi-asserted-by":"crossref","first-page":"12571","DOI":"10.1007\/s00521-022-07333-y","article-title":"Transformers only look once with nonlinear combination for real-time object detection","volume":"34","author":"Xia Ruiyang","year":"2022","unstructured":"Ruiyang Xia, Guoquan Li, Zhengwen Huang, Yu Pang, and Man Qi. 2022. Transformers only look once with nonlinear combination for real-time object detection. Neural Computing and Applications 34, 15 (2022), 12571\u201312585.","journal-title":"Neural Computing and Applications"},{"key":"e_1_3_2_114_2","article-title":"DKTNet: Dual-key transformer network for small object detection","author":"Xu Shoukun","year":"2023","unstructured":"Shoukun Xu, Jianan Gu, Yining Hua, and Yi Liu. 2023. DKTNet: Dual-key transformer network for small object detection. Neurocomputing 525 (2023), 29\u201341.","journal-title":"Neurocomputing"},{"issue":"18","key":"e_1_3_2_115_2","doi-asserted-by":"crossref","first-page":"6993","DOI":"10.3390\/s22186993","article-title":"FEA-swin: Foreground enhancement attention swin transformer network for accurate UAV-based dense object detection","volume":"22","author":"Xu Wenyu","year":"2022","unstructured":"Wenyu Xu, Chaofan Zhang, Qi Wang, and Pangda Dai. 2022. FEA-swin: Foreground enhancement attention swin transformer network for accurate UAV-based dense object detection. Sensors 22, 18 (2022), 6993.","journal-title":"Sensors"},{"issue":"23","key":"e_1_3_2_116_2","doi-asserted-by":"crossref","first-page":"4779","DOI":"10.3390\/rs13234779","article-title":"An improved swin transformer-based model for remote sensing object detection and instance segmentation","volume":"13","author":"Xu Xiangkai","year":"2021","unstructured":"Xiangkai Xu, Zhejun Feng, Changqing Cao, Mengyuan Li, Jin Wu, Zengyan Wu, Yajie Shang, and Shubing Ye. 2021. An improved swin transformer-based model for remote sensing object detection and instance segmentation. Remote Sensing 13, 23 (2021), 4779.","journal-title":"Remote Sensing"},{"key":"e_1_3_2_117_2","doi-asserted-by":"crossref","first-page":"6856","DOI":"10.1109\/JSTARS.2022.3198577","article-title":"Dual network structure with interweaved global-local feature hierarchy for transformer-based object detection in remote sensing image","volume":"15","author":"Xue Jingqian","year":"2022","unstructured":"Jingqian Xue, Da He, Mengwei Liu, and Qian Shi. 2022. Dual network structure with interweaved global-local feature hierarchy for transformer-based object detection in remote sensing image. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 15 (2022), 6856\u20136866.","journal-title":"IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing"},{"issue":"3","key":"e_1_3_2_118_2","first-page":"036501","article-title":"DeepLesion: Automated mining of large-scale lesion annotations and universal lesion detection with deep learning","volume":"5","author":"Yan Ke","year":"2018","unstructured":"Ke Yan, Xiaosong Wang, Le Lu, and Ronald M. Summers. 2018. DeepLesion: Automated mining of large-scale lesion annotations and universal lesion detection with deep learning. Journal of Medical Imaging 5, 3 (2018), 036501\u2013036501.","journal-title":"Journal of Medical Imaging"},{"key":"e_1_3_2_119_2","doi-asserted-by":"crossref","DOI":"10.1109\/TCSVT.2023.3234311","article-title":"Unifying convolution and transformer for efficient concealed object detection in passive millimeter-wave images","author":"Yang Hao","year":"2023","unstructured":"Hao Yang, Zihan Yang, Anyong Hu, Che Liu, Tie Jun Cui, and Jungang Miao. 2023. Unifying convolution and transformer for efficient concealed object detection in passive millimeter-wave images. IEEE Transactions on Circuits and Systems for Video Technology 33, 8 (2023), 3872\u20133887.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_120_2","first-page":"9657","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yang Ze","year":"2019","unstructured":"Ze Yang, Shaohui Liu, Han Hu, Liwei Wang, and Stephen Lin. 2019. Reppoints: Point set representation for object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 9657\u20139666."},{"key":"e_1_3_2_121_2","first-page":"1","article-title":"Real-time object detection network in UAV-vision based on CNN and transformer","volume":"72","author":"Ye Tao","year":"2023","unstructured":"Tao Ye, Wenyang Qin, Zongyang Zhao, Xiaozhi Gao, Xiangpeng Deng, and Yu Ouyang. 2023. Real-time object detection network in UAV-vision based on CNN and transformer. IEEE Transactions on Instrumentation and Measurement 72 (2023), 1\u201313.","journal-title":"IEEE Transactions on Instrumentation and Measurement"},{"key":"e_1_3_2_122_2","first-page":"1","volume-title":"Proceedings of the 2018 7th Mediterranean Conference on Embedded Computing (MECO)","author":"Yudin Dmitry","year":"2018","unstructured":"Dmitry Yudin and Dmitry Slavioglo. 2018. Usage of fully convolutional network with clustering for traffic light detection. In Proceedings of the 2018 7th Mediterranean Conference on Embedded Computing (MECO). IEEE, 1\u20136."},{"key":"e_1_3_2_123_2","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1016\/j.neucom.2022.04.062","article-title":"NLFFTNet: A non-local feature fusion transformer network for multi-scale object detection","volume":"493","author":"Zeng Kai","year":"2022","unstructured":"Kai Zeng, Qian Ma, Jiawen Wu, Sijia Xiang, Tao Shen, and Lei Zhang. 2022. NLFFTNet: A non-local feature fusion transformer network for multi-scale object detection. Neurocomputing 493 (2022), 15\u201327.","journal-title":"Neurocomputing"},{"key":"e_1_3_2_124_2","unstructured":"Chi Zhang Lijuan Liu Xiaoxue Zang Frederick Liu Hao Zhang Xinying Song and Jindong Chen. 2022. DETR++: Taming your multi-scale detection transformer. Retrieved from https:\/\/arxiv.org\/abs\/2206.02977"},{"key":"e_1_3_2_125_2","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1007\/978-3-030-58555-6_16","volume-title":"ECCV 2020: Proceedings of the 16th European Conference on Computer Vision, Part XV 16","author":"Zhang Hongkai","year":"2020","unstructured":"Hongkai Zhang, Hong Chang, Bingpeng Ma, Naiyan Wang, and Xilin Chen. 2020. Dynamic R-CNN: Towards high quality object detection via dynamic training. In ECCV 2020: Proceedings of the 16th European Conference on Computer Vision, Part XV 16. Springer, 260\u2013275."},{"key":"e_1_3_2_126_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Zhang Hao","unstructured":"Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel Ni, and Heung-Yeung Shum. 2023. DINO: DETR with improved denoising anchor boxes for end-to-end object detection. In Proceedings of the 11th International Conference on Learning Representations."},{"key":"e_1_3_2_127_2","first-page":"9759","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Shifeng","year":"2020","unstructured":"Shifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei, and Stan Z. Li. 2020. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9759\u20139768."},{"issue":"8","key":"e_1_3_2_128_2","doi-asserted-by":"crossref","first-page":"5535","DOI":"10.1109\/TGRS.2019.2900302","article-title":"Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection","volume":"57","author":"Zhang Yuanlin","year":"2019","unstructured":"Yuanlin Zhang, Yuan Yuan, Yachuang Feng, and Xiaoqiang Lu. 2019. Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection. IEEE Transactions on Geoscience and Remote Sensing 57, 8 (2019), 5535\u20135548.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_2_129_2","article-title":"TransVOD: End-to-end video object detection with spatial-temporal transformers","author":"Zhou Qianyu","year":"2023","unstructured":"Qianyu Zhou, Xiangtai Li, Lu He, Yibo Yang, Guangliang Cheng, Yunhai Tong, Lizhuang Ma, and Dacheng Tao. 2023. TransVOD: End-to-end video object detection with spatial-temporal transformers. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 6 (2023), 7853\u20137869.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_130_2","unstructured":"Xingyi Zhou Dequan Wang and Philipp Kr\u00e4henb\u00fchl. 2019. Objects as points. Retrieved from https:\/\/arxiv.org\/abs\/1904.07850"},{"key":"e_1_3_2_131_2","article-title":"Deformable DETR: Deformable transformers for end-to-end object detection","author":"Zhu Xizhou","year":"2021","unstructured":"Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2021. Deformable DETR: Deformable transformers for end-to-end object detection. In Proceedings of the International Conference on Learning Representations (ICLR) .","journal-title":"Proceedings of the International Conference on Learning Representations (ICLR)"},{"issue":"1","key":"e_1_3_2_132_2","doi-asserted-by":"crossref","first-page":"2448","DOI":"10.1080\/09540091.2022.2125499","article-title":"SRDD: A lightweight end-to-end object detection with transformer","volume":"34","author":"Zhu Yuan","year":"2022","unstructured":"Yuan Zhu, Qingyuan Xia, and Wen Jin. 2022. SRDD: A lightweight end-to-end object detection with transformer. Connection Science 34, 1 (2022), 2448\u20132465.","journal-title":"Connection Science"},{"key":"e_1_3_2_133_2","first-page":"6748","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zong Zhuofan","year":"2023","unstructured":"Zhuofan Zong, Guanglu Song, and Yu Liu. 2023. DETRs with collaborative hybrid assignments training. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 6748\u20136758."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3758090","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,10]],"date-time":"2025-09-10T12:54:24Z","timestamp":1757508864000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3758090"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,10]]},"references-count":132,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3758090"],"URL":"https:\/\/doi.org\/10.1145\/3758090","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,10]]},"assertion":[{"value":"2024-12-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-24","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}