{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T17:51:37Z","timestamp":1785952297018,"version":"3.56.0"},"reference-count":40,"publisher":"Oxford University Press (OUP)","issue":"4","license":[{"start":{"date-parts":[[2026,4,2]],"date-time":"2026-04-02T00:00:00Z","timestamp":1775088000000},"content-version":"vor","delay-in-days":1,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100009558","name":"University Natural Science Research Project of Anhui Province","doi-asserted-by":"publisher","award":["No. 2024AH051852"],"award-info":[{"award-number":["No. 2024AH051852"]}],"id":[{"id":"10.13039\/501100009558","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Temporary Exercise Project of Tongling University","award":["No. 2025GZDLSJ17"],"award-info":[{"award-number":["No. 2025GZDLSJ17"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Small object detection in UAV (Unmanned Aerial Vehicle) aerial imagery faces significant challenges, including insufficient feature representation, severe background noise interference, and inadequate multi-scale fusion. To address these issues, this study proposes UMS-DET (Unmanned Aerial Vehicle Multi-scale Small-object Detector), a specialized framework engineered to optimize multi-scale representation and small object discrimination in UAV views. Firstly, we design the Multi-Scale Context-Gated Network backbone, which strengthens multi-scale feature representation through hierarchical context synergy aggregation and adaptive gated fusion mechanisms, all while reducing the parameter count. Subsequently, we introduce the Sparse Hierarchical Frequency Feature Interaction encoder. This module integrates sparse window attention with frequency-domain enhancement techniques to effectively suppress background noise and improve discriminative capability for small targets. Furthermore, we propose the Small Object-Focused Cross-scale Feature Fusion module, which leverages high-resolution shallow features to preserve fine-grained details, facilitating effective multi-scale integration. Experimental results demonstrate that UMS-DET achieves AP$_{50}$ improvements of 5.2% and 3.0% on the VisDrone and Dataset for Object deTection in Aerial images (DOTA) datasets, respectively. Notably, the model reduces parameters by 22.1% and achieves a real-time inference speed of 67.3 Frames Per Second. The proposed method exhibits superior accuracy and robustness in dense small object scenarios and complex environments, offering an efficient and practical solution for UAV intelligent perception systems.<\/jats:p>","DOI":"10.1093\/jcde\/qwag037","type":"journal-article","created":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T11:45:46Z","timestamp":1775043946000},"page":"290-306","source":"Crossref","is-referenced-by-count":1,"title":["UMS-DET: A frequency-enhanced transformer-based UAV multi-scale small object detector for aerial imagery"],"prefix":"10.1093","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-5591-2103","authenticated-orcid":false,"given":"Wei","family":"Hu","sequence":"first","affiliation":[{"name":"School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications , No. 10 Xitucheng Road, Haidian District, Beijing 100876 ,","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2769-5418","authenticated-orcid":false,"given":"Feixiang","family":"Du","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Computing and Science, Swinburne University of Technology , Sarawak Campus, Kuching 93350 ,","place":["Malaysia"]},{"name":"School of Electrical Engineering, Tongling University , No.1335 Cuihu 4th Road, Tongling, Anhui 244061 ,","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-9353-7658","authenticated-orcid":false,"given":"Xiaodong","family":"Qin","sequence":"additional","affiliation":[{"name":"School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications , No. 10 Xitucheng Road, Haidian District, Beijing 100876 ,","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joel C M","family":"Than","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Computing and Science, Swinburne University of Technology , Sarawak Campus, Kuching 93350 ,","place":["Malaysia"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2026,4,2]]},"reference":[{"key":"2026042709230520600_bib1","doi-asserted-by":"publisher","first-page":"82","DOI":"10.1093\/jcde\/qwaf080","article-title":"Wcformer: A wavelet-enhanced cnn-transformer hybrid network for bearing fault diagnosis using multi-sensor signal fusion","volume":"12","author":"Cao","year":"2025","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026042709230520600_bib2","first-page":"213","article-title":"End-to-end object detection with transformers","volume-title":"European Conference on Computer Vision","author":"Carion","year":"2020"},{"key":"2026042709230520600_bib3","first-page":"5513","article-title":"Unireplknet: A universal perception large-kernel convnet for audio video point cloud time-series and image recognition","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ding","year":"2024"},{"key":"2026042709230520600_bib4","first-page":"213","article-title":"Visdrone-det2019: The vision meets drone object detection in image challenge results","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops","author":"Du","year":"2019"},{"key":"2026042709230520600_bib5","doi-asserted-by":"publisher","first-page":"121366","DOI":"10.1016\/j.ins.2024.121366","article-title":"Lud-yolo: A novel lightweight object detection network for unmanned aerial vehicle","volume":"686","author":"Fan","year":"2025","journal-title":"Information Sciences"},{"key":"2026042709230520600_bib6","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1093\/jcde\/qwaf117","article-title":"Fault diagnosis of uav sensors based on multi-auxiliary task learning with few samples","volume":"12","author":"Fang","year":"2025","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026042709230520600_bib7","doi-asserted-by":"publisher","first-page":"97","DOI":"10.1093\/jcde\/qwag028","article-title":"Enhanced yolo11 for tiny object detection based on multiscale information interaction and fusion in uav aerial images","volume":"13","author":"Gao","year":"2026","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026042709230520600_bib8","first-page":"7132","article-title":"Squeeze-and-excitation networks","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Hu","year":"2018"},{"key":"2026042709230520600_bib9","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TIM.2024.3381272","article-title":"Mffsodnet: Multi-scale feature fusion small object detection network for uav aerial images","volume":"73","author":"Jiang","year":"2024","journal-title":"IEEE Transactions on Instrumentation and Measurement"},{"key":"2026042709230520600_bib10","doi-asserted-by":"publisher","first-page":"1066","DOI":"10.1016\/j.procs.2022.01.135","article-title":"A review of yolo algorithm developments","volume":"199","author":"Jiang","year":"2022","journal-title":"Procedia Computer Science"},{"key":"2026042709230520600_bib11","doi-asserted-by":"publisher","first-page":"5496","DOI":"10.3390\/s24175496","article-title":"Drone-detr: Efficient small object detection for remote sensing image using enhanced rt-detr model","volume":"24","author":"Kong","year":"2024","journal-title":"Sensors"},{"key":"2026042709230520600_bib12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1093\/jcde\/qwaf081","article-title":"Scribblesam: Weakly supervised salient object detection and localization in remote sensing images using transformer and segment anything model","volume":"12","author":"Lee","year":"2025","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026042709230520600_bib13","first-page":"13619","article-title":"Dn-detr: Accelerate detr training by introducing query denoising","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li","year":"2022"},{"key":"2026042709230520600_bib14","first-page":"18558","article-title":"Lite detr: An interleaved multi-scale encoder for efficient detr","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li","year":"2023"},{"key":"2026042709230520600_bib15","doi-asserted-by":"crossref","first-page":"1410","DOI":"10.1007\/s11263-024-02247-9","article-title":"Lsknet: A foundation lightweight backbone for remote sensing","volume":"133","author":"Li","year":"2025","journal-title":"International Journal of Computer Vision"},{"key":"2026042709230520600_bib16","first-page":"14420","article-title":"Efficientvit: Memory efficient vision transformer with cascaded group attention","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu","year":"2023"},{"key":"2026042709230520600_bib17","doi-asserted-by":"publisher","first-page":"1211","DOI":"10.1109\/JSTARS.2023.3234161","article-title":"A cnn-transformer hybrid model based on cswin transformer for uav image object detection","volume":"16","author":"Lu","year":"2023","journal-title":"IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing"},{"key":"2026042709230520600_bib18","article-title":"Polaformer: Polarity-aware linear attention for vision transformers","author":"Meng","year":"2025","journal-title":"The Thirteenth International Conference on Learning Representations"},{"key":"2026042709230520600_bib19","doi-asserted-by":"publisher","first-page":"108411","DOI":"10.1016\/j.patcog.2021.108411","article-title":"Visual vs internal attention mechanisms in deep neural networks for image classification and object detection","volume":"123","author":"Obeso","year":"2022","journal-title":"Pattern Recognition"},{"key":"2026042709230520600_bib20","first-page":"14541","article-title":"Fast vision transformers with hilo attention","volume":"35","author":"Pan","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026042709230520600_bib21","first-page":"91","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren","year":"2015","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026042709230520600_bib22","first-page":"111","article-title":"Restoring images in adverse weather conditions via histogram transformer","volume-title":"European Conference on Computer Vision","author":"Sun","year":"2024"},{"key":"2026042709230520600_bib23","first-page":"443","article-title":"No more strided convolutions or pooling: A new cnn building block for low-resolution images and small objects","volume-title":"Joint European Conference on Machine Learning and Knowledge Discovery in Databases","author":"Sunkara","year":"2022"},{"key":"2026042709230520600_bib24","first-page":"15909","article-title":"Repvit: Revisiting mobile cnn from vit perspective","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang","year":"2024"},{"key":"2026042709230520600_bib25","first-page":"390","article-title":"Cspnet: A new backbone that can enhance learning capability of cnn","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Wang","year":"2020"},{"key":"2026042709230520600_bib26","doi-asserted-by":"publisher","first-page":"3123","DOI":"10.1109\/TPAMI.2023.3341806","article-title":"Crossformer++: A versatile vision transformer hinging on cross-scale attention","volume":"46","author":"Wang","year":"2023","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026042709230520600_bib27","doi-asserted-by":"publisher","first-page":"7376","DOI":"10.3390\/s24227376","article-title":"Dv-detr: Improved uav aerial small target detection algorithm based on rt-detr","volume":"24","author":"Wei","year":"2024","journal-title":"Sensors"},{"key":"2026042709230520600_bib28","first-page":"3","article-title":"Cbam: Convolutional block attention module","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Woo","year":"2018"},{"key":"2026042709230520600_bib29","article-title":"Token statistics transformer: Linear-time attention via variational rate reduction","author":"Wu","year":"2025","journal-title":"The Thirteenth International Conference on Learning Representations"},{"key":"2026042709230520600_bib30","first-page":"3974","article-title":"Dota: A large-scale dataset for object detection in aerial images","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Xia","year":"2018"},{"key":"2026042709230520600_bib31","doi-asserted-by":"publisher","first-page":"223","DOI":"10.1007\/s44196-024-00632-3","article-title":"Small object detection in uav images based on yolov8n","volume":"17","author":"Xu","year":"2024","journal-title":"International Journal of Computational Intelligence Systems"},{"key":"2026042709230520600_bib32","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TIM.2022.3162274","article-title":"Dense and small object detection in uav-vision based on a global-local feature enhanced network","volume":"71","author":"Ye","year":"2022","journal-title":"IEEE Transactions on Instrumentation and Measurement"},{"key":"2026042709230520600_bib33","first-page":"4484","article-title":"Mambaout: Do we really need mamba for vision?","volume-title":"Proceedings of the Computer Vision and Pattern Recognition Conference","author":"Yu","year":"2025"},{"key":"2026042709230520600_bib34","article-title":"Rt-detr++ for uav object detection","author":"Yuan","year":"2025","journal-title":"arXiv preprint arXiv:2509.09157"},{"key":"2026042709230520600_bib35","first-page":"6317","article-title":"Small object detection via coarse-to-fine proposal generation and imitation learning","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yuan","year":"2023"},{"key":"2026042709230520600_bib36","doi-asserted-by":"crossref","first-page":"5966","DOI":"10.1109\/TCSVT.2025.3532243","article-title":"Semantic differentiation aids oriented small object detection","volume":"35","author":"Yuan","year":"2025","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"2026042709230520600_bib37","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1093\/jcde\/qwaf121","article-title":"Pgem-detr: Physics-guided enhancement mechanism for drone-based object detection in adverse visual environments","volume":"13","author":"Zhang","year":"2026","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026042709230520600_bib38","first-page":"16965","article-title":"Detrs beat yolos on real-time object detection","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhao","year":"2024"},{"key":"2026042709230520600_bib39","first-page":"2952","article-title":"Adapt or perish: Adaptive sparse transformer with attentive feature refinement for image restoration","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhou","year":"2024"},{"key":"2026042709230520600_bib40","article-title":"Deformable detr: Deformable transformers for end-to-end object detection","author":"Zhu","year":"2021","journal-title":"International Conference on Learning Representations"}],"container-title":["Journal of Computational Design and Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jcde\/advance-article-pdf\/doi\/10.1093\/jcde\/qwag037\/67725586\/qwag037.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jcde\/article-pdf\/13\/4\/290\/67725586\/qwag037.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jcde\/article-pdf\/13\/4\/290\/67725586\/qwag037.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,27]],"date-time":"2026-04-27T13:23:22Z","timestamp":1777296202000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jcde\/article\/13\/4\/290\/8572530"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,1]]},"references-count":40,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,1]]}},"URL":"https:\/\/doi.org\/10.1093\/jcde\/qwag037","relation":{},"ISSN":["2288-5048"],"issn-type":[{"value":"2288-5048","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026,4]]},"published":{"date-parts":[[2026,4,1]]}}}