{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T05:17:56Z","timestamp":1785388676371,"version":"3.55.0"},"reference-count":54,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2023,11,28]],"date-time":"2023-11-28T00:00:00Z","timestamp":1701129600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,11,28]],"date-time":"2023-11-28T00:00:00Z","timestamp":1701129600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2024,2]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Object detection is one of the core tasks of computer vision, and bounding box (bbox) regression is one of the basic tasks of object detection. In recent years of related research, bbox regression is often used in the Intersection over Union (IoU) loss and its improved version. In this paper, for the first time, we introduce the Dice coefficient into the regression loss calculation and propose a new measure which is superior to and can replace the IoU. We define three properties of the new measure and prove the theory by mathematical reasoning and analysis of the existing work. This paper also proposes the N-IoU regression loss family. And the superiority of the N-IoU regression loss family is proved by designing simulation experiments and comparative experiments. The main results of this paper are: (1) The proposed new measure is better than IoU which can be used to evaluate bounding box regression, and the three properties of the new measure can be used as a broad criterion for the design of regression loss functions; and (2) we propose N-IoU loss. The parameter <jats:italic>n<\/jats:italic> of N-IOU can be debugged, which can be widely adapted to different application scenarios with higher flexibility, and the regression performance is better.<\/jats:p>","DOI":"10.1007\/s00521-023-09133-4","type":"journal-article","created":{"date-parts":[[2023,11,28]],"date-time":"2023-11-28T17:02:07Z","timestamp":1701190927000},"page":"3049-3063","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":52,"title":["N-IoU: better IoU-based bounding box regression loss for object detection"],"prefix":"10.1007","volume":"36","author":[{"given":"Keke","family":"Su","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3235-6557","authenticated-orcid":false,"given":"Lihua","family":"Cao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Botong","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ning","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Di","family":"Wu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiyu","family":"Han","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,11,28]]},"reference":[{"key":"9133_CR1","doi-asserted-by":"crossref","unstructured":"Girshick R (2015) Fast r-cnn. In: International Conference on Computer vision(ICCV), pp 1440\u20131448","DOI":"10.1109\/ICCV.2015.169"},{"issue":"15","key":"9133_CR2","first-page":"234","volume":"31","author":"MA Rahman","year":"2016","unstructured":"Rahman MA, Wang Y (2016) Optimizing intersection-over-union in deep neural networks for image segmentation. Int Symp vis Comput 31(15):234\u2013244","journal-title":"Int Symp vis Comput"},{"key":"9133_CR3","doi-asserted-by":"crossref","unstructured":"Yu J, Jiang Y, Wang Z, Cao Z, Huang T (2016) Unitbox: an advanced object detection network. In: ACM International Conference on Multimedia, pp 516\u2013520","DOI":"10.1145\/2964284.2967274"},{"key":"9133_CR4","doi-asserted-by":"crossref","unstructured":"Rezatofighi H, Tsoi, N, Gwak J, Sadeghian A, Reid I, Savarese S (2019) Generalized intersection over union: a metric and a loss for bounding box regression. In: International Conference on Computer Vision (ICCV), pp 658\u2013666","DOI":"10.1109\/CVPR.2019.00075"},{"key":"9133_CR5","doi-asserted-by":"crossref","unstructured":"Zheng Z, Wang P, Liu W, Li J, Ye R, Ren D (2020) Distance-IoU loss: faster and better learning for bounding box regression. In: Association for the Advancement of Artificial Intelligence (AAAI), pp 12993\u201313000","DOI":"10.1609\/aaai.v34i07.6999"},{"key":"9133_CR6","unstructured":"He J, Erfani S, Ma X, Bailey J, Chi Y, Hua XS (2022) Alpha-IoU: a family of power intersection over union losses for bounding box regression. arXiv:2110.13675v2"},{"key":"9133_CR7","doi-asserted-by":"crossref","unstructured":"Zhang YF, Ren W, Zhang Z, Jia Z, Wang L, Tan T (2021) Focal and efficient IoU loss for accurate bounding box regression. arXiv:2101.08158","DOI":"10.1016\/j.neucom.2022.07.042"},{"key":"9133_CR8","unstructured":"Wu S, Yang J, Yu H, Gou L, Li X (2022) Gaussian guided IoU: a better metric for balanced learning on object detection. In: IET Computer Vision"},{"key":"9133_CR9","doi-asserted-by":"crossref","unstructured":"Wang K, Zhang L (2020) Single-shot two-pronged detector with rectified IoU loss. In: ACM International Conference Multimedia, pp 1311\u20131319","DOI":"10.1145\/3394171.3413691"},{"key":"9133_CR10","unstructured":"Redmon J, Farhadi A (2018) Yolov3: an incremental improvement. arXiv:1804.02767"},{"key":"9133_CR11","doi-asserted-by":"crossref","unstructured":"Liu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu C-Y, Berg AC (2016) Ssd: single shot multibox detector. In: European Conference on Computer Vision (ECCV), pp 21\u201337","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"9133_CR12","unstructured":"Ren S, He K, Girshick R, Sun J (2015) Faster R-CNN: towards real-time object detection with region proposal networks. In: Advances in Neural Information Processing Systems (NeurIPS), pp 91\u201399"},{"key":"9133_CR13","unstructured":"Ge Z, Liu S, Wang F, Li Z, Sun J (2021) Yolox: exceeding yolo series in 2021. arXiv:2107.08430"},{"key":"9133_CR14","unstructured":"Jocher G, Chaurasia A, Qiu J (2023) YOLO by ultralytics"},{"key":"9133_CR15","doi-asserted-by":"crossref","unstructured":"Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A, Zagoruyko S (2020) End-to-end object detection with transformers. In: European Conference on Computer Vision (ECCV)","DOI":"10.1007\/978-3-030-58452-8_13"},{"issue":"2","key":"9133_CR16","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","volume":"88","author":"M Everingham","year":"2010","unstructured":"Everingham M, Gool LV, Williams CK, Winn J, Zisserman A (2010) The pascal visual object classes (VOC) challenge. Int J Comput Vis 88(2):303\u2013338","journal-title":"Int J Comput Vis"},{"key":"9133_CR17","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: European Conference on Computer Vision (ECCV), pp 740\u2013755","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"9133_CR18","doi-asserted-by":"crossref","unstructured":"Girshick R, Donahue J, Darrell T, Malik J (2014) Rich feature hierarchies for accurate object detection and semantic segmentation. In: Conference on Computer Vision and Pattern Recognition (CVPR), pp 580\u2013587","DOI":"10.1109\/CVPR.2014.81"},{"key":"9133_CR19","unstructured":"Fu C-Y, Liu W, Ranga A, Tyagi A, Berg AC (2017) Dssd : deconvolutional single shot detector. arXiv:1701.06659"},{"key":"9133_CR20","doi-asserted-by":"crossref","unstructured":"Jeong J, Park H, Kwak N (2017) Enhancement of SSD by concatenating feature maps for object detection. arXiv:1705.09587","DOI":"10.5244\/C.31.76"},{"key":"9133_CR21","doi-asserted-by":"crossref","unstructured":"Redmon J, Divvala S, Girshick R, Farhadi A (2016) You only look once: unified, real-time object detection. In: Conference on Computer Vision and Pattern Recognition (CVPR), pp 779\u2013788","DOI":"10.1109\/CVPR.2016.91"},{"key":"9133_CR22","doi-asserted-by":"crossref","unstructured":"Redmon J, Farhadi A (2017) Yolo9000: better, faster, stronger. In: Conference on Computer Vision and Pattern Recognition (CVPR), pp 7263\u20137271","DOI":"10.1109\/CVPR.2017.690"},{"key":"9133_CR23","unstructured":"Bochkovskiy A, Wang C, Liao HM (2020) Yolov4: optimal speed and accuracy of object detection. arXiv:2004.10934"},{"key":"9133_CR24","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Goyal P, Girshick R, He K, Doll\u00e1r P (2017) Focal loss for dense object detection. In: International Conference on Computer Vision (ICCV), pp. 2980\u20132988","DOI":"10.1109\/ICCV.2017.324"},{"key":"9133_CR25","doi-asserted-by":"crossref","unstructured":"Tian Z, Shen C, Chen H, He T (2019) Fcos: fully convolutional one-stage object detection. In: International Conference on Computer Vision (ICCV), pp 9627\u20139636","DOI":"10.1109\/ICCV.2019.00972"},{"key":"9133_CR26","doi-asserted-by":"crossref","unstructured":"Law H, Deng J (2018) Cornernet: detecting objects as paired keypoints. In: European Conference on Computer Vision (ECCV), pp 734\u2013750","DOI":"10.1007\/978-3-030-01264-9_45"},{"key":"9133_CR27","doi-asserted-by":"crossref","unstructured":"Duan K, Bai S, Xie L, Qi H, Huang Q, Tian Q (2019) Centernet: keypoint triplets for object detection. In: International Conference on Computer Vision (ICCV), pp 6569\u20136578","DOI":"10.1109\/ICCV.2019.00667"},{"key":"9133_CR28","doi-asserted-by":"crossref","unstructured":"He K, Gkioxari G, Doll\u00e1r P, Girshick R (2017) Mask r-cnn. In: International Conference on Computer Vision (ICCV), pp 2961\u20132969","DOI":"10.1109\/ICCV.2017.322"},{"key":"9133_CR29","doi-asserted-by":"crossref","unstructured":"Cai Z, Vasconcelos N (2017) Cascade r-cnn: delving into high quality object detection. arXiv:1712.00726","DOI":"10.1109\/CVPR.2018.00644"},{"key":"9133_CR30","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2014) Spatial pyramid pooling in deep convolutional networks for visual recognition. In: European Conference on Computer Vision (ECCV), pp 346\u2013361","DOI":"10.1007\/978-3-319-10578-9_23"},{"key":"9133_CR31","doi-asserted-by":"crossref","unstructured":"Chen K, Pang J, Wang J, Xiong Y, Li X, Sun S, Feng W, Liu Z, Shi J, Ouyang W, al (2019) Hybrid task cascade for instance segmentation. In: Conference on Computer Vision and Pattern Recognition (CVPR), pp 4974\u201349831","DOI":"10.1109\/CVPR.2019.00511"},{"key":"9133_CR32","doi-asserted-by":"crossref","unstructured":"Song G, Liu Y, Wang X (2020) Revisiting the sibling head in object detector. In: Conference on Computer Vision and Pattern Recognition (CVPR), pp 11563\u201311572","DOI":"10.1109\/CVPR42600.2020.01158"},{"key":"9133_CR33","doi-asserted-by":"crossref","unstructured":"Duan K, Xie L, Qi H, Bai S, Huang Q, Tian Q (2020) CPNDET: corner proposal network for anchor-free, two-stage object detection. In: European Conference on Computer Vision (ECCV)","DOI":"10.1007\/978-3-030-58580-8_24"},{"key":"9133_CR34","unstructured":"Zhou X, Koltun V, Kr\u00e4henb\u00fchl P (2021) Probabilistic two-stage detection. arXiv:2103.07461"},{"key":"9133_CR35","doi-asserted-by":"crossref","unstructured":"Zhu C, He Y, Savvides M (2019) Feature selective anchor-free module for single-shot object detection. In: Conference on Computer Vision and Pattern Recognition (CVPR), pp 840\u2013849","DOI":"10.1109\/CVPR.2019.00093"},{"key":"9133_CR36","doi-asserted-by":"crossref","unstructured":"Wang J, Chen K, Yang S, Loy CC, Lin D (2019) Region proposal by guided anchoring. In: Conference on Computer Vision and Pattern Recognition (CVPR), pp 2960\u20132969","DOI":"10.1109\/CVPR.2019.00308"},{"key":"9133_CR37","doi-asserted-by":"crossref","unstructured":"Xie S, Tu Z (2015) Holistically-nested edge detection. In: International Conference on Computer Vision (ICCV)","DOI":"10.1109\/ICCV.2015.164"},{"key":"9133_CR38","doi-asserted-by":"crossref","unstructured":"Li J, Cheng B, Feris R, Xiong J, Huang T, Hwu WM, Shi H (2021) Pseudo-IoU: improving label assignment in anchor-free object detection. In: Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2378\u20132387","DOI":"10.1109\/CVPRW53098.2021.00270"},{"key":"9133_CR39","doi-asserted-by":"crossref","unstructured":"Gao Y, Wang Q, Tang X, Wang H, Ding F, Li J, Hu Y (2022) Decoupled IoU regression for object detection. arXiv:2202.00866","DOI":"10.1145\/3474085.3475707"},{"issue":"4","key":"9133_CR40","doi-asserted-by":"publisher","first-page":"51","DOI":"10.3390\/jlpea12040051","volume":"12","author":"N Ravi","year":"2022","unstructured":"Ravi N, Naqvi S, El-Sharkawy M (2022) BIOU: an improved bounding box regression for object detection. J Low Power Electron 12(4):51","journal-title":"J Low Power Electron"},{"key":"9133_CR41","doi-asserted-by":"publisher","first-page":"24","DOI":"10.1007\/s11554-023-01287-7","volume":"20","author":"F Gao","year":"2023","unstructured":"Gao F, Cai C, Jia R, Hu X (2023) Improved Yolox for pedestrian detection in crowded scenes. J Real-Time Image Proc 20:24","journal-title":"J Real-Time Image Proc"},{"key":"9133_CR42","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1016\/j.neucom.2022.05.052","volume":"500","author":"Y Shen","year":"2022","unstructured":"Shen Y, Zhang F, Liu D, Pu W, Zhang Q (2022) Manhattan-distance IoU loss for fast and accurate bounding box regression for object detection. Neurocomputing 500:99\u2013114","journal-title":"Neurocomputing"},{"key":"9133_CR43","unstructured":"Ma S, Xu Y (2023) MPDIoU: a loss for efficient and accurate bounding box regression. arXiv:2307.07662v1"},{"key":"9133_CR44","unstructured":"Gevorgyan Z (2022) SIoU loss: more powerful learning for bounding box regression. arXiv:2205.12740"},{"key":"9133_CR45","unstructured":"Tong Z, Chen Y, Xu Z, Yu R (2023) Wise-IoU: bounding box regression loss with dynamic focusing mechanism. arXiv:2301.10051v3"},{"key":"9133_CR46","unstructured":"Shruti J (2020) A survey of loss functions for semantic segmentation. In: IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), pp 115\u2013121"},{"key":"9133_CR47","doi-asserted-by":"crossref","unstructured":"Sudre CH, Li W, Vercauteren T, Ourselin S, Cardoso MJ (2017) Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations. In: 3rd MICCAI International Workshop on Deep Learning in Medical Image Analysis (DLMIA), pp 240\u2013248","DOI":"10.1007\/978-3-319-67558-9_28"},{"key":"9133_CR48","doi-asserted-by":"crossref","unstructured":"Salehi SSM, Erdogmus D, Gholipour A (2017) Tversky loss function for image segmentation using 3d fully onvolutional deep networks. In: International Workshop on Machine Learning in Medical Imaging (MLMI), pp 379\u2013387","DOI":"10.1007\/978-3-319-67389-9_44"},{"key":"9133_CR49","doi-asserted-by":"publisher","first-page":"1721","DOI":"10.1109\/ACCESS.2018.2886371","volume":"7","author":"SR Hashemi","year":"2019","unstructured":"Hashemi SR, Salehi SSM, Erdogmus D, Prabhu SP, Warfield SK, Gholipour A (2019) Asymmetric loss functions and deep densely-connected networks for highly-imbalanced medical image segmentation: application to multiple sclerosis lesion detection. IEEE Access 7:1721\u20131735","journal-title":"IEEE Access"},{"key":"9133_CR50","doi-asserted-by":"crossref","unstructured":"Milletari F, Navab N, Ahmadi S-A (2016) V-net: fully convolutional neural networks for volumetric medical image segmentation. In: IEEE International Conference on 3D Vision (3DV), pp 565\u2013571","DOI":"10.1109\/3DV.2016.79"},{"key":"9133_CR51","doi-asserted-by":"crossref","unstructured":"Cheng G, Yuan X, Yao X, Yan K, Zeng Q, Xie X, Han J (2023) Towards large-scale small object detection survey and benchmarks. In: IEEE Transactions on Pattern Analysis and Machine Intelligence (Early Access)","DOI":"10.1109\/TPAMI.2023.3290594"},{"key":"9133_CR52","unstructured":"Kervadec H, Bouchtiba J, Desrosiers C, Granger E, Dolz J, Ayed IB (2019) Boundary loss for highly unbalanced segmentation. In: PMLR, 2019, pp 285\u2013296"},{"key":"9133_CR53","doi-asserted-by":"publisher","first-page":"24","DOI":"10.1016\/j.compmedimag.2019.04.005","volume":"75","author":"SA Taghanaki","year":"2019","unstructured":"Taghanaki SA, Zheng YF, Zhou SK, Georgescu B, Sharma P, Xu DG, Comaniciu D, Hamarneh G (2019) Combo loss: handling input and output imbalance in multi-organ segmentation. Comput Med Imaging Graph 75:24\u201333","journal-title":"Comput Med Imaging Graph"},{"key":"9133_CR54","doi-asserted-by":"crossref","unstructured":"Wong KCL, Moradi M, Tang H, Syeda-Mahmood T (2018) 3d segmentation with exponential logarithmic loss for highly unbalanced object sizes. In: International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), pp 612\u2013619","DOI":"10.1007\/978-3-030-00931-1_70"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-023-09133-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-023-09133-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-023-09133-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,23]],"date-time":"2024-01-23T07:12:00Z","timestamp":1705993920000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-023-09133-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,28]]},"references-count":54,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2024,2]]}},"alternative-id":["9133"],"URL":"https:\/\/doi.org\/10.1007\/s00521-023-09133-4","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,28]]},"assertion":[{"value":"25 February 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 October 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 November 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that there is no conflict of interests, we do not have any possible conflicts of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}