{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T16:00:28Z","timestamp":1784822428853,"version":"3.55.0"},"reference-count":106,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2021,7,28]],"date-time":"2021-07-28T00:00:00Z","timestamp":1627430400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Recent progress in deep learning has led to accurate and efficient generic object detection networks. Training of highly reliable models depends on large datasets with highly textured and rich images. However, in real-world scenarios, the performance of the generic object detection system decreases when (i) occlusions hide the objects, (ii) objects are present in low-light images, or (iii) they are merged with background information. In this paper, we refer to all these situations as challenging environments. With the recent rapid development in generic object detection algorithms, notable progress has been observed in the field of deep learning-based object detection in challenging environments. However, there is no consolidated reference to cover the state of the art in this domain. To the best of our knowledge, this paper presents the first comprehensive overview, covering recent approaches that have tackled the problem of object detection in challenging environments. Furthermore, we present a quantitative and qualitative performance analysis of these approaches and discuss the currently available challenging datasets. Moreover, this paper investigates the performance of current state-of-the-art generic object detection algorithms by benchmarking results on the three well-known challenging datasets. Finally, we highlight several current shortcomings and outline future directions.<\/jats:p>","DOI":"10.3390\/s21155116","type":"journal-article","created":{"date-parts":[[2021,7,28]],"date-time":"2021-07-28T21:21:04Z","timestamp":1627507264000},"page":"5116","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":71,"title":["Survey and Performance Analysis of Deep Learning Based Object Detection in Challenging Environments"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3606-7042","authenticated-orcid":false,"given":"Muhammad","family":"Ahmed","sequence":"first","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"Mindgrage, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0456-6493","authenticated-orcid":false,"given":"Khurram Azeem","family":"Hashmi","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"Mindgrage, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alain","family":"Pagani","sequence":"additional","affiliation":[{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4029-6574","authenticated-orcid":false,"given":"Marcus","family":"Liwicki","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Lule\u00e5 University of Technology, 971 87 Lule\u00e5, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Didier","family":"Stricker","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0536-6867","authenticated-orcid":false,"given":"Muhammad Zeshan","family":"Afzal","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"Mindgrage, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,7,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object detection with discriminatively trained part-based models","volume":"32","author":"Felzenszwalb","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Dai, J., He, K., and Sun, J. (2016, January 26). Instance-aware semantic segmentation via multi-task network cascades. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.343"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Hariharan, B., Arbel\u00e1ez, P., Girshick, R., and Malik, J. (2015, January 7\u201312). Hypercolumns for object segmentation and fine-grained localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298642"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Hariharan, B., Arbel\u00e1ez, P., Girshick, R., and Malik, J. (2014). Simultaneous detection and segmentation. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10584-0_20"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Alberti, C., Ling, J., Collins, M., and Reitter, D. (2019). Fusion of detected objects in text for visual question answering. arXiv.","DOI":"10.18653\/v1\/D19-1219"},{"key":"ref_6","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015, January 6\u201311). Show, attend and tell: Neural image caption generation with visual attention. Proceedings of the International Conference on Machine Learning (ICML 2015), Lille, France. PMLR 37:2048-2057."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1367","DOI":"10.1109\/TPAMI.2017.2708709","article-title":"Image captioning and visual question answering based on attributes and external knowledge","volume":"40","author":"Wu","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"2896","DOI":"10.1109\/TCSVT.2017.2736553","article-title":"T-cnn: Tubelets with convolutional neural networks for object detection from videos","volume":"28","author":"Kang","year":"2017","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, P., Lan, C., Zeng, W., Xing, J., Xue, J., and Zheng, N. (2020, January 16\u201318). Semantics-guided neural networks for efficient skeleton-based human action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00119"},{"key":"ref_10","unstructured":"Vaswani, N., Chowdhury, A.R., and Chellappa, R. (2003, January 18\u201320). Activity recognition using the dynamics of the configuration of interacting objects. Proceedings of the 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Madison, WI, USA."},{"key":"ref_11","unstructured":"Motwani, T.S., and Mooney, R.J. (2012). Improving Video Activity Recognition using Object Recognition and Text Mining, Citeseer. ECAI."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C.L., and Dollar, P. (2019). Microsoft COCO: Common objects in context (2014). arXiv.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The pascal visual object classes (voc) challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_14","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201326). Histograms of oriented gradients for human detection. Proceedings of the 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Felzenszwalb, P.F., Girshick, R.B., and McAllester, D. (2010, January 13\u201318). Cascade object detection with deformable part models. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539906"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 17\u201319). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"2189","DOI":"10.1109\/TPAMI.2012.28","article-title":"Measuring the objectness of image windows","volume":"34","author":"Alexe","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1312","DOI":"10.1109\/TPAMI.2011.231","article-title":"CPMC: Automatic object segmentation using constrained parametric min-cuts","volume":"34","author":"Carreira","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Rahtu, E., Kannala, J., and Blaschko, M. (2011, January 6\u201313). Learning a category independent object detection cascade. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126351"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","article-title":"Selective search for object recognition","volume":"104","author":"Uijlings","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zitnick, C.L., and Doll\u00e1r, P. (2014). Edge boxes: Locating object proposals from edges. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10602-1_26"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Kuo, W., Hariharan, B., and Malik, J. (2015, January 7\u201312). Deepbox: Learning objectness with convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Boston, MA, USA.","DOI":"10.1109\/ICCV.2015.285"},{"key":"ref_23","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv."},{"key":"ref_24","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (July, January 26). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2021, July 21). Ssd: Single Shot Multibox Detector. Available online: http:\/\/gitlinux.net\/assets\/SSD-Single-Shot-MultiBox-Detector.pdf.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_27","first-page":"353","article-title":"Learning to detect a salient object","volume":"33","author":"Liu","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"569","DOI":"10.1109\/TPAMI.2014.2345401","article-title":"Global contrast based salient region detection","volume":"37","author":"Cheng","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ouyang, W., Wang, X., Zeng, X., Qiu, S., Luo, P., Tian, Y., Li, H., Yang, S., Wang, Z., and Loy, C.C. (2015, January 7\u201312). Deepid-net: Deformable deep convolutional neural networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298854"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1109\/MSP.2017.2749125","article-title":"Advanced deep-learning techniques for salient and category-specific object detection: A survey","volume":"35","author":"Han","year":"2018","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201312). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Boston, MA, USA.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_33","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_34","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_37","unstructured":"Lienhart, R., and Maydt, J. (2002, January 22\u201325). An extended set of haar-like features for rapid object detection. Proceedings of the International Conference on Image Processing, Rochester, NY, USA."},{"key":"ref_38","unstructured":"Agarwal, S., Terrail, J.O.D., and Jurie, F. (2018). Recent advances in object detection in the age of deep convolutional neural networks. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Huang, J., Rathod, V., Sun, C., Zhu, M., Korattikara, A., Fathi, A., Fischer, I., Wojna, Z., Song, Y., and Guadarrama, S. (2017, January 21\u201326). Speed\/accuracy trade-offs for modern convolutional object detectors. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.351"},{"key":"ref_40","first-page":"1","article-title":"Visual object recognition","volume":"5","author":"Grauman","year":"2011","journal-title":"Synth. Lect. Artif. Intell. Mach. Learn."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"827","DOI":"10.1016\/j.cviu.2013.04.005","article-title":"50 years of object recognition: Directions forward","volume":"117","author":"Andreopoulos","year":"2013","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_42","unstructured":"Zou, Z., Shi, Z., Guo, Y., and Ye, J. (2019). Object detection in 20 years: A survey. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"128837","DOI":"10.1109\/ACCESS.2019.2939201","article-title":"A survey of deep learning-based object detection","volume":"7","author":"Jiao","year":"2019","journal-title":"IEEE Access"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"3782","DOI":"10.1109\/TITS.2019.2892405","article-title":"A survey on 3d object detection methods for autonomous driving applications","volume":"20","author":"Arnold","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Bharati, P., and Pramanik, A. (2020). Deep Learning Techniques\u2014R-CNN to Mask R-CNN: A Survey. Computational Intelligence in Pattern Recognition, Springer.","DOI":"10.1007\/978-981-13-9042-5_56"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1016\/j.compag.2019.01.012","article-title":"Apple detection during different growth stages in orchards using the improved YOLO-V3 model","volume":"157","author":"Tian","year":"2019","journal-title":"Comput. Electron. Agric."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Lan, W., Dang, J., Wang, Y., and Wang, S. (2018, January 5\u20138). Pedestrian detection based on YOLO network model. Proceedings of the 2018 IEEE International Conference on Mechatronics and Automation (ICMA), Changchun, China.","DOI":"10.1109\/ICMA.2018.8484698"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Sasagawa, Y., and Nagahara, H. (2020). Yolo in the dark-domain adaptation method for merging multiple models. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58589-1_21"},{"key":"ref_49","unstructured":"Sutskever, I., Vinyals, O., and Le, Q.V. (2014). Sequence to sequence learning with neural networks. arXiv."},{"key":"ref_50","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_52","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 6\u201311). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the International Conference on Machine Learning (ICML 2015), Lille, France. PMLR 37:448-456."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Chen, C., Chen, Q., Xu, J., and Koltun, V. (2018, January 18\u201322). Learning to see in the dark. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00347"},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"125459","DOI":"10.1109\/ACCESS.2020.3007481","article-title":"Thermal Object Detection in Difficult Weather Conditions Using YOLO","volume":"8","author":"Pobar","year":"2020","journal-title":"IEEE Access"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Kri\u0161to, M., and Iva\u0161i\u0107-Kos, M. (2019, January 20\u201324). Thermal imaging dataset for person detection. Proceedings of the 2019 42nd International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO), Opatija, Croatia.","DOI":"10.23919\/MIPRO.2019.8757208"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Liu, S., and Huang, D. (2018, January 8\u201314). Receptive field block net for accurate and fast object detection. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_24"},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"123075","DOI":"10.1109\/ACCESS.2020.3007610","article-title":"Making of night vision: Object detection under low-illumination","volume":"8","author":"Xiao","year":"2020","journal-title":"IEEE Access"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Sarin, M., Chandrakar, S., and Patel, R. (2019, January 11\u201312). Face and Human Detection in Low Light for Surveillance Purposes. Proceedings of the 2019 International Conference on Computational Intelligence and Knowledge Economy (ICCIKE), Dubai, UAE.","DOI":"10.1109\/ICCIKE47802.2019.9004249"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Hwang, S., Park, J., Kim, N., Choi, Y., and So Kweon, I. (2015, January 7\u201312). Multispectral pedestrian detection: Benchmark dataset and baseline. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298706"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Nada, H., Sindagi, V.A., Zhang, H., and Patel, V.M. (2018, January 22\u201325). Pushing the limits of unconstrained face detection: A challenge dataset and baseline results. Proceedings of the 2018 IEEE 9th International Conference on Bio metrics Theory, Applications and Systems (BTAS), Redondo Beach, CA, USA.","DOI":"10.1109\/BTAS.2018.8698561"},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1109\/TBIOM.2019.2908436","article-title":"A fast and accurate system for face detection, identification, and verification","volume":"1","author":"Ranjan","year":"2019","journal-title":"IEEE Trans. Biom. Behav. Identity Sci."},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"189855","DOI":"10.1109\/ACCESS.2020.3031191","article-title":"Neural-Network-Based Traffic Sign Detection and Recognition in High-Definition Images Using Region Focusing and Parallelization","volume":"8","author":"Sluga","year":"2020","journal-title":"IEEE Access"},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Goldman, E., Herzig, R., Eisenschtat, A., Goldberger, J., and Hassner, T. (2019, January 16\u201320). Precise detection in densely packed scenes. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00537"},{"key":"ref_64","doi-asserted-by":"crossref","first-page":"193168","DOI":"10.1109\/ACCESS.2020.3032981","article-title":"Object Recognition at Night Scene Based on DCGAN and Faster R-CNN","volume":"8","author":"Wang","year":"2020","journal-title":"IEEE Access"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Ghose, D., Desai, S.M., Bhattacharya, S., Chakraborty, D., Fiterau, M., and Rahman, T. (2019, January 16\u201320). Pedestrian detection in thermal images using saliency maps. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Long Beach, CA, USA.","DOI":"10.1109\/CVPRW.2019.00130"},{"key":"ref_66","unstructured":"Rashed, H., Ramzy, M., Vaquero, V., El Sallab, A., Sistu, G., and Yogamani, S. (November, January 27). Fusemodnet: Real-time camera and lidar based moving object detection for robust low-light autonomous driving. Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops, Seoul, Korea."},{"key":"ref_67","doi-asserted-by":"crossref","first-page":"1467","DOI":"10.1109\/TITS.2019.2911727","article-title":"Automatic traffic sign detection and recognition using SegU-net and a modified tversky loss function with L1-constraint","volume":"21","author":"Kamal","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Wang, Q., Zhang, L., Bertinetto, L., Hu, W., and Torr, P.H. (2019, January 16\u201320). Fast online object tracking and segmentation: A unifying approach. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00142"},{"key":"ref_69","unstructured":"Tu, Z., Ma, Y., Li, Z., Li, C., Xu, J., and Liu, Y. (2020). RGBT salient object detection: A large-scale dataset and benchmark. arXiv."},{"key":"ref_70","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_71","unstructured":"Szegedy, C., Reed, S., Erhan, D., Anguelov, D., and Ioffe, S. (2014). Scalable, high-quality object detection. arXiv."},{"key":"ref_72","unstructured":"Yang, S., Luo, P., Loy, C.C., and Tang, X. (July, January 26). Wider face: A face detection benchmark. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Klare, B.F., Klein, B., Taborsky, E., Blanton, A., Cheney, J., Allen, K., Grother, P., Mah, A., and Jain, A.K. (2015, January 7\u201312). Pushing the frontiers of unconstrained face detection and recognition: Iarpa janus benchmark a. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298803"},{"key":"ref_74","doi-asserted-by":"crossref","unstructured":"Whitelam, C., Taborsky, E., Blanton, A., Maze, B., Adams, J., Miller, T., Kalka, N., Jain, A.K., Duncan, J.A., and Allen, K. (2017, January 21\u201326). Iarpa janus benchmark-b face dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.87"},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Maze, B., Adams, J., Duncan, J.A., Kalka, N., Miller, T., Otto, C., Jain, A.K., Niggel, W.T., Anderson, J., and Cheney, J. (2018, January 20\u201323). Iarpa janus benchmark-c: Face dataset and protocol. Proceedings of the 2018 International Conference on Biometrics (ICB), Queensland, Australia.","DOI":"10.1109\/ICB2018.2018.00033"},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Fan, Q., Brown, L., and Smith, J. (2016, January 19\u201322). A closer look at Faster R-CNN for vehicle detection. Proceedings of the 2016 IEEE Intelligent Vehicles Symposium (IV), Gothenberg, Sweden.","DOI":"10.1109\/IVS.2016.7535375"},{"key":"ref_77","unstructured":"He, Z., and Zhang, L. (November, January 27). Multi-adversarial faster-rcnn for unrestricted object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_78","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1109\/MSP.2017.2765202","article-title":"Generative adversarial networks: An overview","volume":"35","author":"Creswell","year":"2018","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_79","unstructured":"Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. (2017). Improved training of wasserstein gans. arXiv."},{"key":"ref_80","unstructured":"Kopelowitz, E., and Engelhard, G. (2019). Lung Nodules Detection and Segmentation Using 3D Mask-RCNN. arXiv."},{"key":"ref_81","doi-asserted-by":"crossref","first-page":"6997","DOI":"10.1109\/ACCESS.2020.2964055","article-title":"Vehicle-damage-detection segmentation algorithm based on improved mask RCNN","volume":"8","author":"Zhang","year":"2020","journal-title":"IEEE Access"},{"key":"ref_82","doi-asserted-by":"crossref","first-page":"1427","DOI":"10.1109\/TITS.2019.2913588","article-title":"Deep learning for large-scale traffic-sign detection and recognition","volume":"21","author":"Tabernik","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_83","unstructured":"Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A.L. (2014). Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv."},{"key":"ref_84","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_85","doi-asserted-by":"crossref","unstructured":"Liu, N., Han, J., and Yang, M.H. (2018, January 18\u201322). Picanet: Learning pixel-wise contextual attention for saliency detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00326"},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Deng, Z., Hu, X., Zhu, L., Xu, X., Qin, J., Han, G., and Heng, P.A. R3net: Recurrent residual refinement network for saliency detection. Proceedings of the 27th International Joint Conference on Artificial Intelligence, Stockholm, Sweden, 13\u201319 July 2018.","DOI":"10.24963\/ijcai.2018\/95"},{"key":"ref_87","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_88","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_89","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 18\u201320). Are we ready for autonomous driving? The kitti vision benchmark suite. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA. PMLR 116:171-183.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_90","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_91","doi-asserted-by":"crossref","unstructured":"Salehi, S.S.M., Erdogmus, D., and Gholipour, A. (2017). Tversky loss function for image segmentation using 3D fully convolutional deep networks. International Workshop on Machine Learning in Medical Imaging, Springer.","DOI":"10.1007\/978-3-319-67389-9_44"},{"key":"ref_92","doi-asserted-by":"crossref","first-page":"3663","DOI":"10.1109\/TITS.2019.2931429","article-title":"Traffic sign detection under challenging conditions: A deeper look into performance variations and spectral characteristics","volume":"21","author":"Temel","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_93","doi-asserted-by":"crossref","unstructured":"Bertinetto, L., Valmadre, J., Henriques, J.F., Vedaldi, A., and Torr, P.H. (2016). Fully-convolutional siamese networks for object tracking. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-48881-3_56"},{"key":"ref_94","doi-asserted-by":"crossref","unstructured":"Li, B., Yan, J., Wu, W., Zhu, Z., and Hu, X. (2018, January 18\u201322). High performance visual tracking with siamese region proposal network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00935"},{"key":"ref_95","unstructured":"Kristan, M., Leonardis, A., Matas, J., Felsberg, M., Pflugfelder, R., \u010cehovin Zajc, L., Vojir, T., Bhat, G., Lukezic, A., and Eldesokey, A. (2018, January 8\u201314). The sixth visual object tracking vot2018 challenge results. Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Munich, Germany."},{"key":"ref_96","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_97","doi-asserted-by":"crossref","first-page":"30","DOI":"10.1016\/j.cviu.2018.10.010","article-title":"Getting to know low-light images with the exclusively dark dataset","volume":"178","author":"Loh","year":"2019","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_98","doi-asserted-by":"crossref","first-page":"492","DOI":"10.1109\/TIP.2018.2867951","article-title":"Benchmarking single-image dehazing and beyond","volume":"28","author":"Li","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_99","unstructured":"Jaeger, P.F., Kohl, S.A., Bickelhaupt, S., Isensee, F., Kuder, T.A., Schlemmer, H.P., and Maier-Hein, K.H. (2020, January 7\u20138). Retina U-Net: Embarrassingly simple exploitation of segmentation supervision for medical object detection. Proceedings of the Machine Learning for Health NeurIPS Workshop, Durham, NC, USA."},{"key":"ref_100","doi-asserted-by":"crossref","first-page":"1483","DOI":"10.1109\/TPAMI.2019.2956516","article-title":"Cascade R-CNN: High quality object detection and instance segmentation","volume":"43","author":"Cai","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_101","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. (2017, January 4\u20139). Inception-v4, inception-resnet and the impact of residual connections on learning. Proceedings of the AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"ref_102","doi-asserted-by":"crossref","unstructured":"Zhang, Z. (2018, January 4\u20136). Improved adam optimizer for deep neural networks. Proceedings of the 2018 IEEE\/ACM 26th International Symposium on Quality of Service (IWQoS), Banff, AB, Canada.","DOI":"10.1109\/IWQoS.2018.8624183"},{"key":"ref_103","unstructured":"Powers, D.M. (2020). Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. arXiv."},{"key":"ref_104","doi-asserted-by":"crossref","unstructured":"Blaschko, M.B., and Lampert, C.H. (2008). Learning to localize objects with structured output regression. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-540-88682-2_2"},{"key":"ref_105","doi-asserted-by":"crossref","unstructured":"Zhu, J.Y., Park, T., Isola, P., and Efros, A.A. (2017, January 22\u201329). Unpaired image-to-image translation using cycle-consistent adversarial networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.244"},{"key":"ref_106","doi-asserted-by":"crossref","unstructured":"Nidadavolu, P.S., Villalba, J., and Dehak, N. (2019, January 12\u201317). Cycle-gans for domain adaptation of acoustic features for speaker recognition. Proceedings of the ICASSP 2019\u20142019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683055"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/15\/5116\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:36:20Z","timestamp":1760164580000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/15\/5116"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,28]]},"references-count":106,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2021,8]]}},"alternative-id":["s21155116"],"URL":"https:\/\/doi.org\/10.3390\/s21155116","relation":{"has-preprint":[{"id-type":"doi","id":"10.20944\/preprints202106.0590.v1","asserted-by":"object"}]},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,7,28]]}}}