{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T16:15:07Z","timestamp":1775837707517,"version":"3.50.1"},"reference-count":31,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2021,4,1]],"date-time":"2021-04-01T00:00:00Z","timestamp":1617235200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100003130","name":"Fonds Wetenschappelijk Onderzoek","doi-asserted-by":"publisher","award":["S003817N"],"award-info":[{"award-number":["S003817N"]}],"id":[{"id":"10.13039\/501100003130","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Object detection models are usually trained and evaluated on highly complicated, challenging academic datasets, which results in deep networks requiring lots of computations. However, a lot of operational use-cases consist of more constrained situations: they have a limited number of classes to be detected, less intra-class variance, less lighting and background variance, constrained or even fixed camera viewpoints, etc. In these cases, we hypothesize that smaller networks could be used without deteriorating the accuracy. However, there are multiple reasons why this does not happen in practice. Firstly, overparameterized networks tend to learn better, and secondly, transfer learning is usually used to reduce the necessary amount of training data. In this paper, we investigate how much we can reduce the computational complexity of a standard object detection network in such constrained object detection problems. As a case study, we focus on a well-known single-shot object detector, YoloV2, and combine three different techniques to reduce the computational complexity of the model without reducing its accuracy on our target dataset. To investigate the influence of the problem complexity, we compare two datasets: a prototypical academic (Pascal VOC) and a real-life operational (LWIR person detection) dataset. The three optimization steps we exploited are: swapping all the convolutions for depth-wise separable convolutions, perform pruning and use weight quantization. The results of our case study indeed substantiate our hypothesis that the more constrained a problem is, the more the network can be optimized. On the constrained operational dataset, combining these optimization techniques allowed us to reduce the computational complexity with a factor of 349, as compared to only a factor 9.8 on the academic dataset. When running a benchmark on an Nvidia Jetson AGX Xavier, our fastest model runs more than 15 times faster than the original YoloV2 model, whilst increasing the accuracy by 5% Average Precision (AP).<\/jats:p>","DOI":"10.3390\/jimaging7040064","type":"journal-article","created":{"date-parts":[[2021,4,1]],"date-time":"2021-04-01T10:44:01Z","timestamp":1617273841000},"page":"64","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Investigating the Potential of Network Optimization for a Constrained Object Detection Problem"],"prefix":"10.3390","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8679-6828","authenticated-orcid":false,"given":"Tanguy","family":"Ophoff","sequence":"first","affiliation":[{"name":"EAVISE, PSI, KU Leuven, Jan Pieter De Nayerlaan 5, 2860 Sint-Katelijne-Waver, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"C\u00e9dric","family":"Gullentops","sequence":"additional","affiliation":[{"name":"EAVISE, PSI, KU Leuven, Jan Pieter De Nayerlaan 5, 2860 Sint-Katelijne-Waver, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3667-7406","authenticated-orcid":false,"given":"Kristof","family":"Van Beeck","sequence":"additional","affiliation":[{"name":"EAVISE, PSI, KU Leuven, Jan Pieter De Nayerlaan 5, 2860 Sint-Katelijne-Waver, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7477-8961","authenticated-orcid":false,"given":"Toon","family":"Goedem\u00e9","sequence":"additional","affiliation":[{"name":"EAVISE, PSI, KU Leuven, Jan Pieter De Nayerlaan 5, 2860 Sint-Katelijne-Waver, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,4,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany, 5\u20139 October 2015, Springer.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_5","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). ImageNet: A Large-Scale Hierarchical Image Database. Proceedings of the Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1007\/s11263-014-0733-5","article-title":"The Pascal Visual Object Classes Challenge: A Retrospective","volume":"111","author":"Everingham","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft COCO: Common objects in context. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland, 7 March 2014, Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_9","unstructured":"Gaus, Y.F.A., Bhowmik, N., Isaac-Medina, B.K., and Breckon, T.P. (2020, January 21\u201325). Visible to infrared transfer learning as a paradigm for accessible real-time object detection and classification in infrared imagery. Proceedings of the Counterterrorism, Crime Fighting, Forensics, and Surveillance Technologies IV. International Society for Optics and Photonics, Bellingham, WA, USA."},{"key":"ref_10","unstructured":"Sifre, L., and Mallat, S. (2014). Rigid-Motion Scattering for Image Classification. arXiv."},{"key":"ref_11","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.C. (2018, January 18\u201323). Mobilenetv2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_13","unstructured":"Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J. (2017, January 24\u201326). Pruning Convolutional Neural Networks for Resource Efficient Inference. Proceedings of the 5th International Conference on Learning Representations, Toulon, France."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"He, Y., Liu, P., Wang, Z., Hu, Z., and Yang, Y. (2019, January 16\u201320). Filter Pruning via Geometric Median for Deep Convolutional Neural Networks Acceleration. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00447"},{"key":"ref_15","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2015). Distilling the knowledge in a neural network. arXiv."},{"key":"ref_16","unstructured":"Vanholder, H. (2021, March 31). Efficient Inference with Tensorrt. Available online: https:\/\/on-demand.gputechconf.com\/gtc-eu\/2017\/presentation\/23425-han-vanholder-efficient-inference-with-tensorrt.pdf."},{"key":"ref_17","unstructured":"Zhou, A., Yao, A., Guo, Y., Xu, L., and Chen, Y. (2017, January 24\u201326). Incremental network quantization: Towards lossless cnns with low-precision weights. Proceedings of the 5th International Conference on Learning Representations (ILCR), Toulon, France."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"105742","DOI":"10.1016\/j.compag.2020.105742","article-title":"Using channel pruning-based YOLO v4 deep learning algorithm for the real-time and accurate detection of apple flowers in natural environments","volume":"178","author":"Wu","year":"2020","journal-title":"Comput. Electron. Agric."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Md Zain, Z., Ahmad, H., Pebrianti, D., Mustafa, M., Abdullah, N.R.H., Samad, R., and Mat Noh, M. (2021). Analysis of Pruned Neural Networks (MobileNetV2-YOLO v2) for Underwater Object Detection. Proceedings of the 11th National Technical Seminar on Unmanned System Technolog, Kuantan, Malaysia, 2\u20133 December 2019, Springer.","DOI":"10.1007\/978-981-15-5281-6"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Van Beeck, K., Van Engeland, K., Vennekens, J., and Goedem\u00e9, T. (September, January 29). Abnormal behavior detection in LWIR surveillance of railway platforms. Proceedings of the 2017 14th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Lecce, Italy.","DOI":"10.1109\/AVSS.2017.8078540"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_22","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster R-CNN: Towards real-time object detection with region proposal networks. Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS), Montreal, QC, Canada."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016). Ssd: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, 8\u201316 October 2016, Springer.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Law, H., and Deng, J. (2018, January 8\u201314). Cornernet: Detecting objects as paired keypoints. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_45"},{"key":"ref_25","unstructured":"Ophoff, T. (2021, March 31). Lightnet: Building Blocks to Recreate Darknet Networks in Pytorch. Available online: https:\/\/gitlab.com\/EAVISE\/lightnet."},{"key":"ref_26","unstructured":"LeCun, Y., Denker, J.S., Solla, S.A., Howard, R.E., and Jackel, L.D. (1989, January 27\u201330). Optimal brain damage. Proceedings of the 2nd International Conference on Neural Information Processing Systems (NIPs), Denver, CO, USA."},{"key":"ref_27","unstructured":"Frankle, J., and Carbin, M. (2018). The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv."},{"key":"ref_28","unstructured":"Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2016). Understanding deep learning requires rethinking generalization. arXiv."},{"key":"ref_29","unstructured":"Wallach, H., Larochelle, H., Beygelzimer, A., d\u2019Alch\u00e9 Buc, F., Fox, E., and Garnett, R. (2019). PyTorch: An Imperative Style, High-Performance Deep Learning Library. Advances in Neural Information Processing Systems 32, Curran Associates, Inc."},{"key":"ref_30","unstructured":"Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G.S., Davis, A., Dean, J., and Devin, M. (2021, March 31). TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Available online: tensorflow.org."},{"key":"ref_31","unstructured":"Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T. (2018). Rethinking the value of network pruning. arXiv."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/7\/4\/64\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T13:33:26Z","timestamp":1760362406000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/7\/4\/64"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,1]]},"references-count":31,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2021,4]]}},"alternative-id":["jimaging7040064"],"URL":"https:\/\/doi.org\/10.3390\/jimaging7040064","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,4,1]]}}}