{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T15:15:33Z","timestamp":1784906133712,"version":"3.55.0"},"reference-count":35,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2019,2,19]],"date-time":"2019-02-19T00:00:00Z","timestamp":1550534400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In this paper, we investigate whether fusing depth information on top of normal RGB data for camera-based object detection can help to increase the performance of current state-of-the-art single-shot detection networks. Indeed, depth sensing is easily acquired using depth cameras such as a Kinect or stereo setups. We investigate the optimal manner to perform this sensor fusion with a special focus on lightweight single-pass convolutional neural network (CNN) architectures, enabling real-time processing on limited hardware. For this, we implement a network architecture allowing us to parameterize at which network layer both information sources are fused together. We performed exhaustive experiments to determine the optimal fusion point in the network, from which we can conclude that fusing towards the mid to late layers provides the best results. Our best fusion models significantly outperform the baseline RGB network in both accuracy and localization of the detections.<\/jats:p>","DOI":"10.3390\/s19040866","type":"journal-article","created":{"date-parts":[[2019,2,20]],"date-time":"2019-02-20T03:05:52Z","timestamp":1550631952000},"page":"866","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":71,"title":["Exploring RGB+Depth Fusion for Real-Time Object Detection"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8679-6828","authenticated-orcid":false,"given":"Tanguy","family":"Ophoff","sequence":"first","affiliation":[{"name":"EAVISE, KU Leuven, 2860 Sint-Katelijne-Waver, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3667-7406","authenticated-orcid":false,"given":"Kristof","family":"Van Beeck","sequence":"additional","affiliation":[{"name":"EAVISE, KU Leuven, 2860 Sint-Katelijne-Waver, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7477-8961","authenticated-orcid":false,"given":"Toon","family":"Goedem\u00e9","sequence":"additional","affiliation":[{"name":"EAVISE, KU Leuven, 2860 Sint-Katelijne-Waver, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,2,19]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Felzenszwalb, P., McAllester, D., and Ramanan, D. (2008, January 23\u201328). A discriminatively trained, multiscale, deformable part model. Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition, Anchorage, AK, USA.","DOI":"10.1109\/CVPR.2008.4587597"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Doll\u00e1r, P., Tu, Z., Perona, P., and Belongie, S. (2009). Integral channel features. Proceedings of the British Machine Vision Conference, BMVC Press.","DOI":"10.5244\/C.23.91"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1532","DOI":"10.1109\/TPAMI.2014.2300479","article-title":"Fast feature pyramids for object detection","volume":"36","author":"Appel","year":"2014","journal-title":"IEEE TPAMI"},{"key":"ref_4","first-page":"347","article-title":"Multispectral pedestrian detection: Benchmark dataset and baseline","volume":"20","author":"Hwang","year":"2013","journal-title":"Integr. Comput. Aided Eng."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"De Smedt, F., Puttemans, S., and Goedem\u00e9, T. (2016, January 12\u201315). How to reach top accuracy for a visual pedestrian warning system from a car?. Proceedings of the 2016 Sixth International Conference on Image Processing Theory, Tools and Applications (IPTA), Oulu, Finland.","DOI":"10.1109\/IPTA.2016.7820997"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Spinello, L., and Arras, K.O. (2011, January 25\u201330). People detection in RGB-D data. Proceedings of the 2011 IEEE\/RSJ International Conference on Intelligent Robots and Systems, San Francisco, CA, USA.","DOI":"10.1109\/IROS.2011.6095074"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Jafari, O.H., Mitzel, D., and Leibe, B. (June, January 31). Real-time RGB-D based people detection and tracking for mobile robots and head-worn cameras. Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA), Hong Kong, China.","DOI":"10.1109\/ICRA.2014.6907688"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Choi, W., Pantofaru, C., and Savarese, S. (2011, January 6\u201313). Detecting and tracking people using an rgb-d camera via multiple detector fusion. Proceedings of the 2011 IEEE International Conference on Computer Vision Workshops (ICCV Workshops), Barcelona, Spain.","DOI":"10.1109\/ICCVW.2011.6130370"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Benenson, R., Mathias, M., Timofte, R., and Van Gool, L. (2012, January 16\u201321). Pedestrian detection at 100 frames per second. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248017"},{"key":"ref_10","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012). Imagenet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25, Neural Information Processing Systems Foundation, Inc."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the CVPR 2016, Las Vegas, NA, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_12","unstructured":"Szegedy, C., Toshev, A., and Erhan, D. (2013). Deep neural networks for object detection. Advances in Neural Information Processing Systems 26, Neural Information Processing Systems Foundation, Inc."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 24\u201327). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the CVPR 2014, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_15","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems 28, Neural Information Processing Systems Foundation, Inc."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (July, January 26). You only look once: Unified, real-time object detection. Proceedings of the CVPR 2016, Las Vegas, NA, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 22\u201325). YOLO9000: Better, Faster, Stronger. Proceedings of the CVPR 2017, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_18","unstructured":"Redmon, J., and Farhadi, A. (arXiv, 2018). Yolov3: An incremental improvement, arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016). SSD: Single shot multibox detector. Lecture Notes in Computer Science, Springer.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_20","unstructured":"Wagner, J., Fischer, V., Herman, M., and Behnke, S. (2016, January 27\u201329). Multispectral pedestrian detection using deep fusion convolutional neural networks. Proceedings of the 24th European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN), Bruges, Belgium."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Liu, J., Zhang, S., Wang, S., and Metaxas, D. (arXiv, 2016). Multispectral Deep Neural Networks for Pedestrian Detection, arXiv.","DOI":"10.5244\/C.30.73"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"K\u00f6nig, D., Adam, M., Jarvers, C., Layher, G., Neumann, H., and Teutsch, M. (2017, January 21). Fully Convolutional Region Proposal Networks for Multispectral Person Detection. Proceedings of the CVPR Workshops, Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.36"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Vandersteegen, M., Van Beeck, K., and Goedem\u00e9, T. (2018, January 27\u201329). Real-time multispectral pedestrian detection with a single-pass deep neural network. Proceedings of the 15th International Conference on Image Analysis and Recognition (ICIAR), Varzim, Portugal.","DOI":"10.1007\/978-3-319-93000-8_47"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Schwarz, M., Schulz, H., and Behnke, S. (2015, January 26\u201330). RGB-D object recognition and pose estimation based on pre-trained convolutional neural network features. Proceedings of the 2015 IEEE International Conference on Robotics and Automation (ICRA), Seattle, WA, USA.","DOI":"10.1109\/ICRA.2015.7139363"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Eitel, A., Springenberg, J.T., Spinello, L., Riedmiller, M., and Burgard, W. (October, January 28). Multimodal deep learning for robust rgb-d object recognition. Proceedings of the 2015 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany.","DOI":"10.1109\/IROS.2015.7353446"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Gupta, S., Girshick, R., Arbel\u00e1ez, P., and Malik, J. (2014, January 6\u201312). Learning rich features from RGB-D images for object detection and segmentation. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10584-0_23"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhou, K., Paiement, A., and Mirmehdi, M. (2017, January 8\u201312). Detecting humans in RGB-D data with CNNs. Proceedings of the 2017 Fifteenth IAPR International Conference on Machine Vision Applications (MVA), Nagoya, Japan.","DOI":"10.23919\/MVA.2017.7986862"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Ophoff, T., Van Beeck, K., and Goedem\u00e9, T. (2018, January 27\u201330). Improving Real-Time Pedestrian Detectors with RGB+Depth Fusion. Proceedings of the AVSS\u2014MSS Workshop, Auckland, New Zealand.","DOI":"10.1109\/AVSS.2018.8639110"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). ImageNet: A Large-Scale Hierarchical Image Database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common objects in context. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Bagautdinov, T., Fleuret, F., and Fua, P. (2015, January 8\u201310). Probability Occupancy Maps for Occluded Depth Images. Proceedings of the CVPR 2015, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298900"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for Autonomous Driving? The KITTI Vision Benchmark Suite. Proceedings of the CVPR 2012, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"328","DOI":"10.1109\/TPAMI.2007.1166","article-title":"Stereo Processing by Semiglobal Matching and Mutual Information","volume":"30","author":"Hirschmuller","year":"2008","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_34","unstructured":"Bradski, G. (2018, October 10). The OpenCV Library. Available online: https:\/\/opencv.org\/."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"323","DOI":"10.1016\/j.resconrec.2017.06.022","article-title":"Ease of disassembly of products to support circular economy strategies","volume":"135","author":"Vanegas","year":"2018","journal-title":"Resour. Conserv. Recycl."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/4\/866\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:33:18Z","timestamp":1760185998000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/4\/866"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,2,19]]},"references-count":35,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2019,2]]}},"alternative-id":["s19040866"],"URL":"https:\/\/doi.org\/10.3390\/s19040866","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,2,19]]}}}