{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,17]],"date-time":"2026-02-17T14:15:24Z","timestamp":1771337724617,"version":"3.50.1"},"reference-count":24,"publisher":"MDPI AG","issue":"20","license":[{"start":{"date-parts":[[2022,10,20]],"date-time":"2022-10-20T00:00:00Z","timestamp":1666224000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100006769","name":"Russian Science Foundation","doi-asserted-by":"publisher","award":["18-71-10065"],"award-info":[{"award-number":["18-71-10065"]}],"id":[{"id":"10.13039\/501100006769","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006769","name":"Russian Science Foundation","doi-asserted-by":"publisher","award":["FFZF-2022-0005"],"award-info":[{"award-number":["FFZF-2022-0005"]}],"id":[{"id":"10.13039\/501100006769","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Russian State Research","award":["18-71-10065"],"award-info":[{"award-number":["18-71-10065"]}]},{"name":"Russian State Research","award":["FFZF-2022-0005"],"award-info":[{"award-number":["FFZF-2022-0005"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In this paper, we present a two stages solution to 3D vehicle detection and segmentation. The first stage depends on the combination of EfficientNetB3 architecture with multiparallel residual blocks (inspired by CenterNet architecture) for 3D localization and poses estimation for vehicles on the scene. The second stage takes the output of the first stage as input (cropped car images) to train EfficientNet B3 for the image recognition task. Using predefined 3D Models, we substitute each vehicle on the scene with its match using the rotation matrix and translation vector from the first stage to get the 3D detection bounding boxes and segmentation masks. We trained our models on an open-source dataset (ApolloCar3D). Our method outperforms all published solutions in terms of 6 degrees of freedom error (6 DoF err).<\/jats:p>","DOI":"10.3390\/s22207990","type":"journal-article","created":{"date-parts":[[2022,10,21]],"date-time":"2022-10-21T00:34:30Z","timestamp":1666312470000},"page":"7990","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["3D Vehicle Detection and Segmentation Based on EfficientNetB3 and CenterNet Residual Blocks"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6503-1447","authenticated-orcid":false,"given":"Alexey","family":"Kashevnik","sequence":"first","affiliation":[{"name":"St. Petersburg Federal Research Center of the Russian Academy of Sciences, SPC RAS, 199178 St. Petersburg, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3073-9506","authenticated-orcid":false,"given":"Ammar","family":"Ali","sequence":"additional","affiliation":[{"name":"Information Technology and Programming Faculty, ITMO University, 197101 St. Petersburg, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,10,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhang, H., Ji, H., Zheng, A., Hwang, J.-N., and Hwang, R.-H. (2021, January 11\u201317). Monocular 3D Localization of Vehicles in Road Scenes. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00320"},{"key":"ref_2","unstructured":"Jaesung, C., Kyungdon, J., Fran\u00e7ois, R., Gyumin, S., and Inso, K. (2019, January 22\u201326). Segment2Regress: Monocular 3D Vehicle Localization in Two Stages. Proceedings of the Robotics: Science and Systems (RSS), Breisgau, Germany."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Fadadu, S., Pandey, S., Hegde, D., Shi, Y., Chou, F., Djuric, N., and Vallespi-Gonzalez, C. (2022, January 4\u20138). Multi-View Fusion of Sensor Data for Improved Perception and Prediction in Autonomous Driving. Proceedings of the 2022 IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA.","DOI":"10.1109\/WACV51458.2022.00335"},{"key":"ref_4","unstructured":"Zhu, H., Deng, J., Zhang, Y., Ji, J., Mao, Q., Li, H., and Zhang, Y. (2021). VPFNet: Improving 3D Object Detection with Virtual Point based LiDAR and Stereo Data Fusion. arXiv."},{"key":"ref_5","unstructured":"Su, Z., Tan, P.S., and Wang, Y. (2021). DV-Det: Efficient 3D Point Cloud Object Detection with Dynamic Voxelization. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Noh, J., Lee, S., and Ham, B. (2021, January 19\u201325). HVPR: Hybrid Voxel-Point Representation for Single-stage 3D Object Detection. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01437"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Chen, Y., Dai, H., and Ding, Y. (2022). Pseudo-Stereo for Monocular 3D Object Detection in Autonomous Driving. arXiv.","DOI":"10.1109\/CVPR52688.2022.00096"},{"key":"ref_8","unstructured":"Li, W., Li, Z., Yi, Z., Zhi, Z., Tong, H., and Mu, L. (2021). Progressive Coordinate Transforms for Monocular 3D Object Detection. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Liu, Z., Zhou, D., Lu, F., Fang, J., and Zhang, L. (2021, January 11\u201317). AutoShape: Real-Time Shape-Aware Monocular 3D Object Detection. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01535"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Julca-Aguilar, F., Taylor, J., Bijelic, M., Mannan, F., Tseng, E., and Heide, F. (2021, January 11\u201317). Gated3D: Monocular 3D Object Detection from Temporal Illumination Cues. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00293"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Babolhavaeji, A., and Fanaei, M. (2020, January 16\u201318). Multi-Stage CNN-Based Monocular 3D Vehicle Localization and Orientation Estimation. Proceedings of the 2020 International Conference on Computational Science and Computational Intelligence (CSCI), Las Vegas, NV, USA.","DOI":"10.1109\/CSCI51800.2020.00295"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Li, P., Chen, X., and Shen, S. (2019, January 16\u201320). Stereo R-CNN Based 3D Object Detection for Autonomous Driving. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00783"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chen, X., Kundu, K., Zhang, Z., Ma, H., Fidler, S., and Urtasun, R. (2016, January 27\u201330). Monocular 3D Object Detection for Autonomous Driving. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.236"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Shi, S., Guo, C., Jiang, L., Wang, Z., Shi, J., Wang, X., and Li, H. (2020, January 13\u201319). PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01054"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Shi, W., and Rajkumar, R. (2020, January 13\u201319). Point-GNN: Graph Neural Network for 3D Object Detection in a Point Cloud. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00178"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ke, L., Li, S., Sun, Y., Tai, Y., and Tang, C. (2020). GSNet: Joint Vehicle Pose and Shape Reconstruction with Geometrical and Scene-aware Supervision. arXiv.","DOI":"10.1007\/978-3-030-58555-6_31"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zauss, D., Kreiss, S., and Alahi, A. (2021, January 11\u201317). Keypoint Communities. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01087"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"118493","DOI":"10.1016\/j.eswa.2022.118493","article-title":"An active contour model driven by adaptive local pre-fitting energy function based on Jeffreys divergence for image segmentation","volume":"210","author":"Ge","year":"2022","journal-title":"Expert Syst. Appl."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1016\/j.patrec.2022.04.025","article-title":"A hybrid active contour model based on pre-fitting energy and adaptive functions for fast image segmentation","volume":"158","author":"Ge","year":"2022","journal-title":"Pattern Recogn. Lett."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"115633","DOI":"10.1016\/j.eswa.2021.115633","article-title":"A level set method based on additive bias correction for image segmentation","volume":"185","author":"Weng","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1718","DOI":"10.1109\/TITS.2020.2980855","article-title":"An Efficient and Scalable Simulation Model for Autonomous Vehicles with Economical Hardware","volume":"22","author":"Irfan","year":"2021","journal-title":"IEEE Trans. Intell. Trans. Syst."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Weber, M., F\u00fcrst, M., and Z\u00f6llner, J.M. (2019, January 9\u201312). Direct 3D Detection of Vehicles in Monocular Images with a CNN based 3D Decoder. Proceedings of the 2019 IEEE Intelligent Vehicles Symposium (IV), Paris, France.","DOI":"10.1109\/IVS.2019.8814198"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Song, X., Wang, P., Zhou, D., Zhu, R., Guan, C., Dai, Y., Su, H., Li, H., and Yang, R. (2019). Apollocar3D: A large 3d car instance understanding benchmark for autonomous driving. arXiv.","DOI":"10.1109\/CVPR.2019.00560"},{"key":"ref_24","unstructured":"Tan, M., and Le, Q.V. (2019). EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/20\/7990\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:57:47Z","timestamp":1760144267000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/20\/7990"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,20]]},"references-count":24,"journal-issue":{"issue":"20","published-online":{"date-parts":[[2022,10]]}},"alternative-id":["s22207990"],"URL":"https:\/\/doi.org\/10.3390\/s22207990","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,20]]}}}