{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T10:56:34Z","timestamp":1772794594062,"version":"3.50.1"},"reference-count":63,"publisher":"PeerJ","license":[{"start":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T00:00:00Z","timestamp":1772755200000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"abstract":"<jats:p>To navigate safely through complicated traffic systems, autonomous driving technologies depend on accurate object detection in real time. The algorithms currently being used to identify objects have difficulty finding the right balance between accuracies and computation efficiencies, especially when analyzing high-resolution images under varying weather and lighting conditions. To address these issues, this article presents Real-time Enhanced Vision Analysis Detection Transformer (REVA-DETR), a new approach to real-time object detection created specifically for use in autonomous driving systems. Four new technologies were integrated into the REVA-DETR framework to overcome the limitations of existing technologies: Multi-Scale Dilated Receptive Network (MS-DRNET) was developed to improve the ability to identify objects using multiple scales and obtain more context; Efficient Additive Attention (EAA) enabled for the first time the use of linear-complexity feature interactions with existing attention mechanisms; FusBoost Net was engineered to allow for multi-scale feature fusion; Minimum Point Distance IoU loss (MPD-IoU) permitted for improved bounding box regression; and lastly an experimental evaluation of REVA-DETR on both Berkeley DeepDrive 100K (BDD100K) and Self-supervised Object Detection and Annotation 10M (SODA10M) datasets provided evidence of superior performance relative to the original RT-DETR-R18\u00a0model with mean Average Precision at 50% IoU (mAP50) equal to 0.634 and mean Average Precision at 50%-95% IoU (mAP50-95) equal to 0.407 and therefore equal to improved performance of 4.96% and 24%, respectively, as well as providing continuous real-time operation at 33.2 frames per second (FPS). Additionally, the cross-dataset validation provided evidence of the ability to generalise well and balance the trade-off between detection accuracies and real-time computation needs necessary for effective deployment of autonomous driving systems.<\/jats:p>","DOI":"10.7717\/peerj-cs.3704","type":"journal-article","created":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T08:56:42Z","timestamp":1772787402000},"page":"e3704","source":"Crossref","is-referenced-by-count":0,"title":["Research on image recognition in autonomous driving systems based on the REVA-DETR (real-time enhanced vision analysis detection transformer) algorithm"],"prefix":"10.7717","volume":"12","author":[{"given":"Minhan","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rui","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"4443","published-online":{"date-parts":[[2026,3,6]]},"reference":[{"key":"10.7717\/peerj-cs.3704\/ref-1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2004.10934","article-title":"YOLOv4: optimal speed and accuracy of object detection","author":"Bochkovskiy","year":"2020"},{"key":"10.7717\/peerj-cs.3704\/ref-2","first-page":"11621","article-title":"nuScenes: a multimodal dataset for autonomous driving","author":"Caesar","year":"2020"},{"key":"10.7717\/peerj-cs.3704\/ref-3","first-page":"6154","article-title":"Cascade R-CNN: delving into high quality object detection","author":"Cai","year":"2018"},{"key":"10.7717\/peerj-cs.3704\/ref-4","first-page":"213","article-title":"End-to-end object detection with transformers","author":"Carion","year":"2020"},{"issue":"2","key":"10.7717\/peerj-cs.3704\/ref-5","doi-asserted-by":"publisher","first-page":"1046","DOI":"10.1109\/TIV.2022.3223131","article-title":"Milestones in autonomous driving and intelligent vehicles: survey of surveys","volume":"8","author":"Chen","year":"2023","journal-title":"IEEE Transactions on Intelligent Vehicles"},{"key":"10.7717\/peerj-cs.3704\/ref-6","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1906.07155","article-title":"MMDetection: open MMLab detection toolbox and benchmark","author":"Chen","year":"2019"},{"key":"10.7717\/peerj-cs.3704\/ref-7","first-page":"886","article-title":"Histograms of oriented gradients for human detection","author":"Dalal","year":"2005"},{"key":"10.7717\/peerj-cs.3704\/ref-8","article-title":"An image is worth 16 x 16 words: transformers for image recognition at scale","author":"Dosovitskiy","year":"2021"},{"issue":"1","key":"10.7717\/peerj-cs.3704\/ref-9","doi-asserted-by":"publisher","first-page":"119","DOI":"10.1006\/jcss.1997.1504","article-title":"A decision-theoretic generalization of on-line learning and an application to boosting","volume":"55","author":"Freund","year":"1997","journal-title":"Journal of Computer and System Sciences"},{"key":"10.7717\/peerj-cs.3704\/ref-10","first-page":"3354","article-title":"Are we ready for autonomous driving? The KITTI vision benchmark suite","author":"Geiger","year":"2012"},{"key":"10.7717\/peerj-cs.3704\/ref-11","first-page":"7036","article-title":"NAS-FPN: learning scalable feature pyramid architecture for object detection","author":"Ghiasi","year":"2019"},{"key":"10.7717\/peerj-cs.3704\/ref-12","first-page":"1440","article-title":"Fast R-CNN","author":"Girshick","year":"2015"},{"key":"10.7717\/peerj-cs.3704\/ref-13","first-page":"580","article-title":"Rich feature hierarchies for accurate object detection and semantic segmentation","author":"Girshick","year":"2014"},{"key":"10.7717\/peerj-cs.3704\/ref-14","article-title":"SODA10M: a large-scale 2D self\/semi-supervised object detection dataset for autonomous driving","author":"Han","year":"2021"},{"key":"10.7717\/peerj-cs.3704\/ref-15","first-page":"2961","article-title":"Mask R-CNN","author":"He","year":"2017"},{"issue":"10","key":"10.7717\/peerj-cs.3704\/ref-16","doi-asserted-by":"publisher","first-page":"2702","DOI":"10.1109\/tpami.2019.2926463","article-title":"The ApolloScape open dataset for autonomous driving and its application","volume":"42","author":"Huang","year":"2020","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"1\u20133","key":"10.7717\/peerj-cs.3704\/ref-17","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1561\/0600000079","article-title":"Computer vision for autonomous vehicles: problems, datasets and state of the art","volume":"12","author":"Janai","year":"2020","journal-title":"Foundations and Trends in Computer Graphics and Vision"},{"issue":"11","key":"10.7717\/peerj-cs.3704\/ref-18","doi-asserted-by":"publisher","first-page":"1066","DOI":"10.1016\/j.procs.2022.01.135","article-title":"A review of yolo algorithm developments","volume":"199","author":"Jiang","year":"2022","journal-title":"Procedia Computer Science"},{"key":"10.7717\/peerj-cs.3704\/ref-19","article-title":"YOLO by ultralytics","author":"Jocher","year":"2023"},{"key":"10.7717\/peerj-cs.3704\/ref-20","doi-asserted-by":"publisher","first-page":"287","DOI":"10.1109\/ICCPS.2018.00035","article-title":"Autoware on board: enabling autonomous vehicles with embedded systems","author":"Kato","year":"2018"},{"key":"10.7717\/peerj-cs.3704\/ref-21","first-page":"12697","article-title":"PointPillars: fast encoders for object detection from point clouds","author":"Lang","year":"2019"},{"issue":"7553","key":"10.7717\/peerj-cs.3704\/ref-22","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"10.7717\/peerj-cs.3704\/ref-23","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2203.01305","article-title":"DN-DETR: towards a better training for object detection via noise learning","author":"Li","year":"2022"},{"key":"10.7717\/peerj-cs.3704\/ref-24","first-page":"1","article-title":"BEVFormer: learning bird\u2019s-eye-view representation from multi-camera images via spatiotemporal transformers","author":"Li","year":"2022"},{"issue":"8","key":"10.7717\/peerj-cs.3704\/ref-25","doi-asserted-by":"publisher","first-page":"7041","DOI":"10.1109\/tcsvt.2023.3318401","article-title":"Meta-learning based domain prior with application to optical-ISAR image translation","volume":"34","author":"Liao","year":"2024","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"10.7717\/peerj-cs.3704\/ref-26","first-page":"2117","article-title":"Feature pyramid networks for object detection","author":"Lin","year":"2017"},{"key":"10.7717\/peerj-cs.3704\/ref-27","first-page":"2980","article-title":"Focal loss for dense object detection","author":"Lin","year":"2017"},{"key":"10.7717\/peerj-cs.3704\/ref-28","first-page":"21","article-title":"SSD: single shot multibox detector","author":"Liu","year":"2016"},{"key":"10.7717\/peerj-cs.3704\/ref-29","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2201.12329","article-title":"DAB-DETR: dynamic anchor boxes are better queries for DETR","author":"Liu","year":"2022"},{"key":"10.7717\/peerj-cs.3704\/ref-30","first-page":"10012","article-title":"Swin transformer: hierarchical vision transformer using shifted windows","author":"Liu","year":"2021"},{"key":"10.7717\/peerj-cs.3704\/ref-31","first-page":"8759","article-title":"Path aggregation network for instance segmentation","author":"Liu","year":"2018"},{"key":"10.7717\/peerj-cs.3704\/ref-32","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2111.12419","article-title":"NAM: normalization-based attention module","author":"Liu","year":"2021"},{"key":"10.7717\/peerj-cs.3704\/ref-33","first-page":"2774","article-title":"BEVFusion: multi-task multi-sensor fusion with unified bird\u2019s-eye view representation","author":"Liu","year":"2023"},{"key":"10.7717\/peerj-cs.3704\/ref-34","first-page":"531","article-title":"PETR: position embedding transformation for multi-view 3d object detection","author":"Liu","year":"2022"},{"issue":"2","key":"10.7717\/peerj-cs.3704\/ref-35","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1023\/b:visi.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"International Journal of Computer Vision"},{"key":"10.7717\/peerj-cs.3704\/ref-36","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2307.07662","article-title":"Mpdiou: a loss for efficient and accurate bounding box regression","author":"Ma","year":"2023"},{"key":"10.7717\/peerj-cs.3704\/ref-37","first-page":"3651","article-title":"Conditional DETR for fast training convergence","author":"Meng","year":"2021"},{"key":"10.7717\/peerj-cs.3704\/ref-38","first-page":"194","article-title":"Lift, splat, shoot: encoding images from arbitrary camera rigs by implicitly unprojecting to 3D","author":"Philion","year":"2020"},{"key":"10.7717\/peerj-cs.3704\/ref-39","first-page":"779","article-title":"You only look once: unified, real-time object detection","author":"Redmon","year":"2016"},{"key":"10.7717\/peerj-cs.3704\/ref-40","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1804.02767","article-title":"YOLOv3: an incremental improvement","author":"Redmon","year":"2018"},{"issue":"6","key":"10.7717\/peerj-cs.3704\/ref-41","doi-asserted-by":"publisher","first-page":"1137","DOI":"10.1109\/tpami.2016.2577031","article-title":"Faster R-CNN: towards real-time object detection with region proposal networks","volume":"28","author":"Ren","year":"2015","journal-title":"Advances in Neural Information Processing Systems"},{"key":"10.7717\/peerj-cs.3704\/ref-42","first-page":"658","article-title":"Generalized intersection over union: a metric and a loss for bounding box regression","author":"Rezatofighi","year":"2019"},{"issue":"1","key":"10.7717\/peerj-cs.3704\/ref-43","doi-asserted-by":"publisher","first-page":"908","DOI":"10.1109\/tvt.2018.2884525","article-title":"V2V routing in a VANET based on the autoregressive integrated moving average model","volume":"68","author":"Sun","year":"2018","journal-title":"IEEE Transactions on Vehicular Technology"},{"key":"10.7717\/peerj-cs.3704\/ref-44","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-34372-9","volume-title":"Computer vision: algorithms and applications","author":"Szeliski","year":"2022"},{"key":"10.7717\/peerj-cs.3704\/ref-45","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2003.06404","article-title":"A survey of end-to-end driving: architectures and training methods","author":"Tampuu","year":"2020"},{"key":"10.7717\/peerj-cs.3704\/ref-46","first-page":"10781","article-title":"EfficientDet: scalable and efficient object detection","author":"Tan","year":"2020"},{"key":"10.7717\/peerj-cs.3704\/ref-47","first-page":"5998","article-title":"Attention is all you need","author":"Vaswani","year":"2017"},{"key":"10.7717\/peerj-cs.3704\/ref-48","doi-asserted-by":"publisher","first-page":"511","DOI":"10.1109\/CVPR.2001.990517","article-title":"Rapid object detection using a boosted cascade of simple features","volume":"1","author":"Viola","year":"2001"},{"key":"10.7717\/peerj-cs.3704\/ref-49","doi-asserted-by":"publisher","first-page":"15909","DOI":"10.1109\/CVPR52733.2024.01506","article-title":"RepViT: revisiting mobile CNN from ViT perspective","author":"Wang","year":"2024"},{"key":"10.7717\/peerj-cs.3704\/ref-50","first-page":"3007","article-title":"CARAFE: content-aware reassembly of features","author":"Wang","year":"2019"},{"key":"10.7717\/peerj-cs.3704\/ref-51","first-page":"180","article-title":"DETR3D: 3D object detection from multi-view images via 3D-to-2D queries","author":"Wang","year":"2022"},{"key":"10.7717\/peerj-cs.3704\/ref-52","first-page":"14408","article-title":"InternImage: exploring large-scale vision foundation models with deformable convolutions","author":"Wang","year":"2023"},{"key":"10.7717\/peerj-cs.3704\/ref-53","first-page":"913","article-title":"FCOS3D: fully convolutional one-stage monocular 3d object detection","author":"Wang","year":"2021"},{"issue":"6","key":"10.7717\/peerj-cs.3704\/ref-54","doi-asserted-by":"publisher","first-page":"8065","DOI":"10.1109\/TITS.2025.3558085","article-title":"Human-Factors-in-Aviation-Loop: multimodal deep learning for pilot situation awareness analysis using gaze position and flight control data","volume":"26","author":"Xu","year":"2025","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"10.7717\/peerj-cs.3704\/ref-55","first-page":"2633","article-title":"Bdd100k: a diverse driving dataset for heterogeneous multitask learning","author":"Yu","year":"2020"},{"key":"10.7717\/peerj-cs.3704\/ref-56","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1511.07122","article-title":"Multi-scale context aggregation by dilated convolutions","author":"Yu","year":"2015"},{"key":"10.7717\/peerj-cs.3704\/ref-57","doi-asserted-by":"publisher","first-page":"58443","DOI":"10.1109\/access.2020.2983149","article-title":"A survey of autonomous driving: common practices and emerging technologies","volume":"8","author":"Yurtsever","year":"2020","journal-title":"IEEE Access"},{"key":"10.7717\/peerj-cs.3704\/ref-58","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2203.03605","article-title":"DINO: DETR with improved denoising anchor boxes for end-to-end object detection","author":"Zhang","year":"2022"},{"key":"10.7717\/peerj-cs.3704\/ref-59","first-page":"16965","article-title":"Detrs beat yolos on real-time object detection","author":"Zhao","year":"2024"},{"issue":"7","key":"10.7717\/peerj-cs.3704\/ref-60","doi-asserted-by":"publisher","first-page":"12993","DOI":"10.1609\/aaai.v34i07.6999","article-title":"Distance-IoU loss: faster and better learning for bounding box regression","volume":"34","author":"Zheng","year":"2020","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"issue":"8","key":"10.7717\/peerj-cs.3704\/ref-61","doi-asserted-by":"publisher","first-page":"8574","DOI":"10.1109\/tcyb.2021.3095305","article-title":"Enhancing geometric factors in model learning and inference for object detection and instance segmentation","volume":"52","author":"Zheng","year":"2022","journal-title":"IEEE Transactions on Cybernetics"},{"key":"10.7717\/peerj-cs.3704\/ref-62","first-page":"4490","article-title":"VoxelNet: end-to-end learning for point cloud-based 3D object detection","author":"Zhou","year":"2018"},{"key":"10.7717\/peerj-cs.3704\/ref-63","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.04159","article-title":"Deformable DETR: deformable transformers for end-to-end object detection","author":"Zhu","year":"2020"}],"container-title":["PeerJ Computer Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/peerj.com\/articles\/cs-3704.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/peerj.com\/articles\/cs-3704.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/peerj.com\/articles\/cs-3704.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/peerj.com\/articles\/cs-3704.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T08:56:46Z","timestamp":1772787406000},"score":1,"resource":{"primary":{"URL":"https:\/\/peerj.com\/articles\/cs-3704"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,6]]},"references-count":63,"alternative-id":["10.7717\/peerj-cs.3704"],"URL":"https:\/\/doi.org\/10.7717\/peerj-cs.3704","archive":["CLOCKSS","LOCKSS","Portico"],"relation":{},"ISSN":["2376-5992"],"issn-type":[{"value":"2376-5992","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,6]]},"article-number":"e3704"}}