{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T19:57:07Z","timestamp":1786564627319,"version":"3.56.0"},"reference-count":42,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,9,19]],"date-time":"2024-09-19T00:00:00Z","timestamp":1726704000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>When it comes to interpreting visual input, intelligent systems make use of contextual scene learning, which significantly improves both resilience and context awareness. The management of enormous amounts of data is a driving force behind the growing interest in computational frameworks, particularly in the context of autonomous cars.<\/jats:p><\/jats:sec><jats:sec><jats:title>Method<\/jats:title><jats:p>The purpose of this study is to introduce a novel approach known as Deep Fused Networks (DFN), which improves contextual scene comprehension by merging multi-object detection and semantic analysis.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>To enhance accuracy and comprehension in complex situations, DFN makes use of a combination of deep learning and fusion techniques. With a minimum gain of 6.4% in accuracy for the SUN-RGB-D dataset and 3.6% for the NYU-Dv2 dataset.<\/jats:p><\/jats:sec><jats:sec><jats:title>Discussion<\/jats:title><jats:p>Findings demonstrate considerable enhancements in object detection and semantic analysis when compared to the methodologies that are currently being utilized.<\/jats:p><\/jats:sec>","DOI":"10.3389\/fnbot.2024.1427786","type":"journal-article","created":{"date-parts":[[2024,9,19]],"date-time":"2024-09-19T16:24:31Z","timestamp":1726763071000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":67,"title":["Multi-modal remote perception learning for object sensory data"],"prefix":"10.3389","volume":"18","author":[{"given":"Nouf Abdullah","family":"Almujally","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Adnan Ahmed","family":"Rafique","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Naif","family":"Al Mudawi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Abdulwahab","family":"Alazeb","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohammed","family":"Alonazi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Asaad","family":"Algarni","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ahmad","family":"Jalal","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hui","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2024,9,19]]},"reference":[{"key":"ref1","doi-asserted-by":"publisher","first-page":"1143","DOI":"10.1007\/s42835-020-00650-z","article-title":"Multi-objects detection and segmentation for scene understanding based on texton forest and kernel sliding perceptron","volume":"16","author":"Ahmed","year":"2021","journal-title":"J. Electr. Eng. Technol."},{"key":"ref2","doi-asserted-by":"publisher","first-page":"14819","DOI":"10.1109\/ACCESS.2022.3148036","article-title":"CNN-based object recognition and tracking system to assist visually impaired people","volume":"10","author":"Ashiq","year":"2022","journal-title":"IEEE Access"},{"key":"ref3","article-title":"Centroid based concept learning for rgb-d indoor scene classification","volume":"2019","author":"Ayub","year":"2019","journal-title":"arXiv preprint"},{"key":"ref4","first-page":"7892","article-title":"Road: reality oriented adaptation for semantic segmentation of urban scenes","volume":"2018","author":"Chen","year":"2018","journal-title":"Proc. IEEE Conf. Comput. Vision Pattern Recogn."},{"key":"ref5","doi-asserted-by":"crossref","first-page":"2096","DOI":"10.3390\/rs15082096","article-title":"A multi-feature fusion and attention network for multi-scale object detection in remote sensing images","volume":"15","author":"Cheng","year":"2023","journal-title":"Remote Sens."},{"key":"ref6","first-page":"10972","article-title":"On human scene-sketch and its complementarity with photo and text","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chowdhury","year":"2023"},{"key":"ref7","doi-asserted-by":"publisher","first-page":"85","DOI":"10.1007\/s13748-019-00203-0","article-title":"Convolutional neural network: a review of models, methodologies and applications to object detection","volume":"9","author":"Dhillon","year":"2020","journal-title":"Prog. Artif. Intell."},{"key":"ref8","doi-asserted-by":"publisher","first-page":"9243","DOI":"10.1007\/s11042-022-13644-y","article-title":"Object detection using YOLO: challenges, architectural successors, datasets and applications","volume":"82","author":"Diwan","year":"2023","journal-title":"Multimed. Tools Appl."},{"key":"ref9","first-page":"11836","article-title":"Ranslate-to-recognize networks for rgb-d scene recognition","volume-title":"Proc. of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Du","year":"2019"},{"key":"ref10","first-page":"2014","article-title":"On the importance of visual context for data augmentation in scene understanding","volume":"436","author":"Dvornik","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref11","doi-asserted-by":"crossref","first-page":"1341","DOI":"10.1109\/TITS.2020.2972974","article-title":"Deep multi-modal object detection and semantic segmentation for autonomous driving: datasets, methods, and challenges","volume":"22","author":"Feng","year":"2020","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref9001","doi-asserted-by":"crossref","first-page":"42131","DOI":"10.1109\/ACCESS.2020.2976686","article-title":"Large-scale synthetic urban dataset for aerial scene understanding","volume":"8","author":"Gao","year":"2020","journal-title":"IEEE Access"},{"key":"ref12","doi-asserted-by":"publisher","first-page":"103661","DOI":"10.1016\/j.compind.2022.103661","article-title":"Deep learning-based object detection in augmented reality: a systematic review","volume":"139","author":"Ghasemi","year":"2022","journal-title":"Comput. Ind."},{"key":"ref13","doi-asserted-by":"publisher","first-page":"4267","DOI":"10.1007\/s00371-022-02589-w","article-title":"Multi-level feature fusion pyramid network for object detection","volume":"39","author":"Guo","year":"2023","journal-title":"Vis. Comput."},{"key":"ref14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1155\/2020\/7286187","article-title":"Learning feature fusion in deep learning-based object detector","volume":"2020","author":"Hassan","year":"2020","journal-title":"J. Eng."},{"key":"ref15","first-page":"2094","article-title":"Prompting multi-modal image segmentation with semantic grouping","volume-title":"Proceedings of the AAAI conference on artificial intelligence","author":"He","year":"2024"},{"key":"ref16","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1016\/j.isprsjprs.2023.04.003","article-title":"AST: adaptive self-supervised transformer for optical remote sensing representation","volume":"200","author":"He","year":"2023","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref17","doi-asserted-by":"publisher","first-page":"3820","DOI":"10.1109\/TPAMI.2020.2992222","article-title":"Contextual translation embedding for visual relationship detection and scene graph generation","volume":"43","author":"Hung","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref18","doi-asserted-by":"publisher","first-page":"147808","DOI":"10.1109\/ACCESS.2020.3016008","article-title":"Leveraging contextual information for monocular depth estimation","volume":"8","author":"Kim","year":"2020","journal-title":"IEEE Access"},{"key":"ref19","doi-asserted-by":"publisher","first-page":"55546","DOI":"10.1109\/ACCESS.2022.3177628","article-title":"YOLO-G: a lightweight network model for improving the performance of military targets detection","volume":"10","author":"Kong","year":"2022","journal-title":"IEEE Access"},{"key":"ref20","doi-asserted-by":"publisher","first-page":"183","DOI":"10.1109\/LGRS.2017.2779469","article-title":"Scene classification based on two-stage deep feature fusion","volume":"15","author":"Liu","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref21","first-page":"45","article-title":"Autonomous vehicle assisted by heads up display (HUD) with augmented reality based on machine learning techniques","volume-title":"Virtual and augmented reality for automobile industry: Innovation vision and applications","author":"Murugan","year":"2022"},{"key":"ref22","doi-asserted-by":"publisher","first-page":"207","DOI":"10.3390\/s17010207","article-title":"Object detection and classification by decision-level fusion for intelligent vehicle systems","volume":"17","author":"Oh","year":"2017","journal-title":"Sensors"},{"key":"ref23","doi-asserted-by":"publisher","first-page":"553","DOI":"10.1007\/s12145-021-00746-8","article-title":"Object detection in satellite images by faster R-CNN incorporated with enhanced ROI pooling (FrRNet-ERoI) framework","volume":"15","author":"Pazhani","year":"2022","journal-title":"Earth Sci. Inf."},{"key":"ref24","doi-asserted-by":"publisher","first-page":"1197","DOI":"10.1007\/s00371-013-0886-1","article-title":"Computer vision-based object recognition for the visually impaired in an indoors environment: a survey","volume":"30","author":"Rabia","year":"2014","journal-title":"Vis. Comput."},{"key":"ref25","first-page":"3234","article-title":"The synthia dataset: a large collection of synthetic images for semantic segmentation of urban scenes","author":"Ros","year":"2016"},{"key":"ref26","doi-asserted-by":"publisher","first-page":"82066","DOI":"10.1109\/ACCESS.2020.2989863","article-title":"FOSNet: an end-to-end trainable deep neural network for scene recognition","volume":"8","author":"Seong","year":"2020","journal-title":"IEEE Access"},{"key":"ref27","first-page":"746","article-title":"Indoor segmentation and support inference from rgbd images","volume-title":"European conference on computer vision","author":"Silberman","year":"2012"},{"key":"ref28","doi-asserted-by":"publisher","first-page":"104117","DOI":"10.1016\/j.imavis.2021.104117","article-title":"Weighted boxes fusion: Ensembling boxes from different object detection models","volume":"107","author":"Solovyev","year":"2021","journal-title":"Image Vis. Comput."},{"key":"ref29","first-page":"600","article-title":"RGB-D scene recognition with object-to-object relation","author":"Song","year":"2017"},{"key":"ref30","doi-asserted-by":"crossref","first-page":"980","DOI":"10.1109\/TIP.2018.2872629","article-title":"Learning effective RGB-D representations for scene recognition","volume":"28","author":"Song","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref31","first-page":"567","article-title":"Sun rgb-d: A rgb-d scene understanding benchmark suite","author":"Song","year":"2015"},{"key":"ref32","doi-asserted-by":"crossref","first-page":"13708","DOI":"10.1109\/ICRA48506.2021.9561110","article-title":"Joint object detection and multi-object tracking with graph neural networks","volume-title":"IEEE international conference on robotics and automation (ICRA)","author":"Wang","year":"2021"},{"key":"ref33","doi-asserted-by":"crossref","first-page":"1167","DOI":"10.1145\/3581783.3611808","article-title":"Cal-SFDA: source-free domain-adaptive semantic segmentation with differentiable expected calibration error","volume-title":"Proceedings of the 31st ACM International Conference on Multimedia","author":"Wang","year":"2023"},{"key":"ref34","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.inffus.2020.05.005","article-title":"Deep feature fusion through adaptive discriminative metric learning for scene recognition","volume":"63","author":"Wang","year":"2020","journal-title":"Information Fusion"},{"key":"ref35","article-title":"P2T: Pyramid pooling transformer for scene understanding","author":"Wu","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell. IEEE."},{"key":"ref36","doi-asserted-by":"publisher","first-page":"185","DOI":"10.1016\/j.neucom.2021.01.021","article-title":"Bi-directional skip connection feature pyramid network and sub-pixel convolution for high-quality object detection","volume":"440","author":"Xiong","year":"2021","journal-title":"Neurocomputing"},{"key":"ref37","doi-asserted-by":"crossref","first-page":"81","DOI":"10.1016\/j.neucom.2019.09.066","article-title":"MSN: modality separation networks for RGB-D scene recognition","volume":"373","author":"Xiong","year":"2020","journal-title":"Neurocomputing"},{"key":"ref38","doi-asserted-by":"publisher","first-page":"28746","DOI":"10.1109\/ACCESS.2020.2968771","article-title":"Remote sensing scene classification based on multi-structure deep features fusion","volume":"8","author":"Xue","year":"2020","journal-title":"IEEE Access"},{"key":"ref39","first-page":"1","article-title":"A small-sized object detection oriented multi-scale feature fusion approach with application to defect detection","volume":"71","author":"Zeng","year":"2022","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref40","doi-asserted-by":"crossref","first-page":"2324","DOI":"10.1109\/TIP.2008.2006658","article-title":"Multiresolution bilateral filtering for image denoising","volume":"17","author":"Zhang","year":"2008","journal-title":"IEEE Trans. Image Process."},{"key":"ref41","doi-asserted-by":"publisher","first-page":"6735","DOI":"10.1109\/TCSVT.2023.3289142","article-title":"Differential feature awareness network within antagonistic learning for infrared-visible object detection","volume":"34","author":"Zhang","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2024.1427786\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,28]],"date-time":"2024-11-28T12:09:54Z","timestamp":1732795794000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2024.1427786\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,19]]},"references-count":42,"alternative-id":["10.3389\/fnbot.2024.1427786"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2024.1427786","relation":{},"ISSN":["1662-5218"],"issn-type":[{"value":"1662-5218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,19]]},"article-number":"1427786"}}