{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T15:03:53Z","timestamp":1776956633086,"version":"3.51.4"},"reference-count":39,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,3,26]],"date-time":"2025-03-26T00:00:00Z","timestamp":1742947200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Comput. Neurosci."],"abstract":"<jats:p>In recent years, the integration of machine vision and neuroscience has provided a new perspective for deeply understanding visual information. This paper proposes an innovative deep learning model, NeuroFusionNet, designed to enhance the understanding of visual information by integrating fMRI signals with image features. Specifically, images are processed by a visual model to extract region-of-interest (ROI) features and contextual information, which are then encoded through fully connected layers. The fMRI signals are passed through 1D convolutional layers to extract features, effectively preserving spatial information and improving computational efficiency. Subsequently, the fMRI features are embedded into a 3D voxel representation to capture the brain's activity patterns in both spatial and temporal dimensions. To accurately model the brain's response to visual stimuli, this paper introduces a Mutli-scale fMRI Timeformer module, which processes fMRI signals at different scales to extract both fine details and global responses. To further optimize the model's performance, we introduce a novel loss function called the fMRI-guided loss. Experimental results show that NeuroFusionNet effectively integrates image and brain activity information, providing more precise and richer visual representations for machine vision systems, with broad potential applications.<\/jats:p>","DOI":"10.3389\/fncom.2025.1545971","type":"journal-article","created":{"date-parts":[[2025,3,26]],"date-time":"2025-03-26T06:54:17Z","timestamp":1742972057000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["NeuroFusionNet: cross-modal modeling from brain activity to visual understanding"],"prefix":"10.3389","volume":"19","author":[{"given":"Kehan","family":"Lang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianwei","family":"Fang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guangyao","family":"Su","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2025,3,26]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"116","DOI":"10.1038\/s41593-021-00962-x","article-title":"A massive 7t fMRI dataset to bridge cognitive neuroscience and artificial intelligence","volume":"25","author":"Allen","year":"2022","journal-title":"Nat. Neurosci"},{"key":"B2","doi-asserted-by":"crossref","first-page":"2197","DOI":"10.1109\/BIBM.2018.8621112","article-title":"\u201cUtilizing mask r-CNN for detection and segmentation of oral diseases,\u201d","volume-title":"2018 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)","author":"Anantharaman","year":"2018"},{"key":"B3","doi-asserted-by":"publisher","first-page":"110117","DOI":"10.1016\/j.compeleceng.2025.110117","article-title":"AI-enabled computational intelligence approach to neurodevelopmental disorders detection using rs-fMRI data","volume":"123","author":"Bandyopadhyay","year":"2025","journal-title":"Comput. Electr. Eng"},{"key":"B4","article-title":"VLMO: unified vision-language pre-training with mixture-of-modality-experts","author":"Bao","year":"2021","journal-title":"arXiv preprint arXiv:2111.02358"},{"key":"B5","article-title":"Mmdetection: open MMLAB detection toolbox and benchmark","author":"Chen","year":"2019","journal-title":"arXiv preprint arXiv:1906.07155"},{"key":"B6","first-page":"1597","article-title":"\u201cA simple framework for contrastive learning of visual representations,\u201d","volume-title":"International Conference on Machine Learning","author":"Chen","year":""},{"key":"B7","first-page":"104","article-title":"\u201cUniter: universal image-text representation learning,\u201d","volume-title":"European Conference on Computer Vision","author":"Chen","year":""},{"key":"B8","doi-asserted-by":"publisher","first-page":"22710","DOI":"10.1109\/CVPR52729.2023.02175","article-title":"\u201cSeeing beyond the brain: conditional diffusion model with sparse masked modeling for vision decoding,\u201d","author":"Chen","year":"2023","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B9","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1016\/S1053-8119(03)00049-1","article-title":"Functional magnetic resonance imaging (fMRI) \u201cbrain reading\u201d: detecting and classifying distributed patterns of fMRI activity in human visual cortex","volume":"19","author":"Cox","year":"2003","journal-title":"Neuroimage"},{"key":"B10","first-page":"36693","article-title":"\u201cFuzzy learning machine,\u201d","author":"Cui","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B11","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1109\/CVPR.2009.5206848","article-title":"\u201cImagenet: a large-scale hierarchical image database,\u201d","volume-title":"2009 IEEE Conference on Computer Vision and Pattern Recognition","author":"Deng","year":"2009"},{"key":"B12","article-title":"Bert: pre-training of deep bidirectional transformers for language understanding","author":"Devlin","year":"2018","journal-title":"arXiv preprint arXiv:1810.04805"},{"key":"B13","article-title":"An image is worth 16x16 words: Transformers for image recognition at scale","author":"Dosovitskiy","year":"2020","journal-title":"arXiv preprint arXiv:2010.11929"},{"key":"B14","article-title":"\u201cFew-shot generation via recalling brain-inspired episodic-semantic memory,\u201d","author":"Duan","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B15","doi-asserted-by":"publisher","first-page":"e1009267","DOI":"10.1371\/journal.pcbi.1009267","article-title":"Unveiling functions of the visual cortex using task-specific deep neural networks","volume":"17","author":"Dwivedi","year":"2021","journal-title":"PLoS Comput. Biol"},{"key":"B16","doi-asserted-by":"publisher","first-page":"5397","DOI":"10.1038\/s41598-018-23618-6","article-title":"Using human brain activity to guide machine learning","volume":"8","author":"Fong","year":"2018","journal-title":"Sci. Rep"},{"key":"B17","article-title":"The algonauts project 2023 challenge: How the human brain makes sense of natural scenes","author":"Gifford","year":"2023","journal-title":"arXiv preprint arXiv:2301.03198"},{"key":"B18","doi-asserted-by":"publisher","first-page":"471","DOI":"10.1038\/nature20101","article-title":"Hybrid computing using a neural network with dynamic external memory","volume":"538","author":"Graves","year":"2016","journal-title":"Nature"},{"key":"B19","first-page":"1462","article-title":"\u201cDraw: a recurrent neural network for image generation,\u201d","volume-title":"International Conference on Machine Learning","author":"Gregor","year":"2015"},{"key":"B20","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/CIVEMSA58715.2024.10586647","article-title":"\u201cMINDLDM: reconstruct visual stimuli from fMRI using latent diffusion model,\u201d","volume-title":"2024 IEEE International Conference on Computational Intelligence and Virtual Environments for Measurement Systems and Applications (CIVEMSA)","author":"Guo","year":"2024"},{"key":"B21","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s43586-021-00018-1","article-title":"Clip and complementary methods","volume":"1","author":"Hafner","year":"2021","journal-title":"Nat. Rev. Methods Prim"},{"key":"B22","doi-asserted-by":"publisher","first-page":"102551","DOI":"10.1016\/j.inffus.2024.102551","article-title":"Coarse to fine-based image-point cloud fusion network for 3D object detection","volume":"112","author":"Hao","year":"2024","journal-title":"Inf. Fusion"},{"key":"B23","doi-asserted-by":"publisher","first-page":"2425","DOI":"10.1126\/science.1063736","article-title":"Distributed and overlapping representations of faces and objects in ventral temporal cortex","volume":"293","author":"Haxby","year":"2001","journal-title":"Science"},{"key":"B24","doi-asserted-by":"publisher","first-page":"686","DOI":"10.1038\/nn1445","article-title":"Predicting the orientation of invisible stimuli from activity in human primary visual cortex","volume":"8","author":"Haynes","year":"2005","journal-title":"Nat. Neurosci"},{"key":"B25","doi-asserted-by":"publisher","first-page":"16000","DOI":"10.1109\/CVPR52688.2022.01553","article-title":"\u201cMasked autoencoders are scalable vision learners,\u201d","author":"He","year":"2022","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B26","doi-asserted-by":"publisher","first-page":"770","DOI":"10.1109\/CVPR.2016.90","article-title":"\u201cDeep residual learning for image recognition,\u201d","author":"He","year":"2016","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"B27","doi-asserted-by":"publisher","first-page":"113175","DOI":"10.1016\/j.knosys.2025.113175","article-title":"Multiscale spectral augmentation for graph contrastive learning for fMRI analysis to diagnose psychiatric disease","volume":"314","author":"Hu","year":"2025","journal-title":"Knowl.-Based Syst"},{"key":"B28","article-title":"Brain-streams: fMRI-to-image reconstruction with multi-modal guidance","author":"Joo","year":"2024","journal-title":"arXiv preprint arXiv:2409.12099"},{"key":"B29","doi-asserted-by":"publisher","first-page":"1","DOI":"10.4018\/IJEHMC.321149","article-title":"Computational framework of inverted fuzzy c-means and quantum convolutional neural network towards accurate detection of ovarian tumors","volume":"14","author":"Kodipalli","year":"2023","journal-title":"Int. J. E-Health Med. Commun"},{"key":"B30","article-title":"Language models are few-shot learners","author":"Mann","year":"2020","journal-title":"arXiv preprint arXiv:2005.14165.14161"},{"key":"B31","doi-asserted-by":"publisher","first-page":"109216","DOI":"10.1016\/j.patcog.2022.109216","article-title":"Hyper-sausage coverage function neuron model and learning algorithm for image classification","volume":"136","author":"Ning","year":"2023","journal-title":"Pattern Recognit"},{"key":"B32","doi-asserted-by":"publisher","first-page":"108873","DOI":"10.1016\/j.patcog.2022.108873","article-title":"Hcfnn: high-order coverage function neural network for image classification","volume":"131","author":"Ning","year":"2022","journal-title":"Pattern Recognit"},{"key":"B33","doi-asserted-by":"publisher","first-page":"102033","DOI":"10.1016\/j.inffus.2023.102033","article-title":"Dilf: Differentiable rendering-based multi-view image-language fusion for zero-shot 3d shape understanding","volume":"102","author":"Ning","year":"2024","journal-title":"Inf. Fusion"},{"key":"B34","doi-asserted-by":"publisher","first-page":"1","DOI":"10.4018\/IJEHMC.321150","article-title":"Detection of antibiotic constituent in aspergillus flavus using quantum convolutional neural network","volume":"14","author":"Sannidhan","year":"2023","journal-title":"Int. J. E-Health Med. Commun"},{"key":"B35","first-page":"36","article-title":"\u201cReconstructing the mind's eye: fMRI-to-image with contrastive learning and diffusion priors,\u201d","author":"Scotti","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B36","article-title":"Using graph convolutional networks to address fMRI small data problems","author":"Screven","year":"2025","journal-title":"arXiv preprint arXiv:2502.17489"},{"key":"B37","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2025.3525725","article-title":"\u201cA unified framework for adversarial patch attacks against visual 3d object detection in autonomous driving,\u201d","author":"Wang","year":"2025","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"B38","doi-asserted-by":"publisher","first-page":"8226","DOI":"10.1109\/WACV57701.2024.00804","article-title":"\u201cDream: visual decoding from reversing human visual system,\u201d","author":"Xia","year":"2024","journal-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision"},{"key":"B39","doi-asserted-by":"publisher","first-page":"107471","DOI":"10.1016\/j.bspc.2024.107471","article-title":"Fuse-former: An interpretability analysis model for rs-fMRI based on multi-scale information fusion interaction","volume":"105","author":"Ye","year":"2025","journal-title":"Biomed. Signal Process. Control"}],"container-title":["Frontiers in Computational Neuroscience"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fncom.2025.1545971\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,3,26]],"date-time":"2025-03-26T06:55:08Z","timestamp":1742972108000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fncom.2025.1545971\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,26]]},"references-count":39,"alternative-id":["10.3389\/fncom.2025.1545971"],"URL":"https:\/\/doi.org\/10.3389\/fncom.2025.1545971","relation":{},"ISSN":["1662-5188"],"issn-type":[{"value":"1662-5188","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,26]]},"article-number":"1545971"}}