{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,27]],"date-time":"2026-02-27T09:00:16Z","timestamp":1772182816014,"version":"3.50.1"},"reference-count":46,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,2,19]],"date-time":"2025-02-19T00:00:00Z","timestamp":1739923200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["FRF-TP-24-021A"],"award-info":[{"award-number":["FRF-TP-24-021A"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["No.FRF-MP-19-007,"],"award-info":[{"award-number":["No.FRF-MP-19-007,"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["No.FRF-TP-20-065A1Z"],"award-info":[{"award-number":["No.FRF-TP-20-065A1Z"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No.U2133218"],"award-info":[{"award-number":["No.U2133218"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neuroinform."],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>Recently, numerous studies have focused on the semantic decoding of perceived images based on functional magnetic resonance imaging (fMRI) activities. However, it remains unclear whether it is possible to establish relationships between brain activities and semantic features of human actions in video stimuli. Here we construct a framework for decoding action semantics by establishing relationships between brain activities and semantic features of human actions.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>To effectively use a small amount of available brain activity data, our proposed method employs a pre-trained image action recognition network model based on an expanding three-dimensional (X3D) deep neural network framework (DNN). To apply brain activities to the image action recognition network, we train regression models that learn the relationship between brain activities and deep-layer image features. To improve decoding accuracy, we join by adding the nonlocal-attention mechanism module to the X3D model to capture long-range temporal and spatial dependence, proposing a multilayer perceptron (MLP) module of multi-task loss constraint to build a more accurate regression mapping approach and performing data enhancement through linear interpolation to expand the amount of data to reduce the impact of a small sample.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results and discussion<\/jats:title><jats:p>Our findings indicate that the features in the X3D-DNN are biologically relevant, and capture information useful for perception. The proposed method enriches the semantic decoding model. We have also conducted several experiments with data from different subsets of brain regions known to process visual stimuli. The results suggest that semantic information for human actions is widespread across the entire visual cortex.<\/jats:p><\/jats:sec>","DOI":"10.3389\/fninf.2025.1526259","type":"journal-article","created":{"date-parts":[[2025,2,19]],"date-time":"2025-02-19T06:51:18Z","timestamp":1739947878000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["An action decoding framework combined with deep neural network for predicting the semantics of human actions in videos from evoked brain activities"],"prefix":"10.3389","volume":"19","author":[{"given":"Yuanyuan","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Manli","family":"Tian","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Baolin","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2025,2,19]]},"reference":[{"key":"B1","article-title":"Brain2Word: Improving brain decoding methods and evaluation","author":"Affolter","year":"2020","journal-title":"Proceedings of the Medical Imaging Meets Neurips Workshop-34th Conference on Neural Information Processing Systems"},{"key":"B2","doi-asserted-by":"crossref","first-page":"202","DOI":"10.1109\/GCCE.2018.8574847","article-title":"Estimation of viewed image categories via CCA using human brain activity","author":"Akamatsu","year":"2018","journal-title":"Proceedings of the 2018 IEEE 7th Global Conference on Consumer Electronics (GCCE)"},{"key":"B3","first-page":"1215","article-title":"Multi-view bayesian generative model for multi-subject fmri data on brain decoding of viewed image categories","author":"Akamatsu","year":"2020","journal-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)"},{"key":"B4","first-page":"6836","article-title":"Vivit: A video vision transformer","author":"Arnab","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision"},{"key":"B5","article-title":"Quo vadis, action recognition? A new model and the kinetics dataset","author":"Carreira","year":"2017","journal-title":"Proceedings of the Inernational Conference on Computer Vision and Pattern Recognition"},{"key":"B6","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/srep27755","article-title":"Comparison of deep neural networks to spatio-temporal cortical dynamics of human visual object recognition reveals hierarchical correspondence.","volume":"6","author":"Cichy","year":"2016","journal-title":"Sci. Rep."},{"key":"B7","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1007\/BF00994018","article-title":"Support-vector networks.","volume":"20","author":"Cortes","year":"1995","journal-title":"Mach. Learn."},{"key":"B8","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.11929","article-title":"An image is worth 16x16 words: Transformers for image recognition at scale.","author":"Dosovitskiy","year":"2021","journal-title":"arXiv [Preprint]"},{"key":"B9","doi-asserted-by":"publisher","first-page":"184","DOI":"10.1016\/j.neuroimage.2016.10.001","article-title":"Seeing it all: Convolutional network layers map the function of the human visual system.","volume":"152","author":"Eickenberg","year":"2016","journal-title":"NeuroImage"},{"key":"B10","doi-asserted-by":"publisher","first-page":"181","DOI":"10.1093\/cercor\/7.2.181","article-title":"Retinotopic organization in human visual cortex and the spatial precision of functional MRI.","volume":"7","author":"Engel","year":"1997","journal-title":"Cereb. Cortex"},{"key":"B11","first-page":"203","article-title":"X3D: Expanding architectures for efficient video recognition","author":"Feichtenhofer","year":"2020","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"B12","first-page":"179","article-title":"The use of multiple measurements in taxonomic problems.","volume":"7","author":"Fisher","year":"2012","journal-title":"Ann. Hum. Genet."},{"key":"B13","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1016\/j.neuron.2014.10.047","article-title":"Prediction as a humanitarian and pragmatic contribution from human cognitive neuroscience.","volume":"85","author":"Gabrieli","year":"2015","journal-title":"Neuron"},{"key":"B14","doi-asserted-by":"publisher","first-page":"10005","DOI":"10.1523\/JNEUROSCI.5023-14.2015","article-title":"Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream.","volume":"35","author":"G\u00fc\u00e7l\u00fc","year":"2015","journal-title":"J. Neurosci."},{"key":"B15","doi-asserted-by":"publisher","first-page":"329","DOI":"10.1016\/j.neuroimage.2015.12.036","article-title":"Increasingly complex representations of natural movies across the dorsal stream are shared between subjects.","volume":"145","author":"G\u00fc\u00e7l\u00fc","year":"2017","journal-title":"NeuroImage"},{"key":"B16","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory.","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"B17","doi-asserted-by":"publisher","DOI":"10.1038\/ncomms15037","article-title":"Generic decoding of seen and imagined objects using hierarchical visual features.","volume":"8","author":"Horikawa","year":"","journal-title":"Nat. Commun."},{"key":"B18","doi-asserted-by":"publisher","DOI":"10.3389\/fncom.2017.00004","article-title":"Hierarchical neural representation of dreamed objects revealed by brain decoding with deep neural network features.","volume":"11","author":"Horikawa","year":"","journal-title":"Front. Comput. Neurosci."},{"key":"B19","doi-asserted-by":"publisher","DOI":"10.3389\/fnsys.2016.00081","article-title":"Decoding the semantic content of natural movies from human brain activity.","volume":"10","author":"Huth","year":"2016","journal-title":"Front. Syst. Neurosci."},{"key":"B20","doi-asserted-by":"publisher","first-page":"1210","DOI":"10.1016\/j.neuron.2012.10.014","article-title":"A continuous semantic space describes the representation of thousands of object and action categories across the human brain.","volume":"76","author":"Huth","year":"2012","journal-title":"Neuron"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1705.06950","article-title":"The kinetics human action video dataset.","author":"Kay","year":"2017","journal-title":"arXiv [Preprint]"},{"key":"B22","doi-asserted-by":"publisher","first-page":"2556","DOI":"10.1109\/TIP.2019.2952088","article-title":"HMDB: A large video database for human motion recognition","author":"Kuehne","year":"2011","journal-title":"Proceedings of the IEEE International Conference Computer Vision"},{"key":"B23","doi-asserted-by":"crossref","DOI":"10.3390\/math11143081","article-title":"Transformer models and convolutional networks with different activation functions for swallow classification using depth video data.","volume":"11","author":"Lai","year":"2023","journal-title":"Mathematics"},{"key":"B24","doi-asserted-by":"crossref","first-page":"1025","DOI":"10.1016\/j.ins.2020.09.012","article-title":"Multi-subject data augmentation for target subject semantic decoding with deep multi-view adversarial learning.","volume":"547","author":"Li","year":"2021","journal-title":"Inf. Sci."},{"key":"B25","doi-asserted-by":"publisher","first-page":"679","DOI":"10.1109\/TPAMI.2024.3429387","article-title":"Human-centric transformer for domain adaptive action recognition.","volume":"47","author":"Lin","year":"2025","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"B26","doi-asserted-by":"crossref","first-page":"576","DOI":"10.1109\/SMC.2018.00107","article-title":"Describing semantic representations of brain activity evoked by visual stimuli","author":"Matsuo","year":"2018","journal-title":"Proceedings of the 2018 IEEE International Conference on Systems, Man, and Cybernetics (SMC)"},{"key":"B27","first-page":"64","article-title":"Recurrent neural networks.","volume":"5","author":"Medsker","year":"2001","journal-title":"Design Appl."},{"key":"B28","doi-asserted-by":"publisher","first-page":"14551","DOI":"10.1523\/JNEUROSCI.6801-10.2011","article-title":"A three-dimensional spatiotemporal receptive field model explains responses of area MT neurons to naturalistic movies.","volume":"31","author":"Nishimoto","year":"2011","journal-title":"J. Neurosci."},{"key":"B29","doi-asserted-by":"publisher","first-page":"1641","DOI":"10.1016\/j.cub.2011.08.031","article-title":"Reconstructing visual experiences from brain activity evoked by natural movies.","volume":"21","author":"Nishimoto","year":"2011","journal-title":"Curr. Biol."},{"key":"B30","doi-asserted-by":"publisher","first-page":"1720","DOI":"10.1109\/TNSRE.2020.3006180","article-title":"Modeling EEG data distribution with a wasserstein generative adversarial network to predict rsvp events.","volume":"28","author":"Panwar","year":"2020","journal-title":"IEEE Trans. Neural Syst. Rehabil. Eng."},{"key":"B31","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1016\/j.patrec.2019.08.007","article-title":"Visual representation decoding from human brain activity using machine learning: A baseline study.","volume":"128","author":"Papadimitriou","year":"2019","journal-title":"Patt. Recognit. Lett."},{"key":"B32","doi-asserted-by":"publisher","first-page":"0460b6","DOI":"10.1088\/1741-2552\/ac1179","article-title":"Modeling and augmenting of fMRI data using deep recurrent variational auto-encoder.","volume":"18","author":"Qiang","year":"2021","journal-title":"J. Neural Eng"},{"key":"B33","doi-asserted-by":"publisher","DOI":"10.3389\/fninf.2018.00062","article-title":"Accurate reconstruction of image stimuli from human functional magnetic resonance imaging based on the decoding model with capsule network architecture.","volume":"12","author":"Qiao","year":"2018","journal-title":"Front. Neuroinformatics"},{"key":"B34","doi-asserted-by":"crossref","first-page":"945","DOI":"10.1016\/j.neuron.2005.05.021","article-title":"Spatiotemporal elements of macaque v1 receptive fields.","volume":"46","author":"Rust","year":"2005","journal-title":"Neuron"},{"key":"B35","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1212.0402","article-title":"UCF101: A dataset of 101 human actions classes from videos in the wild.","author":"Soomro","year":"2012","journal-title":"arXiv [Preprint]"},{"key":"B36","doi-asserted-by":"publisher","first-page":"1025","DOI":"10.1016\/j.neuron.2013.06.034","article-title":"Natural scene statistics account for the representation of scene categories in human visual cortex.","volume":"79","author":"Stansbury","year":"2013","journal-title":"Neuron"},{"key":"B37","doi-asserted-by":"crossref","first-page":"2521","DOI":"10.1109\/ICIP40778.2020.9191262","article-title":"Generation of viewed image captions from human brain activity via unsupervised text latent space","author":"Takada","year":"2020","journal-title":"Proceedings of the 2020 IEEE International Conference on Image Processing (ICIP)"},{"key":"B38","doi-asserted-by":"publisher","DOI":"10.1016\/j.neuroimage.2019.116350","article-title":"Reliability-based voxel selection.","volume":"207","author":"Tarhan","year":"2019","journal-title":"NeuroImage"},{"key":"B39","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-020-16846-w","article-title":"Sociality and interaction envelope organize visual action representations.","volume":"11","author":"Tarhan","year":"2020","journal-title":"Nat. Commun."},{"key":"B40","article-title":"Learning spatiotemporal features with 3D convolutional networks","author":"Tran","year":"2015","journal-title":"Proceedings of the IEEE International Conference on Computer Vision"},{"key":"B41","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1016\/j.neuropsychologia.2019.02.006","article-title":"Distinct representations in occipito-temporal, parietal, and premotor cortex during action perception revealed by fMRI and computational modeling.","volume":"127","author":"Urgen","year":"2019","journal-title":"Neuropsychologia"},{"key":"B42","doi-asserted-by":"publisher","first-page":"223","DOI":"10.1016\/j.neuroimage.2017.06.042","article-title":"Mapping between fMRI responses to movies and their natural language annotations.","volume":"180","author":"Vodrahalli","year":"2018","journal-title":"NeuroImage"},{"key":"B43","doi-asserted-by":"publisher","first-page":"7794","DOI":"10.3390\/bioengineering11060627","article-title":"Non-local neural networks","author":"Wang","year":"2018","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"B44","doi-asserted-by":"crossref","first-page":"4136","DOI":"10.1093\/cercor\/bhx268","article-title":"Neural encoding and decoding with deep learning for dynamic natural vision.","volume":"28","author":"Wen","year":"2018","journal-title":"Cereb. Cortex"},{"key":"B45","doi-asserted-by":"publisher","first-page":"8619","DOI":"10.1073\/pnas.1403112111","article-title":"Performance-optimized hierarchical models predict neural responses in higher visual cortex.","volume":"111","author":"Yamins","year":"2014","journal-title":"Proc. Natl. Acad. Sci."},{"key":"B46","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1710.09412","article-title":"mixup: Beyond empirical risk minimization.","author":"Zhang","year":"2018","journal-title":"arXiv [Preprint]"}],"container-title":["Frontiers in Neuroinformatics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fninf.2025.1526259\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,19]],"date-time":"2025-02-19T06:51:29Z","timestamp":1739947889000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fninf.2025.1526259\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,19]]},"references-count":46,"alternative-id":["10.3389\/fninf.2025.1526259"],"URL":"https:\/\/doi.org\/10.3389\/fninf.2025.1526259","relation":{},"ISSN":["1662-5196"],"issn-type":[{"value":"1662-5196","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,19]]},"article-number":"1526259"}}