{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T04:42:00Z","timestamp":1777696920374,"version":"3.51.4"},"reference-count":62,"publisher":"SAGE Publications","issue":"2","license":[{"start":{"date-parts":[[2025,5,27]],"date-time":"2025-05-27T00:00:00Z","timestamp":1748304000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Intelligent Data Analysis: An International Journal"],"published-print":{"date-parts":[[2026,3]]},"abstract":"<jats:p>Due to the scarcity of annotated real-world data for specific categories, audio-visual generalized zero-shot learning (GZSL) has attracted significant attention. GZSL aims to classify novel classes absent during training while ensuring stable performance on seen classes. However, most existing methods operate implicitly, often neglecting the effective utilization of temporal, spatial, and semantic consistency. To address these challenges, we propose the Temporal Spatial Semantic Fusion network (TSSF). Specifically, we explore both audio and visual modalities using a multi-branch, multi-grained structure comprising a temporal global extraction module, a spatial local refinement module, and a multi-grained fusion module. The temporal global extraction module employs a Transformer-based Spiking Neural Network to extract explicit temporal representations and capture global dependencies. Simultaneously, the spatial local refinement module focuses on spatial information and local details using a window attention mechanism. Furthermore, the temporal and spatial features are hierarchically fused in the multi-grained fusion module, which incorporates both temporal and spatial attention for semantic enrichment. To explore multi-modal interactions, we enhance audio and visual features through cross-modal attention, followed by multi-modal alignment with text embeddings. Experiments on three benchmark audio-visual datasets validate the superiority of our method over state-of-the-art approaches. Notably, TSSF achieves significant improvements of 30.34% and 6.96% in HM and ZSL metrics on the VGG-GZSL dataset.<\/jats:p>","DOI":"10.1177\/1088467x251344931","type":"journal-article","created":{"date-parts":[[2025,5,28]],"date-time":"2025-05-28T02:53:41Z","timestamp":1748400821000},"page":"373-387","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["Temporal spatial semantic fusion network for audio-visual zero-shot learning"],"prefix":"10.1177","volume":"30","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-7721-691X","authenticated-orcid":false,"given":"Ming","family":"Guo","sequence":"first","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Feng","family":"Chen","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chongjun","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2025,5,27]]},"reference":[{"key":"e_1_3_4_2_2","doi-asserted-by":"publisher","DOI":"10.3233\/IDA-205183"},{"key":"e_1_3_4_3_2","doi-asserted-by":"crossref","unstructured":"Pian W Mo S Guo Y et al. Audio-visual class-incremental learning. In: Proceedings of the IEEE\/CVF international conference on computer vision (ICCV) 2023 pp.7799\u20137811.","DOI":"10.1109\/ICCV51070.2023.00717"},{"key":"e_1_3_4_4_2","doi-asserted-by":"crossref","unstructured":"Parida K Matiyali N Guha T et al. Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classification and retrieval of videos. In: Proceedings of the IEEE\/CVF winter conference on applications of computer vision 2020 pp.3251\u20133260.","DOI":"10.1109\/WACV45572.2020.9093438"},{"key":"e_1_3_4_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3430080"},{"key":"e_1_3_4_6_2","doi-asserted-by":"crossref","unstructured":"Wu P Liu J Shi Y et al. Not only look but also listen: learning multimodal violence detection under weak supervision. In: Proceedings of the European conference on computer vision (ECCV) Springer 2020 pp.322\u2013339.","DOI":"10.1007\/978-3-030-58577-8_20"},{"key":"e_1_3_4_7_2","doi-asserted-by":"crossref","unstructured":"Xu B Lu C Guo Y et al. Discriminative multi-modality speech recognition. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 2020 pp.14433\u201314442.","DOI":"10.1109\/CVPR42600.2020.01444"},{"key":"e_1_3_4_8_2","doi-asserted-by":"publisher","DOI":"10.3233\/IDA-230239"},{"key":"e_1_3_4_9_2","doi-asserted-by":"publisher","DOI":"10.3233\/IDA-230082"},{"key":"e_1_3_4_10_2","doi-asserted-by":"publisher","DOI":"10.3233\/IDA-205113"},{"key":"e_1_3_4_11_2","doi-asserted-by":"publisher","DOI":"10.3233\/IDA-215780"},{"key":"e_1_3_4_12_2","unstructured":"Wei Y Hu D Tian Y et al. Learning in audio-visual context: a review analysis and new perspective. arXiv preprint arXiv:2208.09579 2022."},{"key":"e_1_3_4_13_2","doi-asserted-by":"crossref","unstructured":"Mazumder P Singh P Parida KK et al. Avgzslnet: audio-visual generalized zero-shot learning by reconstructing label features from multi-modal embeddings. In: Proceedings of the IEEE\/CVF winter conference on applications of computer vision 2021 pp.3090\u20133099.","DOI":"10.1109\/WACV48630.2021.00313"},{"key":"e_1_3_4_14_2","doi-asserted-by":"crossref","unstructured":"Mercea O-B Riesch L Koepke A et al. Audio-visual generalised zero-shot learning with cross-modal attention and language. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 2022 pp.10553\u201310563.","DOI":"10.1109\/CVPR52688.2022.01030"},{"key":"e_1_3_4_15_2","doi-asserted-by":"crossref","unstructured":"Mercea O-B Hummel T Koepke AS et al. Temporal and cross-modal attention for audio-visual zero-shot learning. In: European conference on computer vision Springer 2022 pp.488\u2013505.","DOI":"10.1007\/978-3-031-20044-1_28"},{"key":"e_1_3_4_16_2","doi-asserted-by":"crossref","unstructured":"Li W Ma Z Deng L-J et al. Modality-fusion spiking transformer network for audio-visual zero-shot learning. In: 2023 IEEE international conference on multimedia and expo (ICME) 2023 pp.426\u2013431.","DOI":"10.1109\/ICME55011.2023.00080"},{"key":"e_1_3_4_17_2","doi-asserted-by":"crossref","unstructured":"Li W Zhao X-L Ma Z et al. Motion-decoupled spiking transformer for audio-visual zero-shot learning. In: Proceedings of the 31st ACM international conference on multimedia 2023 pp.3994\u20134002.","DOI":"10.1145\/3581783.3611759"},{"key":"e_1_3_4_18_2","doi-asserted-by":"crossref","unstructured":"Fei H Wu S Zhang M et al. Enhancing video-language representations with structural spatio-temporal alignment. In: IEEE transactions on pattern analysis and machine intelligence 2024.","DOI":"10.1109\/TPAMI.2024.3393452"},{"key":"e_1_3_4_19_2","doi-asserted-by":"crossref","unstructured":"Wu Y Yang Y. Exploring heterogeneous clues for weakly-supervised audio-visual video parsing. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 2021 pp.1326\u20131335.","DOI":"10.1109\/CVPR46437.2021.00138"},{"key":"e_1_3_4_20_2","doi-asserted-by":"crossref","unstructured":"Liang Y Zhu L Wang X et al. Icocap: improving video captioning by compounding images. In: IEEE transactions on multimedia 2023.","DOI":"10.1109\/TMM.2023.3322329"},{"key":"e_1_3_4_21_2","unstructured":"Huang P-Y Sharma V Xu H et al. Mavil: masked audio-video learners. In: Advances in neural information processing systems vol. 36 2024."},{"key":"e_1_3_4_22_2","doi-asserted-by":"crossref","unstructured":"Chen H Zhang H Wang L et al. Self-supervised audio-visual speaker representation with co-meta learning. In: Proceedings of the IEEE international conference on acoustics speech and signal processing (ICASSP) 2023 pp.1\u20135.","DOI":"10.1109\/ICASSP49357.2023.10096925"},{"key":"e_1_3_4_23_2","doi-asserted-by":"crossref","unstructured":"Liang Y Feng Q Zhu L et al. Seeg: semantic energized co-speech gesture generation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 2022 pp.10473\u201310482.","DOI":"10.1109\/CVPR52688.2022.01022"},{"key":"e_1_3_4_24_2","doi-asserted-by":"crossref","unstructured":"Wu Y Zhu L Yan Y et al. Dual attention matching for audio-visual event localization. In: Proceedings of the IEEE\/CVF international conference on computer vision 2019 pp.6292\u20136300.","DOI":"10.1109\/ICCV.2019.00639"},{"key":"e_1_3_4_25_2","doi-asserted-by":"crossref","unstructured":"Korbar B Tran D Torresani L. SCSampler: sampling salient clips from video for efficient action recognition. In: Proceedings of the IEEE\/CVF international conference on computer vision (ICCV) 2019.","DOI":"10.1109\/ICCV.2019.00633"},{"key":"e_1_3_4_26_2","doi-asserted-by":"crossref","unstructured":"Panda R Chen C-FR Fan Q et al. Adamml: adaptive multi-modal learning for efficient video recognition. In: Proceedings of the IEEE\/CVF international conference on computer vision 2021 pp.7576\u20137585.","DOI":"10.1109\/ICCV48922.2021.00748"},{"key":"e_1_3_4_27_2","doi-asserted-by":"crossref","unstructured":"Alfasly S Lu J Xu C et al. Learnable irrelevant modality dropout for multimodal action recognition on modality-specific annotated videos. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 2022 pp.20208\u201320217.","DOI":"10.1109\/CVPR52688.2022.01957"},{"key":"e_1_3_4_28_2","doi-asserted-by":"crossref","unstructured":"Kazakos E Nagrani A Zisserman A et al. EPIC-fusion: audio-visual temporal binding for egocentric action recognition. In: Proceedings of the IEEE\/CVF international conference on computer vision (ICCV) 2019.","DOI":"10.1109\/ICCV.2019.00559"},{"key":"e_1_3_4_29_2","doi-asserted-by":"crossref","unstructured":"Planamente M Plizzari C Alberti E et al. Domain generalization through audio-visual relative norm alignment in first person action recognition. In: Proceedings of the IEEE\/CVF winter conference on applications of computer vision 2022 pp.1807\u20131818.","DOI":"10.1109\/WACV51458.2022.00024"},{"key":"e_1_3_4_30_2","unstructured":"Romera-Paredes B Torr P. An embarrassingly simple approach to zero-shot learning. In: International conference on machine learning 2015 pp.2152\u20132161."},{"key":"e_1_3_4_31_2","doi-asserted-by":"crossref","unstructured":"Kodirov E Xiang T Gong S. Semantic autoencoder for zero-shot learning. In: Proceedings of the IEEE conference on computer vision and pattern recognition 2017 pp.3174\u20133183.","DOI":"10.1109\/CVPR.2017.473"},{"key":"e_1_3_4_32_2","doi-asserted-by":"crossref","unstructured":"Verma VK Arora G Mishra A et al. Generalized zero-shot learning via synthesized examples. In: Proceedings of the IEEE conference on computer vision and pattern recognition 2018 pp.4281\u20134289.","DOI":"10.1109\/CVPR.2018.00450"},{"key":"e_1_3_4_33_2","doi-asserted-by":"crossref","unstructured":"Roitberg A Martinez M Haurilet M et al. Towards a fair evaluation of zero-shot action recognition using external data. In: Proceedings of the European conference on computer vision (ECCV) workshops 2018.","DOI":"10.1007\/978-3-030-11018-5_8"},{"key":"e_1_3_4_34_2","doi-asserted-by":"crossref","unstructured":"Xian Y Lorenz T Schiele B et al. Feature generating networks for zero-shot learning. In: Proceedings of the IEEE conference on computer vision and pattern recognition 2018 pp.5542\u20135551.","DOI":"10.1109\/CVPR.2018.00581"},{"key":"e_1_3_4_35_2","doi-asserted-by":"crossref","unstructured":"Zhu Y Elhoseiny M Liu B et al. A generative adversarial approach for zero-shot learning from noisy texts. In: Proceedings of the IEEE conference on computer vision and pattern recognition 2018 pp.1004\u20131013.","DOI":"10.1109\/CVPR.2018.00111"},{"key":"e_1_3_4_36_2","doi-asserted-by":"crossref","unstructured":"Schonfeld E Ebrahimi S Sinha S et al. Generalized zero-and few-shot learning via aligned variational autoencoders. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 2019 pp.8247\u20138255.","DOI":"10.1109\/CVPR.2019.00844"},{"key":"e_1_3_4_37_2","doi-asserted-by":"crossref","unstructured":"Zhu Y Xie J Liu B et al. Learning feature-to-feature translator by alternating back-propagation for generative zero-shot learning. In: Proceedings of the IEEE\/CVF international conference on computer vision 2019 pp.9844\u20139854.","DOI":"10.1109\/ICCV.2019.00994"},{"key":"e_1_3_4_38_2","doi-asserted-by":"crossref","unstructured":"Brattoli B Tighe J Zhdanov F et al. Rethinking zero-shot video classification: End-to-end training for realistic applications. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 2020 pp.4613\u20134623.","DOI":"10.1109\/CVPR42600.2020.00467"},{"key":"e_1_3_4_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(97)00011-7"},{"key":"e_1_3_4_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-014-0788-3"},{"key":"e_1_3_4_41_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2017.00682"},{"key":"e_1_3_4_42_2","doi-asserted-by":"crossref","unstructured":"Wang Y Zhang M Chen Y et al. Signed neuron with memory: towards simple accurate and high-efficient ANN-SNN conversion. In: IJCAI 2022 pp.2501\u20132508.","DOI":"10.24963\/ijcai.2022\/347"},{"key":"e_1_3_4_43_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2020.00119"},{"key":"e_1_3_4_44_2","doi-asserted-by":"crossref","unstructured":"Li W Ma Z Deng L-J et al. Reservoir computing transformer for image-text retrieval. In: Proceedings of the ACM international conference on multimedia 2023 pp.5605\u20135613\u2013.","DOI":"10.1145\/3581783.3611758"},{"key":"e_1_3_4_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3233042"},{"key":"e_1_3_4_46_2","unstructured":"Li W Li J Ma M et al. Multi-scale spiking pyramid wireless communication framework for food recognition. In: IEEE transactions on multimedia 2024 pp.1\u201313."},{"key":"e_1_3_4_47_2","doi-asserted-by":"crossref","unstructured":"Li W Xiong R Fan X. Multi-layer probabilistic association reasoning network for image-text retrieval. In: IEEE transactions on circuits and systems for video technology 2024 pp.1\u201315.","DOI":"10.1109\/TCSVT.2024.3394551"},{"key":"e_1_3_4_48_2","unstructured":"Shen S Zhao D Shen G et al. TIM: an efficient temporal interaction module for spiking transformer. In: Proceedings of the thirty-third international joint conference on artificial intelligence IJCAI 2024 pp.3133\u20133141."},{"key":"e_1_3_4_49_2","unstructured":"Yao M Hu J Zhou Z et al. Spike-driven transformer. In: Advances in neural information processing systems vol. 36 2024."},{"key":"e_1_3_4_50_2","doi-asserted-by":"crossref","unstructured":"Xin Y Du J Wang Q et al. Mmap: multi-modal alignment prompt for cross-domain multi-task learning. In: Proceedings of the AAAI conference on artificial intelligence vol. 38 2024 pp.16076\u201316084.","DOI":"10.1609\/aaai.v38i14.29540"},{"key":"e_1_3_4_51_2","doi-asserted-by":"crossref","unstructured":"Xin Y Du J Wang Q et al. Vmt-adapter: parameter-efficient transfer learning for multi-task dense scene understanding. In: Proceedings of the AAAI conference on artificial intelligence vol. 38 2024 pp.16085\u201316093.","DOI":"10.1609\/aaai.v38i14.29541"},{"key":"e_1_3_4_52_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106677"},{"key":"e_1_3_4_53_2","unstructured":"Zhou Z Zhu Y He C et al. Spikformer: when spiking neural network meets transformer. In: The eleventh international conference on learning representations ICLR 2023."},{"key":"e_1_3_4_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-022-12096-8"},{"key":"e_1_3_4_55_2","doi-asserted-by":"crossref","unstructured":"Luo T Wu J He Z et al. WFormer: a transformer-based soft fusion model for robust image watermarking. In: IEEE transactions on emerging topics in computational intelligence 2024.","DOI":"10.1109\/TETCI.2024.3386916"},{"key":"e_1_3_4_56_2","doi-asserted-by":"crossref","unstructured":"Liu Z Lin Y Cao Y et al. Swin transformer: hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE\/CVF international conference on computer vision 2021 pp.10012\u201310022.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_4_57_2","unstructured":"Xin Y Luo S Zhou H et al. Parameter-efficient fine-tuning for pre-trained vision models: a survey. arXiv preprint arXiv:2402.02242 2024."},{"key":"e_1_3_4_58_2","doi-asserted-by":"crossref","unstructured":"Chen H Xie W Vedaldi A et al. Vggsound: a large-scale audio-visual dataset. In: ICASSP 2020-2020 IEEE international conference on acoustics speech and signal processing (ICASSP) IEEE 2020 pp.721\u2013725.","DOI":"10.1109\/ICASSP40776.2020.9053174"},{"key":"e_1_3_4_59_2","unstructured":"Soomro K. UCF101: a dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 2012."},{"key":"e_1_3_4_60_2","doi-asserted-by":"crossref","unstructured":"Caba Heilbron F Escorcia V Ghanem B et al. Activitynet: a large-scale video benchmark for human activity understanding. In: Proceedings of the IEEE conference on computer vision and pattern recognition 2015 pp.961\u2013970.","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_4_61_2","first-page":"4660","article-title":"Labelling unlabelled videos from scratch with multi-modal self-supervision","volume":"33","author":"Asano Y","year":"2020","unstructured":"Asano Y, Patrick M, Rupprecht C, et al. Labelling unlabelled videos from scratch with multi-modal self-supervision. Adv Neural Inf Process Syst 2020; 33: 4660\u20134671.","journal-title":"Adv Neural Inf Process Syst"},{"key":"e_1_3_4_62_2","first-page":"21969","article-title":"Attribute prototype network for zero-shot learning","volume":"33","author":"Xu W","year":"2020","unstructured":"Xu W, Xian Y, Wang J, et al. Attribute prototype network for zero-shot learning. Adv Neural Inf Process Syst 2020; 33: 21969\u201321980.","journal-title":"Adv Neural Inf Process Syst"},{"key":"e_1_3_4_63_2","doi-asserted-by":"crossref","unstructured":"Xian Y Sharma S Schiele B et al. f-vaegan-d2: a feature generating framework for any-shot learning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 2019 pp.10275\u201310284.","DOI":"10.1109\/CVPR.2019.01052"}],"container-title":["Intelligent Data Analysis: An International Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1088467X251344931","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1088467X251344931","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1088467X251344931","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:21:19Z","timestamp":1777454479000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1088467X251344931"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,27]]},"references-count":62,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,3]]}},"alternative-id":["10.1177\/1088467X251344931"],"URL":"https:\/\/doi.org\/10.1177\/1088467x251344931","relation":{},"ISSN":["1088-467X","1571-4128"],"issn-type":[{"value":"1088-467X","type":"print"},{"value":"1571-4128","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,27]]}}}