{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T22:00:06Z","timestamp":1781647206677,"version":"3.54.5"},"reference-count":23,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T00:00:00Z","timestamp":1775779200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T00:00:00Z","timestamp":1775779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>This paper presents a novel multi-view multimodal graph learning framework for distributed audio-visual event classification using synchronized sequences from multi-microphone and multi-camera sensors. Existing approaches often rely on simple aggregation strategies for multi-view multimodal inputs, which fail to adequately capture the complex spatio-temporal relationships both within and across modalities. To address this limitation, we propose a graph attention network architecture with individual frame-level sensor nodes for each microphone and camera, and three types of frame-level spatio-temporal nodes. In this framework, within each temporal frame, audio spatio-temporal nodes connect to microphones, video spatio-temporal nodes to cameras, and audio-video spatio-temporal nodes to all sensor nodes. Temporal edges further interconnect each spatio-temporal node with its corresponding node in preceding frames. This architecture enables dynamic aggregation of sensor node features through attention-weighted mechanisms, generating updated spatio-temporal nodes that capture intra-modal, inter-modal, intra-frame, and inter-frame relational dependencies. Experimental results on the MM-Office and MM-OR datasets demonstrate that the proposed framework significantly outperforms existing baseline methods for audio-visual event classification, highlighting its superior capability in modeling complex spatio-temporal dependencies across distributed sensor networks.<\/jats:p>","DOI":"10.1007\/s11042-026-21512-2","type":"journal-article","created":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T17:51:59Z","timestamp":1775843519000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Hierarchical graph attention networks with spatio-temporal class tokens for distributed audio-visual event classification"],"prefix":"10.1007","volume":"85","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9553-0906","authenticated-orcid":false,"given":"Vijay","family":"John","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yasutomo","family":"Kawanishi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,4,10]]},"reference":[{"key":"21512_CR1","doi-asserted-by":"crossref","unstructured":"Aguilar-Ortega M, Moh\u00edno-Herranz I, Utrilla-Manso M et al (2019) Multi-microphone acoustic events detection and classification for indoor monitoring. In: Proceedings of the 2019 signal processing: algorithms, architectures, arrangements, and applications, pp 261\u2013266","DOI":"10.23919\/SPA.2019.8936807"},{"key":"21512_CR2","doi-asserted-by":"crossref","unstructured":"Ardianto S, Hang HM (2018) Multi-view and multi-modal action recognition with learned fusion. In: Proceedings of Asia-Pacific signal and information processing association annual summit and conference, pp 1601\u20131604","DOI":"10.23919\/APSIPA.2018.8659539"},{"key":"21512_CR3","doi-asserted-by":"crossref","unstructured":"Chang FJ, Radfar M, Mouchtaris A et al (2021) End-to-end multi-channel transformer for speech recognition. In: Proceedings of the 2021 IEEE international conference on acoustics, speech and signal processing, pp 5884\u20135888","DOI":"10.1109\/ICASSP39728.2021.9414123"},{"key":"21512_CR4","unstructured":"Dagdilelis D, Grigoriadis P, Galeazzi R (2025) Multimodal and multiview deep fusion for autonomous marine navigation. arXiv:2505.01615 [cs.CV]"},{"issue":"1","key":"21512_CR5","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1121\/10.0011809","volume":"152","author":"PA Grumiaux","year":"2022","unstructured":"Grumiaux PA, Kiti\u0107 S, Girin L et al (2022) A survey of sound source localization with deep learning methods. J Acoust Soc Am 152(1):107\u2013151","journal-title":"J Acoust Soc Am"},{"key":"21512_CR6","doi-asserted-by":"crossref","unstructured":"Gui N, Ge D, Hu Z (2019) Afs: an attention-based mechanism for supervised feature selection. In: Proceedings of the 33rd AAAI conference on artificial intelligence (455):3705\u20133713","DOI":"10.1609\/aaai.v33i01.33013705"},{"key":"21512_CR7","first-page":"9265","volume":"13","author":"S Hussain","year":"2025","unstructured":"Hussain S, Ali Teevno M, Naseem U et al (2025) Multiview multimodal feature fusion for breast cancer classification using deep learning 13:9265\u20139275","journal-title":"Multiview multimodal feature fusion for breast cancer classification using deep learning"},{"key":"21512_CR8","doi-asserted-by":"crossref","unstructured":"Ishiwatari T, Azuma M, Handa T et al (2022) Audio visual graph attention networks for event detection in sports video. In: Proceedings of the 30th European Signal Processing Conference (EUSIPCO), pp 553\u2013557","DOI":"10.23919\/EUSIPCO55093.2022.9909552"},{"key":"21512_CR9","unstructured":"Jiang Y, Tao R, Huang W et al (2024) Unified audio event detection. arXiv\/240908552"},{"key":"21512_CR10","doi-asserted-by":"crossref","unstructured":"John V, Kawanishi Y (2022) Audio and video-based emotion recognition using multimodal transformers. In: Proceedings of 26th International Conference on Pattern Recognition (ICPR), pp 2582\u20132588","DOI":"10.1109\/ICPR56361.2022.9956730"},{"key":"21512_CR11","doi-asserted-by":"crossref","unstructured":"John V, Kawanishi Y (2024) Generating pseudo-strong labels from weak labels for distributed multi-microphone sound event detection. In: Proceedings of the 27th international conference pattern recognition, pp 98\u2013113","DOI":"10.1007\/978-3-031-78192-6_7"},{"key":"21512_CR12","doi-asserted-by":"crossref","unstructured":"John V, Kawanishi Y (2025) Modelling spatio-temporal dynamics by graph attention network for distributed multi-microphone sound event classification. In: Proceedings of IEEE international conference on advanced visual and signal-based systems","DOI":"10.1109\/AVSS65446.2025.11149784"},{"key":"21512_CR13","unstructured":"Le TDN, Teh KK, Tran HD (2024) Continuous learning of transformer-based audio deepfake detection. arXiv\/240905924"},{"key":"21512_CR14","doi-asserted-by":"crossref","unstructured":"Li Y, Liu M, Drossos K et al (2020) Sound event detection via dilated convolutional recurrent neural networks. In: Proceedings of the IEEE international conference on acoustics, speech and signal processing, pp 286\u2013290","DOI":"10.1109\/ICASSP40776.2020.9054433"},{"key":"21512_CR15","doi-asserted-by":"crossref","unstructured":"Lin YB, Wang YCF (2020) Audiovisual transformer with instance attention for audio-visual event localization. in Proceedings of Asian Conference of Computer Vision pp 274\u2013290","DOI":"10.1007\/978-3-030-69544-6_17"},{"key":"21512_CR16","unstructured":"Nguyen TT, Kawanishi Y, John V et al (2025) Multitsf: Transformer-based sensor fusion for human-centric multi-view and multi-modal action recognition. In: Proceedings of the 2025 IEEE international conference on multimedia and signal processing"},{"key":"21512_CR17","doi-asserted-by":"crossref","unstructured":"Phan H, Le Nguyen H, Ch\u00e9n OY, et al. (2021) Multi-view audio and music classification. In: Proceedings of the 2021 IEEE international conference on acoustics, speech and signal processing, pp 611\u2013615","DOI":"10.1109\/ICASSP39728.2021.9414551"},{"key":"21512_CR18","doi-asserted-by":"crossref","unstructured":"Shah K, Shah A, Lau CP et al (2023) Multi-view action recognition using contrastive learning. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), pp 3381\u20133388","DOI":"10.1109\/WACV56688.2023.00338"},{"key":"21512_CR19","unstructured":"Xiong X, Arnab A, Nagrani A et al (2022) M-m mix: A multimodal multiview transformer ensemble. arXiv:2206.09852 [cs.CV]"},{"key":"21512_CR20","doi-asserted-by":"crossref","unstructured":"Yan S, Xiong X, Arnab A et al (2022) Multiview transformers for video recognition. arXiv:2201.04288 [cs.CV]","DOI":"10.1109\/CVPR52688.2022.00333"},{"key":"21512_CR21","doi-asserted-by":"crossref","unstructured":"Yasuda M, Ohishi Y, Saito S et al (2022) Multi-view and multi-modal event detection utilizing transformer-based multi-sensor fusion. In: Proceedings of the 2022 IEEE international conference on acoustics, speech and signal processing, pp 4638\u20134642","DOI":"10.1109\/ICASSP43922.2022.9746006"},{"issue":"9","key":"21512_CR22","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3687485","volume":"18","author":"W Ying","year":"2024","unstructured":"Ying W, Wang D, Chen H et al (2024) Feature selection as deep sequential generative learning. ACM Trans Knowl Discov Data 18(9):1\u201321","journal-title":"ACM Trans Knowl Discov Data"},{"key":"21512_CR23","unstructured":"Zhang D, Chen J, Bai J et al (2024) Sound event localization and classification using WASN in outdoor environment. arXiv\/240320130"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-026-21512-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-026-21512-2","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-026-21512-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T21:47:42Z","timestamp":1781646462000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-026-21512-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,10]]},"references-count":23,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["21512"],"URL":"https:\/\/doi.org\/10.1007\/s11042-026-21512-2","relation":{},"ISSN":["1573-7721"],"issn-type":[{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,10]]},"assertion":[{"value":"21 October 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 February 2026","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 March 2026","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 April 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing Interests"}}],"article-number":"343"}}