{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T20:20:28Z","timestamp":1776889228265,"version":"3.51.2"},"reference-count":0,"publisher":"IOS Press","isbn-type":[{"value":"9781643686318","type":"electronic"}],"license":[{"start":{"date-parts":[[2025,10,21]],"date-time":"2025-10-21T00:00:00Z","timestamp":1761004800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,10,21]]},"abstract":"<jats:p>Audio-visual speech separation (AVSS) aims to extract a target speech signal from a mixed signal by leveraging both auditory and visual (lip movement) cues. However, most existing AVSS methods exhibit complex architectures and rely on future context, operating offline, which renders them unsuitable for real-time applications. Inspired by the pipeline of RTFSNet, we propose a novel streaming AVSS model, named Swift-Net, which enhances the causal processing capabilities required for real-time applications. Swift-Net adopts a lightweight visual feature extraction module and an efficient fusion module for audio-visual integration. Additionally, Swift-Net employs Grouped SRUs to integrate historical information across different feature spaces, thereby improving the utilization efficiency of historical information. We further propose a causal transformation template to facilitate the conversion of non-causal AVSS models into causal counterparts. Experiments on three standard benchmark datasets (LRS2, LRS3, and VoxCeleb2) demonstrated that under causal conditions, our proposed Swift-Net exhibited outstanding performance, highlighting the potential of this method for processing speech in complex environments.<\/jats:p>","DOI":"10.3233\/faia250845","type":"book-chapter","created":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T09:43:55Z","timestamp":1761126235000},"source":"Crossref","is-referenced-by-count":1,"title":["A Fast and Lightweight Model for Causal Audio-Visual Speech Separation"],"prefix":"10.3233","author":[{"given":"Wendi","family":"Sang","sequence":"first","affiliation":[{"name":"School of Computer Technology and Application, Qinghai University, Xining 810016, China"},{"name":"Intelligent Computing and Application Laboratory of Qinghai Province, Qinghai University, Xining 810016, China"},{"name":"Department of Computer Science and Technology, Institute for AI, BNRist, Tsinghua Laboratory of Brain and Intelligence (THBI), IDG\/McGovern Institute for Brain Research, Tsinghua University, Beijing 100084, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology, Institute for AI, BNRist, Tsinghua Laboratory of Brain and Intelligence (THBI), IDG\/McGovern Institute for Brain Research, Tsinghua University, Beijing 100084, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Runxuan","family":"Yang","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology, Institute for AI, BNRist, Tsinghua Laboratory of Brain and Intelligence (THBI), IDG\/McGovern Institute for Brain Research, Tsinghua University, Beijing 100084, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianqiang","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer Technology and Application, Qinghai University, Xining 810016, China"},{"name":"Intelligent Computing and Application Laboratory of Qinghai Province, Qinghai University, Xining 810016, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaolin","family":"Hu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology, Institute for AI, BNRist, Tsinghua Laboratory of Brain and Intelligence (THBI), IDG\/McGovern Institute for Brain Research, Tsinghua University, Beijing 100084, China"},{"name":"Chinese Institute for Brain Research (CIBR), Beijing 100010, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"7437","container-title":["Frontiers in Artificial Intelligence and Applications","ECAI 2025"],"original-title":[],"link":[{"URL":"https:\/\/ebooks.iospress.nl\/pdf\/doi\/10.3233\/FAIA250845","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T09:43:55Z","timestamp":1761126235000},"score":1,"resource":{"primary":{"URL":"https:\/\/ebooks.iospress.nl\/doi\/10.3233\/FAIA250845"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,21]]},"ISBN":["9781643686318"],"references-count":0,"URL":"https:\/\/doi.org\/10.3233\/faia250845","relation":{},"ISSN":["0922-6389","1879-8314"],"issn-type":[{"value":"0922-6389","type":"print"},{"value":"1879-8314","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,21]]}}}