{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T15:40:57Z","timestamp":1767973257757,"version":"3.49.0"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T00:00:00Z","timestamp":1761177600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T00:00:00Z","timestamp":1761177600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"SERICS"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Discov Computing"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The rapid evolution of audio deepfakes has raised significant challenges for the security and reliability of voice-driven systems. While recent detection frameworks achieve high accuracy in controlled environments, their performance often degrades under real-world conditions involving codec compression, signal preprocessing, or domain shifts. To address these challenges, we propose dynamic knowledge condensation with audio-selective transformer (DK-CAST), a novel tri-stream knowledge distillation framework designed for robust audio deepfake detection. DK-CAST employs a high-capacity XLS-R teacher trained on clean speech to supervise a compact student model operating on degraded and preprocessed audio. The student employs a custom audio-selective transformer with dual-stream encoding, dynamic fusion, and phoneme-gated attention to emphasize linguistically relevant cues. Knowledge is transferred via multi-level supervision, including logits, embeddings, and phoneme posteriors, and modulated through a codec-aware loss weighting scheme. To enhance generalization, DK-CAST also includes a compression-agnostic embedding alignment module based on MMD and Center Loss. Evaluations on ASVspoof 2019-LA and ASVspoof 2021-DF demonstrate state-of-the-art performance, achieving EERs of 0.38 and 2.18%, respectively. Furthermore, DK-CAST maintains strong performance under codec degradation, achieving an EER of 3.01% on ASVspoof 2021-DF when tested under MP3 compression.<\/jats:p>","DOI":"10.1007\/s10791-025-09746-4","type":"journal-article","created":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T17:19:39Z","timestamp":1761239979000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Dynamic knowledge condensation with audio-selective transformer for audio deepfake detection"],"prefix":"10.1007","volume":"28","author":[{"given":"Taiba Maijd","family":"Wani","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Irene","family":"Amerini","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,10,23]]},"reference":[{"issue":"3\u20134","key":"9746_CR1","doi-asserted-by":"publisher","first-page":"153","DOI":"10.1561\/3300000048","volume":"6","author":"TM Wani","year":"2024","unstructured":"Wani TM, Qadri SAA, Wani FA, Amerini I, et al. Navigating the soundscape of deception: a comprehensive survey on audio deepfake generation, detection, and future horizons. Foundations Trends\u00ae Privacy Secur. 2024;6(3\u20134):153\u2013345.","journal-title":"Foundations Trends\u00ae Privacy Secur"},{"key":"9746_CR2","unstructured":"Yi J, Wang C, Tao J, Zhang X, Zhang CY, Zhao Y, Audio deepfake detection: a survey. arXiv preprint 2023. arXiv:2308.14970."},{"key":"9746_CR3","doi-asserted-by":"publisher","first-page":"1001063","DOI":"10.3389\/fdata.2022.1001063","volume":"5","author":"Z Khanjani","year":"2023","unstructured":"Khanjani Z, Watson G, Janeja VP. Audio deepfakes: a survey. Front Big Data. 2023;5:1001063.","journal-title":"Front Big Data"},{"key":"9746_CR4","doi-asserted-by":"crossref","unstructured":"Todisco M, Wang X, Vestman V, Sahidullah M, Delgado H, Nautsch A, Yamagishi J, Evans N, Kinnunen T, Lee KA. Asvspoof 2019: future horizons in spoofed and fake audio detection. 2019. arXiv preprint arXiv:1904.05441.","DOI":"10.21437\/Interspeech.2019-2249"},{"key":"9746_CR5","doi-asserted-by":"publisher","first-page":"2507","DOI":"10.1109\/TASLP.2023.3285283","volume":"31","author":"X Liu","year":"2023","unstructured":"Liu X, Wang X, Sahidullah M, Patino J, Delgado H, Kinnunen T, et al. Asvspoof 2021: towards spoofed and deepfake speech detection in the wild. IEEE\/ACM Trans Audio Speech Language Process. 2023;31:2507\u201322.","journal-title":"IEEE\/ACM Trans Audio Speech Language Process"},{"issue":"7","key":"9746_CR6","doi-asserted-by":"publisher","first-page":"1989","DOI":"10.3390\/s25071989","volume":"25","author":"B Zhang","year":"2025","unstructured":"Zhang B, Cui H, Nguyen V, Whitty M. Audio deepfake detection: what has been achieved and what lies ahead. Sensors (Basel Switzerland). 2025;25(7):1989.","journal-title":"Sensors (Basel Switzerland)"},{"key":"9746_CR7","doi-asserted-by":"crossref","unstructured":"M\u00fcller NM, Kawa P, Hu S, Neu M, Williams J, Sperl P, B\u00f6ttinger K. A new approach to voice authenticity. 2024. arXiv preprint arXiv:2402.06304.","DOI":"10.21437\/Interspeech.2024-31"},{"key":"9746_CR8","doi-asserted-by":"publisher","first-page":"1009","DOI":"10.1007\/s11042-016-4277-2","volume":"77","author":"M Zakariah","year":"2018","unstructured":"Zakariah M, Khan MK, Malik H. Digital multimedia audio forensics: past, present and future. Multimed Tools Appl. 2018;77:1009\u201340.","journal-title":"Multimed Tools Appl"},{"key":"9746_CR9","unstructured":"Bousba A. Ai-based online api for fake speech detection. Ph.D. thesis, University Larbi T\u00e9bessi\u2013T\u00e9bessa. 2024."},{"issue":"1","key":"9746_CR10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.3390\/computers14010001","volume":"14","author":"D Ghiur\u0103u","year":"2024","unstructured":"Ghiur\u0103u D, Popescu DE. Distinguishing reality from AI: Approaches for detecting synthetic content. Computers. 2024;14(1):1.","journal-title":"Computers"},{"key":"9746_CR11","unstructured":"Norlin A. The effectiveness of knowledge distillation methods for real-time deepfake solutions. 2024."},{"key":"9746_CR12","doi-asserted-by":"crossref","unstructured":"Chen P, Liu S, Zhao H, Jia J. Distilling knowledge via knowledge review. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 2021. p. 5008\u20135017.","DOI":"10.1109\/CVPR46437.2021.00497"},{"key":"9746_CR13","first-page":"33716","volume":"35","author":"T Huang","year":"2022","unstructured":"Huang T, You S, Wang F, Qian C, Xu C. Knowledge distillation from a stronger teacher. Adv Neural Inf Process Syst. 2022;35:33716\u201327.","journal-title":"Adv Neural Inf Process Syst"},{"key":"9746_CR14","unstructured":"Gong Y, Khurana S, Rouditchenko A, Glass J. Cmkd: Cnn\/transformer-based cross-model knowledge distillation for audio classification. 2022. arXiv preprint arXiv:2203.06760."},{"issue":"7","key":"9746_CR15","doi-asserted-by":"publisher","first-page":"2046","DOI":"10.3390\/s24072046","volume":"24","author":"D Priebe","year":"2024","unstructured":"Priebe D, Ghani B, Stowell D. Efficient speech detection in environmental audio using acoustic recognition and knowledge distillation. Sensors. 2024;24(7):2046.","journal-title":"Sensors"},{"key":"9746_CR16","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2024.108341","volume":"133","author":"B Wang","year":"2024","unstructured":"Wang B, Wu X, Wang F, Zhang Y, Wei F, Song Z. Spatial-frequency feature fusion based deepfake detection through knowledge distillation. Eng Appl Artif Intell. 2024;133:108341.","journal-title":"Eng Appl Artif Intell"},{"key":"9746_CR17","doi-asserted-by":"publisher","first-page":"4905","DOI":"10.1109\/TASLP.2024.3492796","volume":"32","author":"B Wang","year":"2024","unstructured":"Wang B, Tang Y, Wei F, Ba Z, Ren K. Ftdkd: frequency-time domain knowledge distillation for low-quality compressed audio deepfake detection. IEEE\/ACM Trans Audio Speech Language Process. 2024;32:4905\u201318.","journal-title":"IEEE\/ACM Trans Audio Speech Language Process"},{"key":"9746_CR18","doi-asserted-by":"crossref","unstructured":"Lu J, Zhang Y, Wang W, Shang Z, Zhang P. One-class knowledge distillation for spoofing speech detection. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (IEEE, 2024). p. 11251\u201311255.","DOI":"10.1109\/ICASSP48485.2024.10446270"},{"key":"9746_CR19","doi-asserted-by":"crossref","unstructured":"Fan C, Dong S, Xue J, Chen Y, Yi J, Lv Z. Frequency-mix knowledge distillation for fake speech detection. 2024. arXiv preprint arXiv:2406.09664.","DOI":"10.21437\/Interspeech.2024-740"},{"key":"9746_CR20","doi-asserted-by":"crossref","unstructured":"Babu A, Wang C, Tjandra A, Lakhotia K, Xu Q, Goyal N, Singh K, Von\u00a0Platen P, Saraf Y, Pino J et\u00a0al. Xls-r: self-supervised cross-lingual speech representation learning at scale. 2021. arXiv preprint arXiv:2111.09296.","DOI":"10.21437\/Interspeech.2022-143"},{"key":"9746_CR21","doi-asserted-by":"crossref","unstructured":"Teytaut Y, Bouvier B, Roebel A. A study on constraining Connectionist Temporal Classification for temporal audio alignment. In: Interspeech 2022 (ISCA, 2022). p. 5015\u20135019.","DOI":"10.21437\/Interspeech.2022-10940"},{"key":"9746_CR22","unstructured":"Nguyen\u00a0Le TD, Teh KK, Dat\u00a0Tran H. Continuous learning of transformer-based audio deepfake detection. 2024. arXiv e-prints. arXiv\u20132409."},{"key":"9746_CR23","doi-asserted-by":"publisher","first-page":"4360","DOI":"10.1109\/TMM.2023.3321505","volume":"26","author":"Y Ren","year":"2023","unstructured":"Ren Y, Peng H, Li L, Yang Y. Lightweight voice spoofing detection using improved one-class learning and knowledge distillation. IEEE Trans Multimed. 2023;26:4360\u201374.","journal-title":"IEEE Trans Multimed"},{"key":"9746_CR24","doi-asserted-by":"crossref","unstructured":"Liu P, Zhang Z, Yang Y. End-to-end spoofing speech detection and knowledge distillation under noisy conditions. In: 2021 International Joint Conference on Neural Networks (IJCNN) (IEEE, 2021). p. 1\u20137.","DOI":"10.1109\/IJCNN52387.2021.9534312"},{"key":"9746_CR25","doi-asserted-by":"crossref","unstructured":"An W, Li R, Ge H, Li M, Li H. An End-to-End Audio Transformer with Multi-student Knowledge Distillation algorithm for Deepfake Speech Detection. In: Proceedings of the 2024 13th International Conference on Computing and Pattern Recognition. 2024. p. 366\u2013371.","DOI":"10.1145\/3704323.3704334"},{"key":"9746_CR26","doi-asserted-by":"crossref","unstructured":"Zhang K, Hua Z, Zhang Y, Guo Y, Xiang T. Robust AI-synthesized speech detection using feature decomposition learning and synthesizer feature augmentation. In: IEEE Transactions on Information Forensics and Security. 2024.","DOI":"10.1109\/TIFS.2024.3520001"},{"key":"9746_CR27","first-page":"1","volume":"99","author":"J Xue","year":"2024","unstructured":"Xue J, Fan C, Yi J, Zhou J, Lv Z. Dynamic ensemble teacher-student distillation framework for light-weight fake audio detection. IEEE Signal Process Lett. 2024;99:1\u20135.","journal-title":"IEEE Signal Process Lett"},{"key":"9746_CR28","doi-asserted-by":"publisher","first-page":"2453","DOI":"10.1109\/TASLP.2024.3389643","volume":"32","author":"C Fan","year":"2024","unstructured":"Fan C, Ding M, Tao J, Fu R, Yi J, Wen Z, et al. Dual-branch knowledge distillation for noise-robust synthetic speech detection. IEEE\/ACM Trans Audio Speech Language Process. 2024;32:2453\u201366.","journal-title":"IEEE\/ACM Trans Audio Speech Language Process"},{"key":"9746_CR29","first-page":"1075","volume":"39","author":"K Zhang","year":"2025","unstructured":"Zhang K, Hua Z, Lan R, Guo Y, Zhang Y, Xu G. Multi-view collaborative learning network for speech deepfake detection. Proc AAAI Conf Artif Intell. 2025;39:1075\u201383.","journal-title":"Proc AAAI Conf Artif Intell."},{"key":"9746_CR30","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2024.129256","volume":"620","author":"S Usmani","year":"2025","unstructured":"Usmani S, Kumar S, Sadhya D. Spatio-temporal knowledge distilled video vision transformer (stkd-vvit) for multimodal deepfake detection. Neurocomputing. 2025;620:129256.","journal-title":"Neurocomputing"},{"key":"9746_CR31","doi-asserted-by":"crossref","unstructured":"Nuha HH, Absa AA. Noise reduction and speech enhancement using wiener filter. In: 2022 International Conference on Data Science and Its Applications (ICoDSA) (IEEE, 2022). p. 177\u2013180.","DOI":"10.1109\/ICoDSA55874.2022.9862912"},{"issue":"1","key":"9746_CR32","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1121\/10.0005314","volume":"150","author":"RM Corey","year":"2021","unstructured":"Corey RM, Singer AC. Modeling the effects of dynamic range compression on signals in noise. J Acoust Soc Am. 2021;150(1):159\u201370.","journal-title":"J Acoust Soc Am"},{"key":"9746_CR33","first-page":"1066","volume":"39","author":"K Zhang","year":"2025","unstructured":"Zhang K, Hua Z, Lan R, Zhang Y, Guo Y. Phoneme-level feature discrepancies: a key to detecting sophisticated speech deepfakes. Proc AAAI Conf Artif Intell. 2025;39:1066\u201374.","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"9746_CR34","unstructured":"Hinton G, Vinyals O, Dean J. Distilling the knowledge in a neural network. 2015. arXiv preprint arXiv:1503.02531."},{"key":"9746_CR35","doi-asserted-by":"crossref","unstructured":"Reimao R, Tzerpos V. For: A dataset for synthetic speech detection. In: 2019 International Conference on Speech Technology and Human-Computer Dialogue (SpeD) (IEEE, 2019). p. 1\u201310.","DOI":"10.1109\/SPED.2019.8906599"},{"key":"9746_CR36","doi-asserted-by":"crossref","unstructured":"Wani TM, Qadri SAA, Comminiello D, Amerini I. Detecting audio deepfakes: integrating CNN and BiLSTM with multi-feature concatenation. In: Proceedings of the 2024 ACM Workshop on Information Hiding and Multimedia Security 2024. p. 271\u2013276.","DOI":"10.1145\/3658664.3659647"},{"key":"9746_CR37","unstructured":"Le TDN, Teh KK, Tran HD. Continuous learning of transformer-based audio deepfake detection. 2024. arXiv preprint arXiv:2409.05924."},{"key":"9746_CR38","doi-asserted-by":"crossref","unstructured":"Tak H, Patino J, Todisco M, Nautsch A, Evans N, Larcher A. End-to-end anti-spoofing with rawnet2. In: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (IEEE, 2021). p. 6369\u20136373.","DOI":"10.1109\/ICASSP39728.2021.9414234"},{"key":"9746_CR39","doi-asserted-by":"crossref","unstructured":"Tak H, Todisco M, Wang X, Jung Jw, Yamagishi J, Evans N. Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation. 2022. arXiv preprint arXiv:2202.12233.","DOI":"10.21437\/Odyssey.2022-16"},{"key":"9746_CR40","unstructured":"Gao C, Postiglione M, Gortner I, Kraus S, Subrahmanian VS. Perturbed public voices ($$p^2 v$$): A dataset for robust audio deepfake detection. 2025. arXiv preprint arXiv:2508.10949."}],"container-title":["Discover Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10791-025-09746-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10791-025-09746-4","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10791-025-09746-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T12:33:12Z","timestamp":1767961992000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10791-025-09746-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,23]]},"references-count":40,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["9746"],"URL":"https:\/\/doi.org\/10.1007\/s10791-025-09746-4","relation":{},"ISSN":["2948-2992"],"issn-type":[{"value":"2948-2992","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,23]]},"assertion":[{"value":"26 May 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 October 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 October 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable. This study does not involve human participants relies only on publicly available datasets.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable. This manuscript does not contain any individual person\u2019s data in any form (including individual details, images, or videos).","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"231"}}