{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T16:21:31Z","timestamp":1781713291984,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":33,"publisher":"ACM","funder":[{"name":"EU Horizon Europe","award":["101070093"],"award-info":[{"award-number":["101070093"]}]},{"name":"German Ministry of Education and Research BMBF","award":["03RU2U151D"],"award-info":[{"award-number":["03RU2U151D"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,6,30]]},"DOI":"10.1145\/3733567.3735565","type":"proceedings-article","created":{"date-parts":[[2025,7,14]],"date-time":"2025-07-14T09:27:42Z","timestamp":1752485262000},"page":"55-62","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Calibrating POI-based Synthetic Speech Detection"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-5290-7695","authenticated-orcid":false,"given":"Thomas","family":"Le Roux","sequence":"first","affiliation":[{"name":"Fraunhofer IDMT, Ilmenau, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5559-6508","authenticated-orcid":false,"given":"Luca","family":"Cuccovillo","sequence":"additional","affiliation":[{"name":"Fraunhofer IDMT, Ilmenau, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4777-6335","authenticated-orcid":false,"given":"Patrick","family":"Aichroth","sequence":"additional","affiliation":[{"name":"Fraunhofer IDMT, Ilmenau, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,7,14]]},"reference":[{"key":"e_1_3_3_2_2_2","first-page":"38","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)","author":"Agarwal Shruti","year":"2019","unstructured":"Shruti Agarwal, Hany Farid, Yuming Gu, Mingming He, Koki Nagano, and Hao Li. 2019. Protecting world leaders against deep fakes. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Long Beach, California, USA, 38\u201345."},{"key":"e_1_3_3_2_3_2","first-page":"2278","volume-title":"Annual Conference of the International Speech Communication Association (ISCA Interspeech)","author":"Babu Arun","year":"2022","unstructured":"Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli. 2022. XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale. In Annual Conference of the International Speech Communication Association (ISCA Interspeech). Incheon, Korea, 2278\u20132282."},{"key":"e_1_3_3_2_4_2","volume-title":"Annual Conference on Neural Information Processing Systems (NeurIPS)","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. In Annual Conference on Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/342009.335388"},{"key":"e_1_3_3_2_6_2","doi-asserted-by":"crossref","unstructured":"Sanyuan Chen Chengyi Wang Zhengyang Chen Yu Wu Shujie Liu Zhuo Chen Jinyu Li Naoyuki Kanda Takuya Yoshioka Xiong Xiao Jian Wu Long Zhou Shuo Ren Yanmin Qian Yao Qian Jian Wu Michael Zeng Xiangzhan Yu and Furu Wei. 2022. WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing. IEEE Journal of Selected Topics in Signal Processing 16 6 (2022) 1505\u20131518.","DOI":"10.1109\/JSTSP.2022.3188113"},{"key":"e_1_3_3_2_7_2","first-page":"5178","volume-title":"International Conference on Machine Learning (ICML)","volume":"202","author":"Chen Sanyuan","year":"2023","unstructured":"Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, Wanxiang Che, Xiangzhan Yu, and Furu Wei. 2023. BEATs: Audio Pre-Training with Acoustic Tokenizers. In International Conference on Machine Learning (ICML) , Vol.\u00a0202. PMLR, Honolulu, Hawaii, USA, 5178\u20135193."},{"key":"e_1_3_3_2_8_2","first-page":"4409","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)","author":"Cuccovillo Luca","year":"2024","unstructured":"Luca Cuccovillo, Milica Gerhardt, and Patrick Aichroth. 2024. Audio Transformer for Synthetic Speech Detection via Multi-Formant Analysis. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Seattle, Washington, USA, 4409\u20134417."},{"key":"e_1_3_3_2_9_2","first-page":"1","volume-title":"IEEE International Workshop on Information Forensics and Security (WIFS)","author":"Cuccovillo Luca","year":"2022","unstructured":"Luca Cuccovillo, Christoforos Papastergiopoulos, Anastasios Vafeiadis, Artem Yaroshchuk, Patrick Aichroth, Konstantinos Votis, and Dimitrios Tzovaras. 2022. Open Challenges in Synthetic Speech Detection. In IEEE International Workshop on Information Forensics and Security (WIFS). Shanghai, China, 1\u20136."},{"key":"e_1_3_3_2_10_2","unstructured":"EBU. 2023. R128-2020: Loudness Normalization and Permitted Maximum Level of Audio Signals. https:\/\/tech.ebu.ch\/docs\/r\/r128.pdf [Accessed 09-04-2025]."},{"key":"e_1_3_3_2_11_2","first-page":"70","volume-title":"ISCA Odyssey: The Speaker and Language Recognition Workshop","author":"Ge Wanying","year":"2022","unstructured":"Wanying Ge, Massimiliano Todisco, and Nicholas Evans. 2022. Explainable Deepfake and Spoofing Detection: An Attack Analysis Using SHapley Additive exPlanations. In ISCA Odyssey: The Speaker and Language Recognition Workshop. Beijing, China, 70\u201376."},{"key":"e_1_3_3_2_12_2","doi-asserted-by":"crossref","unstructured":"Romain Hennequin Anis Khlif Felix Voituret and Manuel Moussallam. 2020. Spleeter: A fast and efficient music source separation tool with pre-trained models. Journal of Open Source Software 5 50 (2020) 2154.","DOI":"10.21105\/joss.02154"},{"key":"e_1_3_3_2_13_2","doi-asserted-by":"crossref","unstructured":"Wei-Ning Hsu Benjamin Bolte Yao-Hung\u00a0Hubert Tsai Kushal Lakhotia Ruslan Salakhutdinov and Abdelrahman Mohamed. 2021. HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units. IEEE\/ACM Transactions on Audio Speech and Language Processing 29 (2021) 3451\u20133460.","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_3_3_2_14_2","doi-asserted-by":"crossref","unstructured":"Jee-weon Jung Yihan Wu Xin Wang Ji-Hoon Kim Soumi Maiti Yuta Matsunaga Hye-jin Shim Jinchuan Tian Nicholas Evans Joon\u00a0Son Chung et\u00a0al. 2025. SpoofCeleb: Speech Deepfake Detection and SASV In The Wild. IEEE Open Journal of Signal Processing 6 (2025) 68\u201377.","DOI":"10.1109\/OJSP.2025.3529377"},{"key":"e_1_3_3_2_15_2","doi-asserted-by":"crossref","unstructured":"Kaylo\u00a0T Littlejohn Cheol\u00a0Jun Cho Jessie\u00a0R Liu Alexander\u00a0B Silva Bohan Yu Vanessa\u00a0R Anderson Cady\u00a0M Kurtz-Miott Samantha Brosler Anshul\u00a0P Kashyap Irina\u00a0P Hallinan et\u00a0al. 2025. A streaming brain-to-voice neuroprosthesis to restore naturalistic communication. Nature Neuroscience 28 4 (2025) 1\u201311.","DOI":"10.1038\/s41593-025-01905-6"},{"key":"e_1_3_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2008.17"},{"key":"e_1_3_3_2_17_2","first-page":"1","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Luong Hieu-Thi","year":"2025","unstructured":"Hieu-Thi Luong, Haoyang Li, Lin Zhang, Kong\u00a0Aik Lee, and Eng\u00a0Siong Chng. 2025. LlamaPartialSpoof: An LLM-driven fake speech dataset simulating disinformation generation. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Hyderabad, India, 1\u20135."},{"key":"e_1_3_3_2_18_2","doi-asserted-by":"crossref","unstructured":"Yi Ma Shuai Wang Tianchi Liu and Haizhou Li. 2025. ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification. IEEE Signal Processing Letters 32 (2025) 731\u2013735.","DOI":"10.1109\/LSP.2025.3530850"},{"key":"e_1_3_3_2_19_2","first-page":"2783","volume-title":"Annual Conference of the International Speech Communication Association (ISCA Interspeech)","author":"M\u00fcller Nicolas","year":"2022","unstructured":"Nicolas M\u00fcller, Pavel Czempin, Franziska Diekmann, Adam Froghyar, and Konstantin B\u00f6ttinger. 2022. Does Audio Deepfake Detection Generalize?. In Annual Conference of the International Speech Communication Association (ISCA Interspeech). Incheon, Korea, 2783\u20132787."},{"key":"e_1_3_3_2_20_2","first-page":"55","volume-title":"Automatic Speaker Verification and Spoofing Countermeasures Challenge (ASVspoof)","author":"M\u00fcller Nicolas","year":"2021","unstructured":"Nicolas M\u00fcller, Franziska Dieckmann, Pavel Czempin, Roman Canals, Konstantin B\u00f6ttinger, and Jennifer Williams. 2021. Speech is Silver, Silence is Golden: What do ASVspoof-trained Models Really Learn?. In Automatic Speaker Verification and Spoofing Countermeasures Challenge (ASVspoof). Brno, Czech Republic, 55\u201360."},{"key":"e_1_3_3_2_21_2","first-page":"2616","volume-title":"Annual Conference of the International Speech Communication Association (ISCA Interspeech)","author":"Nagraniy Arsha","year":"2017","unstructured":"Arsha Nagraniy, Joon\u00a0Son Chungy, and Andrew Zisserman. 2017. VoxCeleb: A large-scale speaker identification dataset. In Annual Conference of the International Speech Communication Association (ISCA Interspeech). Stockholm, Sweden, 2616\u20132620."},{"key":"e_1_3_3_2_22_2","unstructured":"NBCNews. 2024. Why AI-generated audio is so hard to detect. https:\/\/www.nbcnews.com\/tech\/misinformation\/ai-generated-audio-detect-tool-model-rcna136634 [Accessed 09-04-2025]."},{"key":"e_1_3_3_2_23_2","unstructured":"NPR.org. 2024. How AI deepfakes polluted elections in 2024. https:\/\/www.npr.org\/2024\/12\/21\/nx-s1-5220301\/deepfakes-memes-artificial-intelligence-elections [Accessed 09-04-2025]."},{"key":"e_1_3_3_2_24_2","first-page":"1","volume-title":"IEEE International Workshop on Information Forensics and Security (WIFS)","author":"Pianese Alessandro","year":"2022","unstructured":"Alessandro Pianese, Davide Cozzolino, Giovanni Poggi, and Luisa Verdoliva. 2022. Deepfake audio detection by speaker verification. In IEEE International Workshop on Information Forensics and Security (WIFS). Shanghai, China, 1\u20136."},{"key":"e_1_3_3_2_25_2","first-page":"289","volume-title":"ACM Workshop on Information Hiding and Multimedia Security (IH&MMSec)","author":"Pianese Alessandro","year":"2024","unstructured":"Alessandro Pianese, Davide Cozzolino, Giovanni Poggi, and Luisa Verdoliva. 2024. Training-Free Deepfake Voice Recognition by Leveraging Large-Scale Pre-Trained Models. In ACM Workshop on Information Hiding and Multimedia Security (IH&MMSec). Baiona, Spain, 289\u2013294."},{"key":"e_1_3_3_2_26_2","first-page":"28492","volume-title":"International conference on machine learning (ICML)","author":"Radford Alec","year":"2023","unstructured":"Alec Radford, Jong\u00a0Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning (ICML). PMLR, Honolulu, Hawaii, USA, 28492\u201328518."},{"key":"e_1_3_3_2_27_2","first-page":"1","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Sivaraman Ganesh","year":"2025","unstructured":"Ganesh Sivaraman, Hemlata Tak, and Elie Khoury. 2025. Investigating voiced and unvoiced regions of speech for audio deepfake detection. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Hyderabad, India, 1\u20135."},{"key":"e_1_3_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414234"},{"key":"e_1_3_3_2_29_2","unstructured":"Silero Team. 2024. Silero VAD: pre-trained enterprise-grade Voice Activity Detector (VAD) Number Detector and Language Classifier. https:\/\/github.com\/snakers4\/silero-vad"},{"key":"e_1_3_3_2_30_2","unstructured":"theconversation.com. 2024. The apocalypse that wasn\u2019t: AI was everywhere in 2024\u2019s elections but deepfakes and misinformation were only part of the picture. https:\/\/theconversation.com\/the-apocalypse-that-wasnt-ai-was-everywhere-in-2024s-elections-but-deepfakes-and-misinformation-were-only-part-of-the-picture-244225 [Accessed 09-04-2025]."},{"key":"e_1_3_3_2_31_2","unstructured":"TheGuardian. 2024. Company that sent fake Biden robocalls in New Hampshire agrees to $1M fine. https:\/\/www.theguardian.com\/technology\/article\/2024\/aug\/22\/fake-biden-robocalls-fine-lingo-telecom [Accessed 09-04-2025]."},{"key":"e_1_3_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Xin Wang Junichi Yamagishi Massimiliano Todisco H\u00e9ctor Delgado Andreas Nautsch Nicholas Evans Md Sahidullah Ville Vestman Tomi Kinnunen Kong\u00a0Aik Lee et\u00a0al. 2020. ASVspoof 2019: A large-scale public database of synthesized converted and replayed speech. Computer Speech & Language 64 (2020) 101114.","DOI":"10.1016\/j.csl.2020.101114"},{"key":"e_1_3_3_2_33_2","doi-asserted-by":"crossref","unstructured":"Yang Xiao and Rohan\u00a0Kumar Das. 2025. XLSR-Mamba: A dual-column bidirectional state space model for spoofing attack detection. IEEE Signal Processing Letters 32 (2025) 1276\u20131280.","DOI":"10.1109\/LSP.2025.3547861"},{"key":"e_1_3_3_2_34_2","first-page":"6765","volume-title":"ACM International Conference on Multimedia","author":"Zhang Qishan","year":"2024","unstructured":"Qishan Zhang, Shuangbing Wen, and Tao Hu. 2024. Audio Deepfake Detection with Self-Supervised XLS-R and SLS Classifier. In ACM International Conference on Multimedia. Melbourne VIC, Australia, 6765\u20136773."}],"event":{"name":"MAD'25: 4th ACM International Workshop on Multimedia AI against Disinformation","location":"Chicago USA","acronym":"MAD'25","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 4th ACM International Workshop on Multimedia AI against Disinformation"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3733567.3735565","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,14]],"date-time":"2025-07-14T09:28:52Z","timestamp":1752485332000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3733567.3735565"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,30]]},"references-count":33,"alternative-id":["10.1145\/3733567.3735565","10.1145\/3733567"],"URL":"https:\/\/doi.org\/10.1145\/3733567.3735565","relation":{},"subject":[],"published":{"date-parts":[[2025,6,30]]},"assertion":[{"value":"2025-07-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}