{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,16]],"date-time":"2026-01-16T03:31:45Z","timestamp":1768534305806,"version":"3.49.0"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"8","funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2022YFE0138600"],"award-info":[{"award-number":["2022YFE0138600"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,8,31]]},"abstract":"<jats:p>AI-empowered edge computing has given rise to a new paradigm and effectively facilitated the promotion and development of multimedia applications. The speech assistant is one of the significant services provided by multimedia applications, which aims to offer intelligent interactive experiences between humans and machines. However, malicious attackers may exploit spoofed speeches to deceive speech assistants, posing great challenges to the security of multimedia applications. The limited resources of multimedia terminal devices hinder their ability to effectively load speech spoofing detection models. Furthermore, processing and analyzing speech in the cloud can result in poor real-time performance and potential privacy risks. Existing speech spoofing detection methods rely heavily on annotated data and exhibit poor generalization capabilities for unseen spoofed speeches. To address these challenges, this article first proposes the Coordinate Attention Network (CA2Net) that consists of coordinate attention blocks and Res2Net blocks. CA2Net can simultaneously extract temporal and spectral speech feature information and represent multi-scale speech features at a granularity level. Besides, a contrastive learning-based speech spoofing detection framework named GEMINI is proposed. GEMINI can be effectively deployed on edge nodes and autonomously learn speech features with strong generalization capabilities. GEMINI first performs data augmentation on speech signals and extracts conventional acoustic features to enhance the feature robustness. Subsequently, GEMINI utilizes the proposed CA2Net to further explore the discriminative speech features. Then, a tensor-based multi-attention comparison model is employed to maximize the consistency between speech contexts. GEMINI continuously updates CA2Net with contrastive learning, which enables CA2Net to effectively represent speech signals and accurately detect spoofed speeches. Extensive experiments on the ASVspoof2019 dataset show that GEMINI reduces the Equal Error Rate and tandem Detection Cost Function by up to 96.75% and 96.35% in the physical access scenario, and by up to 86.62% and 87.71% in the logical access scenario compared to peer methods.<\/jats:p>","DOI":"10.1145\/3698773","type":"journal-article","created":{"date-parts":[[2024,10,7]],"date-time":"2024-10-07T11:29:47Z","timestamp":1728300587000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Contrastive Learning-Based Speech Spoofing Detection for Multimedia Security in Edge Intelligence"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-4282-7997","authenticated-orcid":false,"given":"Jiaqi","family":"Sun","sequence":"first","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2641-1708","authenticated-orcid":false,"given":"Xianjun","family":"Deng","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8993-8940","authenticated-orcid":false,"given":"Shenghao","family":"Liu","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3580-432X","authenticated-orcid":false,"given":"Xiaoxuan","family":"Fan","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-7551-635X","authenticated-orcid":false,"given":"Yongling","family":"Huang","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6706-4859","authenticated-orcid":false,"given":"Yuanyuan","family":"He","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6853-5878","authenticated-orcid":false,"given":"Celimuge","family":"Wu","sequence":"additional","affiliation":[{"name":"The University of Electro-Communications, Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1831-0309","authenticated-orcid":false,"given":"Jong","family":"Park","sequence":"additional","affiliation":[{"name":"Seoul National University of Science and Technology, Seoul, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,8,12]]},"reference":[{"key":"e_1_3_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/UPCON56432.2022.9986418"},{"key":"e_1_3_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10097131"},{"key":"e_1_3_1_4_1","doi-asserted-by":"crossref","unstructured":"Salahaldeen Duraibi Wasim Alhamdani and Frederick T. Sheldon. 2020. Voice Feature Learning Using Convolutional Neural Networks Designed to Avoid Replay Attacks. In Proceedings of the 2020 IEEE Symposium Series on Computational Intelligence (SSCI \u201920) 1845\u20131851.","DOI":"10.1109\/SSCI47803.2020.9308489"},{"key":"e_1_3_1_5_1","doi-asserted-by":"crossref","unstructured":"Xiaotao Feng Xiaogang Zhu Qing-Long Han Wei Zhou Sheng Wen and Yang Xiang. 2023. Detecting Vulnerability on IoT Device Firmware: A Survey. IEEE CAA J. Autom. Sinica 10 1 (2023) 25\u201341.","DOI":"10.1109\/JAS.2022.105860"},{"key":"e_1_3_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746722"},{"key":"e_1_3_1_7_1","doi-asserted-by":"crossref","unstructured":"Shanghua Gao Ming-Ming Cheng Kai Zhao Xin-Yu Zhang Ming-Hsuan Yang and Philip H. S. Torr. 2021. Res2Net: A New Multi-Scale Backbone Architecture. IEEE Trans. Pattern Anal. Mach. Intell. 43 2 (2021) 652\u2013662.","DOI":"10.1109\/TPAMI.2019.2938758"},{"key":"e_1_3_1_8_1","doi-asserted-by":"publisher","unstructured":"Baoshen Guo Shuai Wang Yi Ding Guang Wang Suining He Desheng Zhang and Tian He. 2021. Concurrent Order Dispatch for Instant Delivery with Time-Constrained Actor-Critic Reinforcement Learning. In Proceedings of the 2021 IEEE Real-Time Systems Symposium (RTSS \u201921) 176\u2013187. DOI: 10.1109\/RTSS52674.2021.00026","DOI":"10.1109\/RTSS52674.2021.00026"},{"key":"e_1_3_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_1_10_1","doi-asserted-by":"crossref","unstructured":"Xiangyu Hu Wanlun Ma Chao Chen Sheng Wen Jun Zhang Yang Xiang and Gaolei Fei. 2022. Event Detection in Online Social Network: Methodologies State-of-Art and Evolution. Comput. Sci. Rev. 46 C (2022) 25 pages.","DOI":"10.1016\/j.cosrev.2022.100500"},{"key":"e_1_3_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICNP.2017.8117585"},{"key":"e_1_3_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3462244.3479883"},{"key":"e_1_3_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3460820"},{"key":"e_1_3_1_14_1","doi-asserted-by":"publisher","unstructured":"Maoran Jiang and Wei Gong. 2024. Bidirectional Bluetooth Backscatter with Edges. IEEE Trans. Mobile Comput. 23 2 (2024) 1601\u20131612. DOI: 10.1109\/TMC.2023.3241202","DOI":"10.1109\/TMC.2023.3241202"},{"key":"e_1_3_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/APSIPAASC58517.2023.10317474"},{"key":"e_1_3_1_16_1","doi-asserted-by":"crossref","unstructured":"Matan Karo Arie Yeredor and Itshak Lapidot. 2024. Compact Time-Domain Representation for Logical Access Spoofed Audio. IEEE\/ACM Trans. Audio Speech Lang. Process. 32 (2024) 946\u2013958.","DOI":"10.1109\/TASLP.2023.3341000"},{"key":"e_1_3_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10096672"},{"key":"e_1_3_1_18_1","doi-asserted-by":"crossref","unstructured":"Tomi Kinnunen H\u00e9ctor Delgado Nicholas W. D. Evans Kong Aik Lee Ville Vestman Andreas Nautsch Massimiliano Todisco Xin Wang Md. Sahidullah Junichi Yamagishi et\u00a0al. 2020. Tandem Assessment of Spoofing Countermeasures and Automatic Speaker Verification: Fundamentals. IEEE ACM Trans. Audio Speech Lang. Process. 28 (2020) 2195\u20132210.","DOI":"10.1109\/TASLP.2020.3009494"},{"key":"e_1_3_1_19_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-1794"},{"key":"e_1_3_1_20_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-1768"},{"key":"e_1_3_1_21_1","doi-asserted-by":"crossref","unstructured":"Jiguo Li Xinfeng Zhang Jizheng Xu Siwei Ma and Wen Gao. 2021. Learning to Fool the Speaker Recognition. ACM Trans. Multimedia Comput. Commun. Appl. 17 3s (2021) 21 pages.","DOI":"10.1145\/3468673"},{"key":"e_1_3_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2020.2993293"},{"key":"e_1_3_1_23_1","doi-asserted-by":"crossref","unstructured":"Chang Liu Zhen-Hua Ling and Ling-Hui Chen. 2023. Pronunciation Dictionary-Free Multilingual Speech Synthesis Using Learned Phonetic Representations. IEEE\/ACM Trans. on Audio Speech and Lang. Process. 31 (2023) 3706\u20133716.","DOI":"10.1109\/TASLP.2023.3313424"},{"key":"e_1_3_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475198"},{"key":"e_1_3_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414670"},{"key":"e_1_3_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME51207.2021.9428313"},{"key":"e_1_3_1_27_1","doi-asserted-by":"crossref","unstructured":"Khalid Mahmood Malik Ali Javed Hafiz Malik and Aun Irtaza. 2020. A Light-Weight Replay Detection Framework for Voice Controlled IoT Devices. IEEE J. Selected Topics Signal Process. 14 5 (2020) 982\u2013996.","DOI":"10.1109\/JSTSP.2020.2999828"},{"key":"e_1_3_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICIINFS.2018.8721379"},{"key":"e_1_3_1_29_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-472"},{"key":"e_1_3_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7953088"},{"key":"e_1_3_1_31_1","doi-asserted-by":"crossref","unstructured":"Xiaochuan Sun Jingchang Fu Biao Wei Zhigang Li Yingqi Li and Ning Wang. 2023. A Self-Attentional ResNet-LightGBM Model for IoT-Enabled Voice Liveness Detection. IEEE Internet Things J. 10 9 (2023) 8257\u20138270.","DOI":"10.1109\/JIOT.2022.3230992"},{"key":"e_1_3_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054322"},{"key":"e_1_3_1_33_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-2249"},{"key":"e_1_3_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE55515.2023.00269"},{"key":"e_1_3_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00044"},{"key":"e_1_3_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN52387.2021.9533328"},{"key":"e_1_3_1_37_1","doi-asserted-by":"crossref","unstructured":"Jun Zhang Lei Pan Qing-Long Han Chao Chen Sheng Wen and Yang Xiang. 2022. Deep Learning Based Attack Detection for Cyber-Physical System Cybersecurity: A Survey. IEEE\/CAA J. Automatica Sinica 9 3 (2022) 377\u2013391.","DOI":"10.1109\/JAS.2021.1004261"},{"key":"e_1_3_1_38_1","doi-asserted-by":"crossref","unstructured":"Lin Zhang Xin Wang Erica Cooper Nicholas Evans and Junichi Yamagishi. 2023. The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance. IEEE\/ACM Trans. Audio Speech Lang. Process. 31 (2023) 813\u2013825.","DOI":"10.1109\/TASLP.2022.3233236"},{"key":"e_1_3_1_39_1","first-page":"1","volume-title":"Proceedings of the 2022 IEEE International Conference on Multimedia and Expo (ICME \u201922)","author":"Zhang Zhenyu","year":"2022","unstructured":"Zhenyu Zhang, Xianfeng Zhao, and Xiaowei Yi. 2022. Improving Robustness of Speech Anti-Spoofing System Using ResNeXt withNeighbor Filters. In Proceedings of the 2022 IEEE International Conference on Multimedia and Expo (ICME \u201922), 1\u20136."},{"key":"e_1_3_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSP58490.2023.10248504"},{"key":"e_1_3_1_41_1","doi-asserted-by":"publisher","unstructured":"Jinli Zhao Ziqi Zhang Hao Yu Haoran Ji Peng Li Wei Xi Jinyue Yan and Chengshan Wang. 2023. Cloud-Edge Collaboration-Based Local Voltage Control for DGs with Privacy Preservation. IEEE Trans. Industrial Inform. 19 1 (2023) 98\u2013108. DOI: 10.1109\/TII.2022.3172901","DOI":"10.1109\/TII.2022.3172901"},{"key":"e_1_3_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM48880.2022.9796782"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3698773","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,12]],"date-time":"2025-08-12T20:36:55Z","timestamp":1755031015000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698773"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,12]]},"references-count":41,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2025,8,31]]}},"alternative-id":["10.1145\/3698773"],"URL":"https:\/\/doi.org\/10.1145\/3698773","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,12]]},"assertion":[{"value":"2024-04-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-20","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}