{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T11:17:20Z","timestamp":1771240640491,"version":"3.50.1"},"reference-count":33,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2026,2,14]],"date-time":"2026-02-14T00:00:00Z","timestamp":1771027200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Henan Provincial Department of Public Security","award":["262102320015"],"award-info":[{"award-number":["262102320015"]}]},{"name":"Henan Police College Scientific Research Projects Achievements","award":["HNJY-2025-09"],"award-info":[{"award-number":["HNJY-2025-09"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Speech-driven face generation aims to synthesize a face image that matches a speaker\u2019s identity from speech alone. However, existing methods typically trade identity fidelity for visual quality and rely on large end-to-end generators that are difficult to train and tune. We propose Vox2Face, a speech-driven face generation framework centered on an explicit identity space rather than direct speech-to-image mapping. A pretrained speaker encoder first extracts speech embeddings, which are distilled and metric-aligned to the ArcFace hyperspherical identity space, transforming cross-modal regression into a geometrically interpretable speech-to-identity alignment problem. On this unified identity representation, we reused an identity-conditioned diffusion model as the generative backbone and synthesized diverse, high-resolution faces in the Stable Diffusion latent space. To better exploit this prior, we introduce a discriminator-free diffusion self-consistency loss that treats denoising residuals as an implicit critique of speech-predicted identity embeddings and updates only the speech-to-identity mapping and lightweight LoRA adapters, encouraging speech-derived identities to lie on the high-probability identity manifold of the diffusion model. Experiments on the HQ-VoxCeleb dataset show that Vox2Face improves the ArcFace cosine similarity from 0.295 to 0.322, boosts R@10 retrieval accuracy from 29.8% to 32.1%, and raises the VGGFace Score from 18.82 to 23.21 over a strong diffusion baseline. These results indicate that aligning speech to a unified identity space and reusing a strong identity-conditioned diffusion prior is an effective method to jointly improve identity fidelity and visual quality.<\/jats:p>","DOI":"10.3390\/info17020200","type":"journal-article","created":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T10:14:32Z","timestamp":1771236872000},"page":"200","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Vox2Face: Speech-Driven Face Generation via Identity-Space Alignment and Diffusion Self-Consistency"],"prefix":"10.3390","volume":"17","author":[{"given":"Qiming","family":"Ma","sequence":"first","affiliation":[{"name":"School of Information Network Security, People\u2019s Public Security University of China, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yizhen","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information Network Security, People\u2019s Public Security University of China, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiang","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Information Network Security, People\u2019s Public Security University of China, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiadi","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Information Network Security, People\u2019s Public Security University of China, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gang","family":"Cheng","sequence":"additional","affiliation":[{"name":"School of Information Network Security, People\u2019s Public Security University of China, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jia","family":"Feng","sequence":"additional","affiliation":[{"name":"School of Traffic Management, People\u2019s Public Security University of China, Beijing 100038, China"},{"name":"Department of Traffic Management Engineering, Henan Police College, Zhengzhou 450046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rong","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information Network Security, People\u2019s Public Security University of China, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fanliang","family":"Bu","sequence":"additional","affiliation":[{"name":"School of Information Network Security, People\u2019s Public Security University of China, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,2,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1709","DOI":"10.1016\/j.cub.2003.09.005","article-title":"\u2018Putting the Face to the Voice\u2019: Matching Identity across Modality","volume":"13","author":"Kamachi","year":"2003","journal-title":"Curr. Biol."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Nagrani, A., Albanie, S., and Zisserman, A. (2018). Seeing voices and hearing faces: Cross-modal biometric matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2018.00879"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Wang, J., Liu, L., and Wang, J. (2025). Fine-portraitist: Visualizing the Speaker\u2019s Face Portrait during Speech Listening. Proceedings of the ICASSP 2025\u20142025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP49660.2025.10889904"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Jawahar, C.V., Li, H., Mori, G., and Schindler, K. (2019). On Learning Associations of Faces and Voices. Computer Vision\u2013ACCV 2018, Springer International Publishing.","DOI":"10.1007\/978-3-030-20873-8"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhang, L., Rao, A., and Agrawala, M. (2023). Adding Conditional Control to Text-to-Image Diffusion Models. Proceedings of the 2023 IEEE\/CVF International Conference on Computer Vision (ICCV), IEEE.","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"ref_7","unstructured":"Ye, H., Zhang, J., Liu, S., Han, X., and Yang, W. (2023). IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Papantoniou, F.P., Lattas, A., Moschoglou, S., Deng, J., Kainz, B., and Zafeiriou, S. (2024). Arc2face: A Foundation Model for Id-Consistent Human Faces, Springer. European Conference on Computer Vision.","DOI":"10.1007\/978-3-031-72913-3_14"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Oh, T.-H., Dekel, T., Kim, C., Mosseri, I., Freeman, W.T., Rubinstein, M., and Matusik, W. (2019). Speech2face: Learning the face behind a voice. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2019.00772"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lin, W.-W., and Mak, M.-W. (2020, January 25\u201329). Wav2Spk: A Simple DNN Architecture for Learning Speaker Embeddings from Waveforms. Proceedings of the Interspeech 2020, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-1287"},{"key":"ref_11","unstructured":"Wen, Y., Raj, B., and Singh, R. (2019). Face reconstruction from voice using generative adversarial networks. Proceedings of the 33rd International Conference on Neural Information Processing Systems, Curran Associates Inc."},{"key":"ref_12","unstructured":"Wang, J., Hu, X., Liu, L., Liu, W., Yu, M., and Xu, T. (2020). Attention-based residual speech portrait model for speech to face generation. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2892","DOI":"10.1016\/j.procs.2024.09.381","article-title":"Face From Voice: PyTorch Adaptation of Speech2Face Framework","volume":"246","author":"Pietkun","year":"2024","journal-title":"Procedia Comput. Sci."},{"key":"ref_14","unstructured":"Jalalifar, S.A., Hasani, H., and Aghajan, H. (2018). Speech-driven facial reenactment using conditional generative adversarial networks. arXiv."},{"key":"ref_15","unstructured":"Wang, J., Liu, L., Wang, J., and Cheng, H.V. (2023). Realistic Speech-to-Face Generation with Speech-Conditioned Latent Diffusion Model with Face Prior. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Chen, W., Xu, K., Dou, Y., and Gao, T. (2024). Voice-to-Face Generation: Couple of Self-Supervised Representation Learning with Diffusion Model. Proceedings of the 2024 IEEE International Conference on Multimedia and Expo (ICME), IEEE.","DOI":"10.1109\/ICME57554.2024.10687417"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"27262","DOI":"10.1038\/s41598-024-76407-9","article-title":"Scalable multimodal approach for face generation and super-resolution using a conditional diffusion model","volume":"14","author":"Abotaleb","year":"2024","journal-title":"Sci. Rep."},{"key":"ref_18","unstructured":"Kim, S.-B., Senocak, A., Ha, H., and Oh, T.-H. (2024). Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Arandjelovi\u0107, R., and Zisserman, A. (2017). Look, Listen and Learn. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), IEEE.","DOI":"10.1109\/ICCV.2017.73"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Aytar, Y., Vondrick, C., and Torralba, A. (2016). SoundNet: Learning Sound Representations from Unlabeled Video. arXiv.","DOI":"10.1109\/CVPR.2016.18"},{"key":"ref_21","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning Transferable Visual Models From Natural Language Supervision. Proceedings of the International Conference on Machine Learning 2021, Online."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"5962","DOI":"10.1109\/TPAMI.2021.3087709","article-title":"ArcFace: Additive Angular Margin Loss for Deep Face Recognition","volume":"44","author":"Deng","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_23","unstructured":"Ho, J., Jain, A., and Abbeel, P. (2020). Denoising Diffusion Probabilistic Models. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E.L., Ghasemipour, S.K.S., Ayan, B.K., Mahdavi, S.S., and Lopes, R.G. (2022). Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. arXiv.","DOI":"10.1145\/3528233.3530757"},{"key":"ref_25","unstructured":"Wang, Q., Bai, X., Wang, H., Qin, Z., and Chen, A. (2024). InstantID: Zero-shot Identity-Preserving Generation in Seconds. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Tomaevi\u0107, D., Boutros, F., Lin, C., Damer, N., \u0160truc, V., and Peer, P. (2025). ID-Booth: Identity-consistent Face Generation with Diffusion Models. Proceedings of the 2025 IEEE 19th International Conference on Automatic Face and Gesture Recognition (FG), IEEE.","DOI":"10.1109\/FG61629.2025.11099217"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Li, Z., Cao, M., Wang, X., Qi, Z., Cheng, M.-M., and Shan, Y. (2024). PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding. Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE.","DOI":"10.1109\/CVPR52733.2024.00825"},{"key":"ref_28","unstructured":"Poole, B., Jain, A., Barron, J.T., and Mildenhall, B. (2022). DreamFusion: Text-to-3D using 2D Diffusion. arXiv."},{"key":"ref_29","unstructured":"Hong, S., Ahn, D., and Kim, S. (2023). Debiasing Scores and Prompts of 2D Diffusion for View-consistent Text-to-3D Generation. arXiv."},{"key":"ref_30","unstructured":"Oord, A.V.D., Li, Y., and Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv."},{"key":"ref_31","unstructured":"Bai, Y., Ma, T., Wang, L., and Zhang, Z. (2020, January 10\u201311). Speech Fusion to Face: Bridging the Gap Between Human\u2019s Vocal Characteristics and Facial Imaging. Proceedings of the 30th ACM International Conference on Multimedia 2020, Istanbul, Turkey."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Gong, Y., Chung, Y.-A., and Glass, J.R. (2021). AST: Audio Spectrogram Transformer. arXiv.","DOI":"10.21437\/Interspeech.2021-698"},{"key":"ref_33","unstructured":"Hugging Face (2026, January 23). Diffusers: State-of-the-Art Diffusion Models for Image and Audio Generation in PyTorch. Available online: https:\/\/github.com\/huggingface\/diffusers."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/17\/2\/200\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T10:34:51Z","timestamp":1771238091000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/17\/2\/200"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,14]]},"references-count":33,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2026,2]]}},"alternative-id":["info17020200"],"URL":"https:\/\/doi.org\/10.3390\/info17020200","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,14]]}}}