{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T05:08:35Z","timestamp":1779340115198,"version":"3.51.4"},"reference-count":51,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2026,5,9]],"date-time":"2026-05-09T00:00:00Z","timestamp":1778284800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computation"],"abstract":"<jats:p>To ensure the reliable operation of speech systems across diverse environments, noise addition methods have emerged as the standard solution. However, existing methods offer limited coverage of real-world scenes and depend on pre-existing noise libraries and scene metadata. This paper presents prompt-based Dynamic Generative Scene-based Noise Addition (DGSNA), a novel approach driven by generative language models that integrates Dynamic Generation of Scene-based Information (DGSI) with Scene-based Noise Addition for Speech (SNAS). The DGSI module, with a BET (Background, Examples, Task) prompt framework, dynamically generates logic-compliant scene-based information, including scene dimensions, sound sources, and microphone positions, thereby addressing the challenges of scene enumeration and detailed description. Complementing this, the SNAS module employs a Time\u2013Frequency Diffusion-based (TFD) Text-to-Audio model to synthesize scene-specific noise. By integrating this noise with clean speech via Room Impulse Response (RIR) filters, the module streamlines the traditionally labor-intensive process of replicating diverse acoustic environments. Experimental results show that DGSNA significantly enhances the robustness of speech recognition and keyword spotting models, achieving relative improvements of up to 11.32%. Furthermore, DGSNA is highly compatible with existing noise addition techniques.<\/jats:p>","DOI":"10.3390\/computation14050109","type":"journal-article","created":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T10:50:52Z","timestamp":1778496652000},"page":"109","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["DGSNA: Dynamic Generative Scene-Based Noise Addition Method"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-5449-3145","authenticated-orcid":false,"given":"Zihao","family":"Chen","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Guangdong University of Technology, Guangzhou 510006, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1977-2249","authenticated-orcid":false,"given":"Zhentao","family":"Lin","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Guangdong University of Technology, Guangzhou 510006, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8596-8333","authenticated-orcid":false,"given":"Bi","family":"Zeng","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Guangdong University of Technology, Guangzhou 510006, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Linyi","family":"Huang","sequence":"additional","affiliation":[{"name":"China Electronic Product Reliability and Environmental, Testing Research Institute, Guangzhou 511370, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jia","family":"Cai","sequence":"additional","affiliation":[{"name":"China Electronic Product Reliability and Environmental, Testing Research Institute, Guangzhou 511370, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,5,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Koyama, S., De Sena, E., Samarasinghe, P., Thomas, M.R., and Antonacci, F. (2025). Past, Present, and Future of Spatial Audio and Room Acoustics. Proceedings of the ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP49660.2025.10890366"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"361","DOI":"10.1109\/TCDS.2025.3598687","article-title":"Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction","volume":"18","author":"Hao","year":"2025","journal-title":"IEEE Trans. Cogn. Dev. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Nguyen, T.B., Pham, N.Q., and Waibel, A. (2025). Cocktail-Party Audio-Visual Speech Recognition. arXiv.","DOI":"10.21437\/Interspeech.2025-676"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"557","DOI":"10.17743\/jaes.2021.0009","article-title":"Six-Degrees-of-Freedom Parametric Spatial Audio Based on One Monaural Room Impulse Response","volume":"69","author":"Arend","year":"2021","journal-title":"J. Audio Eng. Soc."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Koyama, Y., Shigemi, K., Takahashi, M., Shimada, K., Takahashi, N., Tsunoo, E., Takahashi, S., and Mitsufuji, Y. (2022). Spatial Data Augmentation with Simulated Room Impulse Responses for Sound Event Localization and Detection. Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP43922.2022.9746754"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Wang, Z., Liu, Z., Zhu, X., Zhu, Y., Liu, M., Chen, J., Xiao, L., Weng, C., and Xie, L. (2025). FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching. arXiv.","DOI":"10.21437\/Interspeech.2025-1745"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"61","DOI":"10.4316\/AECE.2018.03009","article-title":"A Proposal of a Novel Method for Generating Discrete Analog Uniform Noise","volume":"18","author":"Gazivoda","year":"2018","journal-title":"Adv. Electr. Comput. Eng."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Akter, S., Khalil, K., and Bayoumi, M. (2025). A High Performance and Efficient Method for Enhancing Randomness in Linear Feedback Shift Registers (LFSR). Proceedings of the 2025 IEEE International Symposium on Circuits and Systems (ISCAS), IEEE.","DOI":"10.1109\/ISCAS56072.2025.11044167"},{"key":"ref_9","unstructured":"Borji, A., and Lin, S. (2019). White Noise Analysis of Neural Networks. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Toki\u0107, D., and Juri\u0161i\u0107, D. (2022). High-Precision Fractional-Order Integrator for Generating Pink Noise from White Noise. Proceedings of the 2022 45th Jubilee International Convention on Information, Communication and Electronic Technology (MIPRO), IEEE.","DOI":"10.23919\/MIPRO55190.2022.9803506"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"127","DOI":"10.3847\/1538-3881\/ab037c","article-title":"An Easy Algorithm to Generate Colored Noise Sequences","volume":"157","author":"Xu","year":"2019","journal-title":"Astron. J."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Jin, Z., Qian, Y., Liang, X., and Geng, H. (2025). A Multi-View Fusion Approach for Enhancing Speech Signals via Short-Time Fractional Fourier Transform. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, IEEE.","DOI":"10.24963\/ijcai.2025\/613"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Tu, Z., Deadman, J., Ma, N., and Barker, J. (2022). Auditory-Based Data Augmentation for End-to-End Automatic Speech Recognition. Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP43922.2022.9746252"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Park, D.S., Chan, W., Zhang, Y., Chiu, C.C., Zoph, B., Cubuk, E.D., and Le, Q.V. (2019). SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. arXiv.","DOI":"10.21437\/Interspeech.2019-2680"},{"key":"ref_15","first-page":"197","article-title":"An Image-Source Method for Modelling Sound in Arbitrary Enclosed Spaces","volume":"11","author":"Dance","year":"1995","journal-title":"WIT Trans. Built Environ."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Scheibler, R., Bezzam, E., and Dokmani\u0107, I. (2018). Pyroomacoustics: A Python Package for Audio Room Simulation and Array Processing Algorithms. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP.2018.8461310"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1429","DOI":"10.1109\/TASL.2009.2035038","article-title":"Diffuse Reverberation Model for Efficient Image-Source Simulation of Room Impulse Responses","volume":"18","author":"Lehmann","year":"2009","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Tang, Z., Chen, L., Wu, B., Yu, D., and Manocha, D. (2020). Improving Reverberant Speech Training Using Diffuse Acoustic Simulation. Proceedings of the ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP40776.2020.9052932"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Chen, J., Chen, J., Chen, H., Wang, Q., Gao, Y., and Du, J. (2025). MEAN-RIR: Multi-Modal Environment-Aware Network for Robust Room Impulse Response Estimation. arXiv.","DOI":"10.1109\/ASRU65441.2025.11434671"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"53","DOI":"10.18196\/jrc.v6i1.24160","article-title":"A Review on Comparative Analysis of Generative Adversarial Networks\u2019 Architectures and Applications","volume":"6","author":"Bhat","year":"2025","journal-title":"J. Robot. Control (JRC)"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Ratnarajah, A., Zhang, S.X., Yu, M., Tang, Z., Manocha, D., and Yu, D. (2022). FAST-RIR: Fast Neural Diffuse Room Impulse Response Generator. Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP43922.2022.9747846"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Barakat, B., and Jaf, S. (2025). Beyond Traditional Classifiers: Evaluating Large Language Models for Robust Hate Speech Detection. Computation, 13.","DOI":"10.3390\/computation13080196"},{"key":"ref_23","first-page":"23802","article-title":"AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head","volume":"38","author":"Huang","year":"2024","journal-title":"Proc. Aaai Conf. Artif. Intell."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Ashqar, H.I., Jaber, A., Alhadidi, T.I., and Elhenawy, M. (2025). Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing. Computation, 13.","DOI":"10.3390\/computation13060133"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Kong, Q., Xu, Y., Iqbal, T., Cao, Y., Wang, W., and Plumbley, M.D. (2019). Acoustic Scene Generation with Conditional SampleRNN. Proceedings of the ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP.2019.8683727"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, X., Iqbal, T., Zhao, J., Huang, Q., Plumbley, M.D., and Wang, W. (2021). Conditional Sound Generation Using Neural Discrete Time-Frequency Representation Learning. arXiv.","DOI":"10.1109\/MLSP52302.2021.9596430"},{"key":"ref_27","first-page":"6840","article-title":"Denoising Diffusion Probabilistic Models","volume":"33","author":"Ho","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"120784","DOI":"10.1016\/j.actamat.2025.120784","article-title":"GrainPaint: A Multi-Scale Diffusion-Based Generative Model for Microstructure Reconstruction of Large-Scale Objects","volume":"288","author":"Hoffman","year":"2025","journal-title":"Acta Mater."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1720","DOI":"10.1109\/TASLP.2023.3268730","article-title":"Diffsound: Discrete Diffusion Model for Text-to-Sound Generation","volume":"31","author":"Yang","year":"2023","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_30","unstructured":"Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., D\u00e9fossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y. (2022). AudioGen: Textually Guided Audio Generation. arXiv."},{"key":"ref_31","unstructured":"Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., Wang, W., and Plumbley, M.D. (2023). AudioLDM: Text-to-Audio Generation with Latent Diffusion Models. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Ghosal, D., Majumder, N., Mehrish, A., and Poria, S. (2023). Text-to-Audio Generation Using Instruction Guided Latid Diffusion Model. Proceedings of the 31st ACM International Conference on Multimedia, IEEE.","DOI":"10.1145\/3581783.3612348"},{"key":"ref_33","first-page":"13916","article-title":"Make-An-Audio: Text-to-Audio Generation with Prompt-Enhanced Diffusion Models","volume":"202","author":"Huang","year":"2023","journal-title":"Proc. Int. Conf. Mach. Learn."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S. (2023). Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation. Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP49357.2023.10095969"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Li, L. (2026). Variational Autoencoder (VAE). Artificial Intelligence for Drug Design, Springer.","DOI":"10.1007\/978-981-95-2525-6_12"},{"key":"ref_36","first-page":"17022","article-title":"HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis","volume":"33","author":"Kong","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Bu, H., Du, J., Na, X., Wu, B., and Zheng, H. (2017). Aishell-1: An Open-Source Mandarin Speech Corpus and a Speech Recognition Baseline. Proceedings of the 2017 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I\/O Systems and Assessment (O-COCOSDA), IEEE.","DOI":"10.1109\/ICSDA.2017.8384449"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Zhang, B., Lv, H., Guo, P., Shao, Q., Yang, C., Xie, L., Xu, X., Bu, H., Chen, X., and Zeng, C. (2022). Wenetspeech: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech Recognition. Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP43922.2022.9746682"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Leroy, D., Coucke, A., Lavril, T., Gisselbrecht, T., and Dureau, J. (2019). Federated Learning for Keyword Spotting. Proceedings of the ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP.2019.8683546"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Gulati, A., Qin, J., Chiu, C.C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., and Wu, Y. (2020). Conformer: Convolution-Augmented Transformer for Speech Recognition. arXiv.","DOI":"10.21437\/Interspeech.2020-3015"},{"key":"ref_41","first-page":"4054","article-title":"WeNet: Production Oriented Streaming and Non-Streaming End-to-End Speech Recognition Toolkit","volume":"21","author":"Yao","year":"2021","journal-title":"Proc. Interspeech"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"28","DOI":"10.1016\/j.neunet.2022.03.003","article-title":"Two-Stage Streaming Keyword Detection and Localization with Multi-Scale Depthwise Temporal Convolution","volume":"150","author":"Hou","year":"2022","journal-title":"Neural Netw."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Wang, J., Xu, M., Hou, J., Zhang, B., Zhang, X.L., Xie, L., and Pan, F. (2023). Wekws: A Production First Small-Footprint End-to-End Keyword Spotting Toolkit. Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP49357.2023.10096736"},{"key":"ref_44","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention Is All You Need. Adv. Neural Inf. Process. Syst., 30."},{"key":"ref_45","first-page":"62991","article-title":"C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models","volume":"36","author":"Huang","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_46","unstructured":"Xiong, B., Chen, B., Wang, C., Luo, D., Xu, D., Liu, D., Yang, F., Li, F., Teng, F., and Wang, F. (2025). BlueLM-2.5-3B Technical Report. arXiv."},{"key":"ref_47","unstructured":"Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Wang, G., Li, H., Zhu, J., and Chen, J. (2024). Yi: Open Foundation Models by 01.ai. arXiv."},{"key":"ref_48","unstructured":"Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., and Xia, X. (2022). GLM-130B: An Open Bilingual Pre-Trained Model. arXiv."},{"key":"ref_49","unstructured":"Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., and Huang, F. (2023). Qwen Technical Report. arXiv."},{"key":"ref_50","unstructured":"Song, J., Meng, C., and Ermon, S. (2020). Denoising Diffusion Implicit Models. arXiv."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Hwang, J., Hira, M., Chen, C., Zhang, X., Ni, Z., Sun, G., Ma, P., Huang, R., Pratap, V., and Zhang, Y. (2023). TorchAudio 2.1: Advancing Speech Recognition, Self-Supervised Learning, and Audio Processing Components for PyTorch. Proceedings of the 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE.","DOI":"10.1109\/ASRU57964.2023.10389648"}],"container-title":["Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2079-3197\/14\/5\/109\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T04:28:31Z","timestamp":1779337711000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2079-3197\/14\/5\/109"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,9]]},"references-count":51,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2026,5]]}},"alternative-id":["computation14050109"],"URL":"https:\/\/doi.org\/10.3390\/computation14050109","relation":{},"ISSN":["2079-3197"],"issn-type":[{"value":"2079-3197","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,9]]}}}