{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T21:06:13Z","timestamp":1777583173900,"version":"3.51.4"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2021,7,21]],"date-time":"2021-07-21T00:00:00Z","timestamp":1626825600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2021,9,30]]},"abstract":"<jats:p>This article studies noised Asian speech enhancement based on the deep neural network (DNN) and its implementation on an app. We use the THCHS-30 speech dataset and the common noise dataset in daily life as training and testing data of the DNN. To stack the frequency data of multiple audio frames to improve the effect of speech enhancement, the system compares the best number of stacked frames during training and testing. At the same time, the influence of training rounds on the PESQ is compared, and the best number of rounds is obtained. On this basis, the best model is implemented on the hearing aid app, and the real-time performance of the device is tested. The experiment shows that based on the DNN, using an appropriate number of rounds for training and using an appropriate number of audio frames stacking to improve the speech enhancement effect, and transplanting this speech enhancement model to the hearing aid app, can effectively improve speech clarity and intelligibility within a reasonable time delay range.<\/jats:p>","DOI":"10.1145\/3439797","type":"journal-article","created":{"date-parts":[[2021,7,21]],"date-time":"2021-07-21T21:12:38Z","timestamp":1626901958000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Deep Neural Network Based Noised Asian Speech Enhancement and Its Implementation on a Hearing Aid App"],"prefix":"10.1145","volume":"20","author":[{"given":"Xiaoqian","family":"Fan","sequence":"first","affiliation":[{"name":"Zhejiang University, Hangzhou, Zhejiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bowen","family":"Yang","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, Zhejiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenzhi","family":"Chen","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, Zhejiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Quanfang","family":"Fan","sequence":"additional","affiliation":[{"name":"Hangzhou Youting Technology Co., Ltd., Hangzhou, Zhejiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,7,21]]},"reference":[{"key":"e_1_2_1_1_1","article-title":"The development and prospect of hearing aids","author":"Ma Xiaoling","year":"2014","unstructured":"Xiaoling Ma , Xun Liu , Sixing Zhang , Mengkang Zhang , Xiushan Cao , Liuming Tian , and Wen Gao . 2014 . The development and prospect of hearing aids . Journal of Minzu University of China (Natural Sciences Edition). Xiaoling Ma, Xun Liu, Sixing Zhang, Mengkang Zhang, Xiushan Cao, Liuming Tian, and Wen Gao. 2014. The development and prospect of hearing aids. Journal of Minzu University of China (Natural Sciences Edition).","journal-title":"Journal of Minzu University of China (Natural Sciences Edition)."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2001.941023"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1979.1163209"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1978.1163086"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2005.860851"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1984.1164453"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0165-1684(01)00128-1"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2003.811544"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1989.266851"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the International Conference on Acoustics, Speech, and Signal Processing. 53\u201356","author":"Xie F.","unstructured":"F. Xie and D. V. Compernolle . 1994. A family of MLP based nonlinear spectral estimators for noise reduction . In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing. 53\u201356 . F. Xie and D. V. Compernolle. 1994. A family of MLP based nonlinear spectral estimators for noise reduction. In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing. 53\u201356."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.2006.18.7.1527"},{"key":"e_1_2_1_12_1","doi-asserted-by":"crossref","unstructured":"G. E. Hinton and R. R. Salakhutdinov. 2006. Reducing the dimensionality of data with neural networks. Science 313 5786 (2006) 504\u2013507.  G. E. Hinton and R. R. Salakhutdinov. 2006. Reducing the dimensionality of data with neural networks. Science 313 5786 (2006) 504\u2013507.","DOI":"10.1126\/science.1127647"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the 2013 INTERSPEECH Conference. 3444\u20133448","author":"Xia B.-Y.","year":"2013","unstructured":"B.-Y. Xia and C.-C. Bao . 2013 . Speech enhancement with weighted denoising auto-encoder . In Proceedings of the 2013 INTERSPEECH Conference. 3444\u20133448 . B.-Y. Xia and C.-C. Bao. 2013. Speech enhancement with weighted denoising auto-encoder. In Proceedings of the 2013 INTERSPEECH Conference. 3444\u20133448."},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the 2013 INTERSPEECH Conference. 436\u2013440","author":"Lu X.-G.","unstructured":"X.-G. Lu , Y. Tsao , S. Matsuda , and C. Hori . 2013. Speech enhancement based on deep denoising autoencoder . In Proceedings of the 2013 INTERSPEECH Conference. 436\u2013440 . X.-G. Lu, Y. Tsao, S. Matsuda, and C. Hori. 2013. Speech enhancement based on deep denoising autoencoder. In Proceedings of the 2013 INTERSPEECH Conference. 436\u2013440."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2014.2364452"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/1577069.1577070"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1561\/2200000006"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1989.266851"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.2006.18.7.1527"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2013.2250961"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2014.2352935"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-22482-4_11"},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Serim Park and Jin Lee. 2017. A fully convolutional neural network for speech enhancement. arXiv:1609.07132.  Serim Park and Jin Lee. 2017. A fully convolutional neural network for speech enhancement. arXiv:1609.07132.","DOI":"10.21437\/Interspeech.2017-1465"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2013.2291240"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2015.2512042"},{"key":"e_1_2_1_26_1","doi-asserted-by":"crossref","unstructured":"Jingxian Tu and Youshen Xia. 2018. Effective Kalman filtering algorithm for distributed multichannel speech enhancement. Neurocomputing 275 (2018) 144\u2013154.  Jingxian Tu and Youshen Xia. 2018. Effective Kalman filtering algorithm for distributed multichannel speech enhancement. Neurocomputing 275 (2018) 144\u2013154.","DOI":"10.1016\/j.neucom.2017.05.048"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10772-017-9406-4"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/3045118.3045303"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001163"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/2959355.2959378"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 2007 IEEE International Symposium on Circuits and Systems. IEEE","author":"Yermeche Z.","unstructured":"Z. Yermeche , B. Sallberg , N. Grbic , and I. Claesson . 2007. Real-time DSP implementation of a subband beamforming algorithm for dual microphone speech enhancement . In Proceedings of the 2007 IEEE International Symposium on Circuits and Systems. IEEE , Los Alamitos, CA. Z. Yermeche, B. Sallberg, N. Grbic, and I. Claesson. 2007. Real-time DSP implementation of a subband beamforming algorithm for dual microphone speech enhancement. In Proceedings of the 2007 IEEE International Symposium on Circuits and Systems. IEEE, Los Alamitos, CA."},{"key":"e_1_2_1_32_1","doi-asserted-by":"crossref","unstructured":"J. M. Valin. 2017. A hybrid DSP\/deep learning approach to real-time full-band speech enhancement. arXiv:1709.08243.  J. M. Valin. 2017. A hybrid DSP\/deep learning approach to real-time full-band speech enhancement. arXiv:1709.08243.","DOI":"10.1109\/MMSP.2018.8547084"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.measurement.2020.107790"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3439797","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3439797","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:01:52Z","timestamp":1750197712000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3439797"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,21]]},"references-count":33,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2021,9,30]]}},"alternative-id":["10.1145\/3439797"],"URL":"https:\/\/doi.org\/10.1145\/3439797","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,7,21]]},"assertion":[{"value":"2020-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-07-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}