{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,17]],"date-time":"2026-02-17T12:12:59Z","timestamp":1771330379436,"version":"3.50.1"},"reference-count":25,"publisher":"Engineering and Technology Publishing","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["jcm"],"published-print":{"date-parts":[[2020]]},"abstract":"<jats:p>Speech separation plays an important role in a speech-related system because it can denoise, extract and enhance speech signal, and after all improve the accuracy and performance of the system. In recent years, many approaches only separate the speech out of commonly high-frequency noise or a particular background sound. We propose a more powerful approach, combining an autoencoder and a bandpass filter to separate speech signals. This combination can extract the speech in the mixture with not only high-frequency noise but also many kinds of different background sounds. Our approach can be flexibly applied for the new background sounds. Experimental results show that our model can extract fastly and effectively the speech signal with 9.01 dB in SIR and 11.26 in SDR. On the other hand, we can adjust the passband to identify the range of frequency at the output signal to apply for particular applications.<\/jats:p>","DOI":"10.12720\/jcm.15.11.841-848","type":"journal-article","created":{"date-parts":[[2020,12,29]],"date-time":"2020-12-29T07:30:16Z","timestamp":1609227016000},"page":"841-848","source":"Crossref","is-referenced-by-count":12,"title":["Speech Separation in the Frequency Domain with Autoencoder"],"prefix":"10.12720","author":[{"name":"University of Science, Ho Chi Minh City, Vietnam","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao D.","family":"Do","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Son T.","family":"Tran","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Duc T.","family":"Chau","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"4977","published-online":{"date-parts":[[2020]]},"reference":[{"key":"ref0","doi-asserted-by":"publisher","unstructured":"[1] N. Yang, M. Usman, X. J. He, M. A. Jan, and L. M. Zhang, \"Time-frequency filter bank: A simple approach for audio and music separation,\" IEEE Access, vol. 5, pp. 27114 - 27125, 2017.","DOI":"10.1109\/ACCESS.2017.2761741"},{"key":"ref1","doi-asserted-by":"publisher","unstructured":"[2] Y. Xie, K. Xie, Z. Wu, and S. Xie, \"Underdetermined blind source separation of speech mixtures based on K-means clustering,\" in Proc. Chinese Control Conference (CCC), Guangzhou, China, 2019, pp. 42-46.","DOI":"10.23919\/ChiCC.2019.8865385"},{"key":"ref2","doi-asserted-by":"publisher","unstructured":"[3] B. Peng, W. Liu, and D. P. Mandic, \"Design of oversampled generalised discrete Fourier transform filter banks for application to subband based blind source separation,\" IET Signal Process., vol. 7, no. 9, pp. 843-853, 2013.","DOI":"10.1049\/iet-spr.2012.0361"},{"key":"ref3","doi-asserted-by":"publisher","unstructured":"[4] C. Osterwise and S. L. Grant, \"On over-determined frequency domain BSS,\" IEEE\/ACM Trans. Audio, Speech, Language Process, vol. 22, no. 5, pp. 956-966, May 2014.","DOI":"10.1109\/TASLP.2014.2307166"},{"key":"ref4","doi-asserted-by":"publisher","unstructured":"[5] S. H. Sardouie, M. B. Shamsollahi, L. Albera, and I. Merlet, \"Denoising of ictal EEG data using semi-blind source separation methods based on time-frequency priors,\" IEEE J. Biomed. Health Inform., vol. 19, no. 3, pp. 839-847, May 2015.","DOI":"10.1109\/JBHI.2014.2336797"},{"key":"ref5","doi-asserted-by":"publisher","unstructured":"[6] B. Rivet, \"Source separation of multimodal data: A second-order approach based on a constrained joint block decomposition of covariance matrices,\" IEEE Signal Process. Lett., vol. 22, no. 6, pp. 681-685, June 2015.","DOI":"10.1109\/LSP.2014.2367158"},{"key":"ref6","doi-asserted-by":"publisher","unstructured":"[7] S. Lee and H. S. Pang, \"Multichannel non-negative matrix factorisation based on alternating least squares for audio source separation system,\" Electron. Lett., vol. 51, no. 3, pp. 197-198, 2015.","DOI":"10.1049\/el.2014.2616"},{"key":"ref7","doi-asserted-by":"publisher","unstructured":"[8] G. S. Fu, R. Phlypo, M. Anderson, X. L. Li, and T. Adal\u0131, \"Blind source separation by entropy rate minimization,\" IEEE Trans. Signal Process, vol. 62, no. 16, pp. 4245-4255, Aug. 2014.","DOI":"10.1109\/TSP.2014.2333563"},{"key":"ref8","doi-asserted-by":"publisher","unstructured":"[9] J. Hofmanis, O. Caspary, V. Louis-Dorr, R. Ranta, and L. Maillard, \"Denoising depth EEG signals during DBS using filtering and subspace decomposition,\" IEEE Trans. Biomed. Eng., vol. 60, no. 10, pp. 2686-2695, Oct. 2013.","DOI":"10.1109\/TBME.2013.2262212"},{"key":"ref9","doi-asserted-by":"publisher","unstructured":"[10] O. Tich\u00fd and V. \u0160m\u00eddl, \"Bayesian blind separation and deconvolution of dynamic image sequences using sparsity priors,\" IEEE Trans. Med. Imag., vol. 34, no. 1, pp. 258-266, Jan. 2015.","DOI":"10.1109\/TMI.2014.2352791"},{"key":"ref10","doi-asserted-by":"publisher","unstructured":"[11] J. Nikunen and T. Virtanen, \"Direction of arrival based spatial covariance model for blind sound source separation,\" IEEE\/ACM Trans. Audio, Speech, Language Process, vol. 22, no. 3, pp. 727-739, Mar. 2014.","DOI":"10.1109\/TASLP.2014.2303576"},{"key":"ref11","doi-asserted-by":"publisher","unstructured":"[12] B. Liu, V. G. Reju, A. W. H. Khong, and V. V. Reddy, \"A GMM post-filter for residual crosstalk suppression in blind source separation,\" IEEE Signal Process. Lett., vol. 21, no. 8, pp. 942-946, Aug. 2014.","DOI":"10.1109\/LSP.2014.2317761"},{"key":"ref12","doi-asserted-by":"publisher","unstructured":"[13] B. Liu, V. G. Reju, and A. W. H. Khong, \"A linear source recovery method for underdetermined mixtures of uncorrelated AR-model signals without sparseness,\" IEEE Trans. Signal Process, vol. 62, no. 19, pp. 4947-4958, Oct. 2014.","DOI":"10.1109\/TSP.2014.2329646"},{"key":"ref13","doi-asserted-by":"publisher","unstructured":"[14] S. Hosseini and Y. Deville, \"Blind separation of parametric nonlinear mixtures of possibly auto correlated and non-stationary sources,\" IEEE Trans. Signal Process, vol. 62, no. 24, pp. 6521-6533, Dec. 2014.","DOI":"10.1109\/TSP.2014.2367474"},{"key":"ref14","doi-asserted-by":"publisher","unstructured":"[15] Y. Zhang, P. Candra, G. Wang, and T. Xia, \"2-D entropy and short-time Fourier transform to leverage GPR data analysis efficiency,\" IEEE Trans. Instrum. Meas., vol. 64, no. 1, pp. 103-111, Jan. 2015.","DOI":"10.1109\/TIM.2014.2331429"},{"key":"ref15","doi-asserted-by":"publisher","unstructured":"[16] G. Okopal, S. Wisdom, and L. Atlas, \"Speech analysis with the strong uncorrelating transform,\" IEEE\/ACM Trans. Audio, Speech, Language Process, vol. 23, no. 11, pp. 1858-1868, Nov. 2015.","DOI":"10.1109\/TASLP.2015.2456426"},{"key":"ref16","doi-asserted-by":"publisher","unstructured":"[17] J. L. Roux and E. Vincent, \"Consistent wiener filtering for audio source separation,\" IEEE Signal Process. Lett., vol. 20, no. 3, pp. 217-220, Mar. 2013.","DOI":"10.1109\/LSP.2012.2225617"},{"key":"ref17","doi-asserted-by":"publisher","unstructured":"[18] Y. G. Jin, J. W. Shin, and N. S. Kim, \"Spectro-temporal filtering for multichannel speech enhancement in short-time Fourier transform domain,\" IEEE Signal Process. Lett., vol. 21, no. 3, pp. 352-355, Mar. 2014.","DOI":"10.1109\/LSP.2014.2302897"},{"key":"ref18","doi-asserted-by":"publisher","unstructured":"[19] R. E. Turner and M. Sahani, \"Time-frequency analysis as probabilistic inference,\" IEEE Trans. Signal Process, vol. 62, no. 23, pp. 6171-6183, Dec. 2014.","DOI":"10.1109\/TSP.2014.2362100"},{"key":"ref19","doi-asserted-by":"publisher","unstructured":"[20] L. Stankovic, S. Stankovic, and M. Dakovic, \"From the STFT to the Wigner distribution [lecture notes],\" IEEE Signal Process. Mag., vol. 31, no. 3, pp. 163-174, May 2014.","DOI":"10.1109\/MSP.2014.2301791"},{"key":"ref20","doi-asserted-by":"publisher","unstructured":"[21] V. K. Mai, D. Pastor, A. A\u00efssa-El-Bey, and R. Le-Bidan, \"Robust estimation of non-stationary noise power spectrum for speech enhancement,\" IEEE\/ACM Trans. Audio, Speech, Language Process, vol. 23, no. 4, pp. 670-682, Apr. 2015.","DOI":"10.1109\/TASLP.2015.2401426"},{"key":"ref21","doi-asserted-by":"publisher","unstructured":"[22] P. Flandrin, \"Time-frequency filtering based on spectrogram zeros,\" IEEE Signal Process. Lett., vol. 22, no. 11, pp. 2137-2141, Nov. 2015.","DOI":"10.1109\/LSP.2015.2463093"},{"key":"ref22","unstructured":"[23] R. B. Blackman and J. W. Tukey, The Measurement of Power Spectra from the Point of View of Communications Engineering, Dover Publications Publishing House, 1959."},{"key":"ref23","unstructured":"[24] T. F. Quatieri, Discrete-time Speech Signal Processing: Principles and Practice, Prentice Hall Publishing House, 2001."},{"key":"ref24","doi-asserted-by":"publisher","unstructured":"[25] E. Vincent, R. Gribonval, and C. F\u00e9votte, \"Performance measurement in blind audio source separation,\" IEEE Transactions on Audio, Speech and Language Processing, Institute of Electrical and Electronics Engineers, vol. 14, no. 4, pp. 1462-1469, 2006","DOI":"10.1109\/TSA.2005.858005"}],"container-title":["Journal of Communications"],"original-title":[],"link":[{"URL":"http:\/\/www.jocm.us\/uploadfile\/2020\/1013\/20201013060720285.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,11,24]],"date-time":"2021-11-24T07:51:31Z","timestamp":1637740291000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.jocm.us\/show-246-1613-1.html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020]]},"references-count":25,"URL":"https:\/\/doi.org\/10.12720\/jcm.15.11.841-848","relation":{},"ISSN":["1796-2021"],"issn-type":[{"value":"1796-2021","type":"print"}],"subject":[],"published":{"date-parts":[[2020]]}}}