{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:10:10Z","timestamp":1750183810249,"version":"3.41.0"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"8","license":[{"start":{"date-parts":[[2023,8,23]],"date-time":"2023-08-23T00:00:00Z","timestamp":1692748800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100019258","name":"Vingroup JSC","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100019258","id-type":"DOI","asserted-by":"crossref"}]},{"name":"PhD Scholarship Programme of Vingroup Innovation Foundation (VINIF), Institute of Big Data","award":["VINIF.2022.TS.037"],"award-info":[{"award-number":["VINIF.2022.TS.037"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,8,31]]},"abstract":"<jats:p>Time-frequency analysis (TFA) is a powerful method to exploit the hidden information of signals, including speech signals. Many techniques in this group were invented and developed to capture the most crucial stationary feature. However, human speech is not stable, and it contains some non-stationary elements. This work aims to design a new algorithm via the TFA technique to extract the trends and changes inside the speech signal in the time-frequency (TF) plane. We design a new algorithm to create a set of atoms for the signal transform, which can analyze the signal in many different view directions via Poly-Linear Chirplet Transform (PLCT). After processing the signal, the proposed method returns a multichannel output in which each channel results from a particular Linear Chirplet Transform (LCT). The feature then is combined with the MFCC feature to form the final representation. Although the size for speech representation rises, our extracted feature contains rich-meaning information to improve the recognition results compared to other features in gender recognition, dialect recognition, and speaker recognition.<\/jats:p>","DOI":"10.1145\/3605549","type":"journal-article","created":{"date-parts":[[2023,8,23]],"date-time":"2023-08-23T10:54:45Z","timestamp":1692788085000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Speech Feature Enhancement based on Time-frequency Analysis"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9014-1506","authenticated-orcid":false,"given":"Duc-Hao","family":"Do","sequence":"first","affiliation":[{"name":"University of Science, Vietnam National University, and FPT University, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-4877-2649","authenticated-orcid":false,"given":"Thanh-Duc","family":"Chau","sequence":"additional","affiliation":[{"name":"University of Science, and Vietnam National University, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4063-7095","authenticated-orcid":false,"given":"Thai-Son","family":"Tran","sequence":"additional","affiliation":[{"name":"University of Science, and Vietnam National University, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,8,23]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2015.2456097"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1994.389741"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/78.382394"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSCTI.2015.7489535"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/78.553486"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.5555\/200604"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.acha.2010.08.002"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","first-page":"719","DOI":"10.1007\/978-3-031-16014-1_56","volume-title":"Computational Collective Intelligence","author":"Do Hao D.","year":"2022","unstructured":"Hao D. Do, Duc T. Chau, and Son T. Tran. 2022. Speech representation using linear chirplet transform and its application in speaker-related recognition. In Computational Collective Intelligence, Ngoc Thanh Nguyen, Yannis Manolopoulos, Richard Chbeir, Adrianna Kozierkiewicz, and Bogdan Trawi\u0144ski (Eds.). Springer International Publishing, Cham, 719\u2013729."},{"key":"e_1_3_1_10_2","first-page":"93","volume-title":"DARPA Workshop on Speech Recognition","author":"Doddington George R. Goudie-Marshall, Kathleen M. Fisher, and William M.","year":"1986","unstructured":"George R. Goudie-Marshall, Kathleen M. Fisher, and William M. Doddington. 1986. The DARPA speech recognition research database: Specifications and status. In DARPA Workshop on Speech Recognition. 93\u201399."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2021.3127678"},{"key":"e_1_3_1_12_2","article-title":"Adam: A method for stochastic optimization","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014).","journal-title":"arXiv preprint arXiv:1412.6980"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1989.1.4.541"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3095054"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CEI52496.2021.9574552"},{"key":"e_1_3_1_16_2","first-page":"51","volume-title":"3rd International Workshop on Worldwide Language Service Infrastructure and 2nd Workshop on Open Infrastructures and Analysis Frameworks for Human Language Technologies (WLSI\/OIAF4HLT@COLING\u201916)","author":"Luong Hieu-Thi","year":"2016","unstructured":"Hieu-Thi Luong and Hai-Quan Vu. 2016. A non-expert Kaldi recipe for Vietnamese speech recognition system. In 3rd International Workshop on Worldwide Language Service Infrastructure and 2nd Workshop on Open Infrastructures and Analysis Frameworks for Human Language Technologies (WLSI\/OIAF4HLT@COLING\u201916), Yohei Murakami, Donghui Lin, Nancy Ide, and James Pustejovsky (Eds.). The COLING 2016 Organizing Committee, 51\u201355. Retrieved from https:\/\/aclanthology.org\/W16-5207\/."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/78.482123"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1992.226187"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.crhy.2019.07.001"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2009.2020355"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPEL.2020.3034585"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/78.678465"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1044\/1092-4388(2011\/11-0223)"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-011-9145-0"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2003.1198804"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICOIACT.2018.8350748"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIE.2018.2868296"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ymssp.2015.09.004"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3605549","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3605549","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:03Z","timestamp":1750182543000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3605549"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,23]]},"references-count":27,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2023,8,31]]}},"alternative-id":["10.1145\/3605549"],"URL":"https:\/\/doi.org\/10.1145\/3605549","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2023,8,23]]},"assertion":[{"value":"2022-08-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-09","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}