{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,31]],"date-time":"2025-12-31T09:42:04Z","timestamp":1767174124782,"version":"build-2238731810"},"reference-count":13,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,1,19]],"date-time":"2023-01-19T00:00:00Z","timestamp":1674086400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,1,19]],"date-time":"2023-01-19T00:00:00Z","timestamp":1674086400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>SincNet architecture has shown significant benefits over traditional Convolutional Neural Networks (CNN), especially for speaker recognition applications. SincNet comprises parameterized Sinc functions as filters in the first layer followed by convolutional layers. Although SincNet is compact in nature and offers top-level understanding of the features extracted, the effect of window function used in SincNet is not thoroughly addressed yet. Hamming and Hann are popularly used as the default time-localized windows to reduce spectral leakage. Hence, a comprehensive investigation of 28 different windowing functions on SincNet architecture towards speaker recognition task using TIMIT dataset was performed in this work. Additionally, \u201ctrainable\u201d\u00a0window functions were configured with tunable parameters to characterize the performance. The paper benchmarks the effect of the time-localized windowing function in terms of the bandwidth, side-lobe suppression, and spectral leakage for the filter banks employed in the first layer of the SincNet architecture. Trainable Gaussian and Cosine-Sum functions exhibited relative improvement of 41.46% and 82.11% in the sentence level classification error rate over Hamming window when employed on SincNet architecture.<\/jats:p>","DOI":"10.1186\/s13636-023-00271-0","type":"journal-article","created":{"date-parts":[[2023,1,19]],"date-time":"2023-01-19T04:04:05Z","timestamp":1674101045000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Trainable windows for SincNet architecture"],"prefix":"10.1186","volume":"2023","author":[{"given":"Prashanth","family":"H C","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2278-9148","authenticated-orcid":false,"given":"Madhav","family":"Rao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dhanya","family":"Eledath","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ramasubramanian","family":"V","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,1,19]]},"reference":[{"key":"271_CR1","unstructured":"M. Ravanelli. Deep learning for distant speech recognition (2017).\u00a0https:\/\/arxiv.org\/pdf\/1712.06086.pdf"},{"key":"271_CR2","doi-asserted-by":"publisher","unstructured":"M. Ravanelli, Y. Bengio, in 2018 IEEE Spoken Language Technology Workshop (SLT). Speaker recognition from raw waveform with sincnet (2018), pp. 1021\u20131028. https:\/\/doi.org\/10.1109\/SLT.2018.8639585","DOI":"10.1109\/SLT.2018.8639585"},{"key":"271_CR3","doi-asserted-by":"publisher","unstructured":"D. Eledath, P. Inbarajan, A. Biradar, S. Mahadeva, V. Ramasubramanian, in 2021 29th European Signal Processing Conference (EUSIPCO), End-to-end speech recognition from raw speech: Multi time-frequency resolution cnn architecture for efficient representation learning (2021), pp. 536\u2013540. https:\/\/doi.org\/10.23919\/EUSIPCO54536.2021.9616171","DOI":"10.23919\/EUSIPCO54536.2021.9616171"},{"key":"271_CR4","unstructured":"M. Ravanelli, Y. Bengio, in Proc. 32nd Conference on Neural Information Processing Systems (NIPS 2018) IRASL workshop, Montreal, Canada, Interpretable convolutional filters with sincnet.\u00a0arXiv\u00a0(2018)"},{"key":"271_CR5","doi-asserted-by":"publisher","unstructured":"S. Mittermaier, L. K\u00fcrzinger, B. Waschneck, G. Rigoll, in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Small-footprint keyword spotting on raw audio data with sinc-convolutions (2020), pp. 7454\u20137458. https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9053395","DOI":"10.1109\/ICASSP40776.2020.9053395"},{"key":"271_CR6","doi-asserted-by":"publisher","unstructured":"D. Onea\u0163\u0103, L. Georgescu, H. Cucu, D. Burileanu, C. Burileanu, in 2020 28th European Signal Processing Conference (EUSIPCO), Revisiting sincnet: An evaluation of feature and network hyperparameters for speaker recognition (2021), pp. 1\u20135. https:\/\/doi.org\/10.23919\/Eusipco47968.2020.9287794","DOI":"10.23919\/Eusipco47968.2020.9287794"},{"key":"271_CR7","doi-asserted-by":"publisher","unstructured":"L. Chowdhury, M. Kamal, N. Hasan, N. Mohammed, in 2021 International Conference of the Biometrics Special Interest Group (BIOSIG), Curricular sincnet: Towards robust deep speaker recognition by emphasizing hard samples in latent space (2021), pp. 1\u20134. https:\/\/doi.org\/10.1109\/BIOSIG52210.2021.9548296","DOI":"10.1109\/BIOSIG52210.2021.9548296"},{"key":"271_CR8","doi-asserted-by":"publisher","unstructured":"N. Zeghidour, N. Usunier, I. Kokkinos, T. Schatz, G. Synnaeve, E. Dupoux, Learning filterbanks from raw speech for phone recognition (2018). https:\/\/doi.org\/10.1109\/ICASSP.2018.8462015","DOI":"10.1109\/ICASSP.2018.8462015"},{"key":"271_CR9","doi-asserted-by":"publisher","unstructured":"P.G. No\u00e9, T. Parcollet, M. Morchid, Cgcnn: Complex gabor convolutional neural network on raw speech (2020). https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9054220","DOI":"10.1109\/ICASSP40776.2020.9054220"},{"key":"271_CR10","unstructured":"P.H. C. Trainable windows in sincnet.\u00a0https:\/\/sites.google.com\/view\/sincnet\/home. Accessed 10 Sept 2022"},{"key":"271_CR11","doi-asserted-by":"publisher","unstructured":"E. Loweimi, P. Bell, S. Renals, in Proceedings of Interspeech 2020, On the robustness and training dynamics of raw waveform models (International Speech Communication Association, 2020), pp. 1001\u20131005. https:\/\/doi.org\/10.21437\/Interspeech.2020-0017. http:\/\/www.interspeech2020.org\/. Interspeech 2020, INTERSPEECH 2020 ; Conference date: 25-10-2020 Through 29-10-2020","DOI":"10.21437\/Interspeech.2020-0017"},{"key":"271_CR12","unstructured":"J. Garofolo, L. Lamel, W. Fisher, J. Fiscus, D. Pallett, N. Dahlgren, V. Zue, Timit acoustic-phonetic continuous speech corpus. Linguist. Data Consortium (1992)"},{"key":"271_CR13","doi-asserted-by":"publisher","unstructured":"X. Liu, M. Sahidullah, T. Kinnunen, in 2021 IEEE International Symposium on Circuits and Systems (ISCAS), Learnable mfccs for speaker verification (2021), pp. 1\u20135. https:\/\/doi.org\/10.1109\/ISCAS51556.2021.9401593","DOI":"10.1109\/ISCAS51556.2021.9401593"}],"updated-by":[{"DOI":"10.1186\/s13636-023-00275-w","type":"correction","label":"Correction","source":"publisher","updated":{"date-parts":[[2023,2,9]],"date-time":"2023-02-09T00:00:00Z","timestamp":1675900800000}}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-023-00271-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-023-00271-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-023-00271-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,11]],"date-time":"2024-07-11T10:41:41Z","timestamp":1720694501000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-023-00271-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,19]]},"references-count":13,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["271"],"URL":"https:\/\/doi.org\/10.1186\/s13636-023-00271-0","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,19]]},"assertion":[{"value":"1 August 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 January 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 January 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"None","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Yes","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"3"}}