{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T21:22:17Z","timestamp":1775856137221,"version":"3.50.1"},"reference-count":46,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2024,10,10]],"date-time":"2024-10-10T00:00:00Z","timestamp":1728518400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,10,10]],"date-time":"2024-10-10T00:00:00Z","timestamp":1728518400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001823","name":"Ministerstvo \u0160kolstv\u00ed, Ml\u00e1de\u017ee a Telov\u00fdchovy","doi-asserted-by":"publisher","award":["90254"],"award-info":[{"award-number":["90254"]}],"id":[{"id":"10.13039\/501100001823","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100009532","name":"Ministerstvo Vnitra Cesk\u00e9 Republiky","doi-asserted-by":"publisher","award":["VJ01010108"],"award-info":[{"award-number":["VJ01010108"]}],"id":[{"id":"10.13039\/100009532","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100009056","name":"University of West Bohemia","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100009056","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Speech Technol"],"published-print":{"date-parts":[[2024,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The current state-of-the-art for various speech processing problems is a sequence-to-sequence model based on a self-attention mechanism known as transformer. The widely used wav2vec\u00a02.0 is a self-supervised transformer model pre-trained on large amounts of unlabeled speech and then fine-tuned for a specific task. The data used for training and fine-tuning, along with the size of the transformer model, play a crucial role in both of these training steps. The most commonly used wav2vec 2.0 models are trained on relatively \u201cclean\u201d data from sources such as the LibriSpeech dataset, but we can expect there to be a benefit in using more realistic data gathered from a variety of acoustic conditions. However, it is not entirely clear how big the difference would be. Investigating this is the main goal of our article. To this end, we utilize wav2vec\u00a02.0 models in three fundamental speech processing tasks: speaker change detection, voice activity detection, and overlapped speech detection, and test them on four real conversation datasets. We compare four wav2vec\u00a02.0 models with different sizes and different data used for pre-training, and we fine-tune them either on in-domain data from the same dataset or on artificial training data created from the LibriSpeech corpus. Our results suggest that richer data that are more similar to the task domain bring better performance than a larger model.<\/jats:p>","DOI":"10.1007\/s10772-024-10140-6","type":"journal-article","created":{"date-parts":[[2024,10,10]],"date-time":"2024-10-10T14:03:15Z","timestamp":1728568995000},"page":"847-859","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Comparison of wav2vec 2.0 models on three speech processing tasks"],"prefix":"10.1007","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7187-8481","authenticated-orcid":false,"given":"Marie","family":"Kune\u0161ov\u00e1","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4153-6560","authenticated-orcid":false,"given":"Zbyn\u011bk","family":"Zaj\u00edc","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8169-2410","authenticated-orcid":false,"given":"Lubo\u0161","family":"\u0160m\u00eddl","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6474-8366","authenticated-orcid":false,"given":"Martin","family":"Karafi\u00e1t","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,10,10]]},"reference":[{"key":"10140_CR1","doi-asserted-by":"publisher","first-page":"2324","DOI":"10.1109\/TASLP.2021.3093817","volume":"29","author":"OH Anidjar","year":"2021","unstructured":"Anidjar, O. H., Lapidot, I., Hajaj, C., Dvir, A., & Gilad, I. (2021). Hybrid speech and text analysis methods for speaker change detection. IEEE\/ACM Transactions on Audio Speech and Language Processing, 29, 2324\u20132338. https:\/\/doi.org\/10.1109\/TASLP.2021.3093817","journal-title":"IEEE\/ACM Transactions on Audio Speech and Language Processing"},{"key":"10140_CR5","doi-asserted-by":"publisher","unstructured":"Aronowitz, H., & Zhu, W. (2020). Context and uncertainty modeling for online speaker change detection. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP) (pp 8379\u20138383). https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9053280","DOI":"10.1109\/ICASSP40776.2020.9053280"},{"key":"10140_CR6","unstructured":"Baevski, A., Zhou, Y., Mohamed, A., Auli, M. (2020). wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 33, 12449\u201312460. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/92d1e1eb1cd6f9fba3227870bb6d7f07-Paper.pdf"},{"key":"10140_CR7","doi-asserted-by":"publisher","unstructured":"Bergelson, E. (2016). Bergelson seedlings HomeBank Corpus. https:\/\/doi.org\/10.21415\/T5PK6D","DOI":"10.21415\/T5PK6D"},{"key":"10140_CR8","doi-asserted-by":"publisher","unstructured":"Boakye, K., Trueba-Hornero, B., Vinyals, O., & Friedland, G. (2008). Overlapped speech detection for improved speaker diarization in multiparty meetings. In 2008 IEEE international conference on acoustics, speech and signal processing (pp. 4353\u20134356). https:\/\/doi.org\/10.1109\/ICASSP.2008.4518619","DOI":"10.1109\/ICASSP.2008.4518619"},{"key":"10140_CR9","doi-asserted-by":"publisher","unstructured":"Bredin, H. (2017). Pyannote.metrics: A toolkit for reproducible evaluation, diagnostic, and error analysis of speaker diarization systems. In Proceedings of the Interspeech (pp. 3587\u20133591). https:\/\/doi.org\/10.21437\/INTERSPEECH.2017-411","DOI":"10.21437\/INTERSPEECH.2017-411"},{"key":"10140_CR10","doi-asserted-by":"publisher","unstructured":"Bredin, H., & Laurent, A. (2021). End-to-end speaker segmentation for overlap-aware resegmentation. In Proceedings of the Interspeech (pp. 3111\u20133115). https:\/\/doi.org\/10.21437\/Interspeech.2021-560","DOI":"10.21437\/Interspeech.2021-560"},{"key":"10140_CR11","doi-asserted-by":"publisher","unstructured":"Bredin, H., Yin, R., Coria, J. M., Gelly, G., Korshunov, P., Lavechin, M., Fustes, D., Titeux, H., Bouaziz, W., & Gill, M. P. (2020). Pyannote.audio: Neural building blocks for speaker diarization. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP) (pp. 7124\u20137128). https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9052974","DOI":"10.1109\/ICASSP40776.2020.9052974"},{"key":"10140_CR12","doi-asserted-by":"publisher","unstructured":"Bullock, L., Bredin, H., & Garcia-Perera, L. P. (2020). Overlap-aware diarization: Resegmentation using neural end-to-end overlapped speech detection. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP) (pp. 7114\u20137118). https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9053096","DOI":"10.1109\/ICASSP40776.2020.9053096"},{"key":"10140_CR13","doi-asserted-by":"publisher","unstructured":"Canavan, A., Graff, D., & Zipperlen, G. (1997). CALLHOME American English Speech, LDC97S42. https:\/\/doi.org\/10.35111\/exq3-x930","DOI":"10.35111\/exq3-x930"},{"issue":"2","key":"10140_CR14","doi-asserted-by":"publisher","first-page":"181","DOI":"10.1007\/S10579-007-9040-X","volume":"41","author":"J Carletta","year":"2007","unstructured":"Carletta, J. (2007). Unleashing the killer corpus: Experiences in creating the multi-everything AMI Meeting Corpus. Language Resources and Evaluation, 41(2), 181\u2013190. https:\/\/doi.org\/10.1007\/S10579-007-9040-X","journal-title":"Language Resources and Evaluation"},{"key":"10140_CR15","doi-asserted-by":"publisher","first-page":"1636","DOI":"10.1109\/TASLP.2024.3366756","volume":"32","author":"Z Chen","year":"2024","unstructured":"Chen, Z., Han, B., Wang, S., Qian, Y. (2024). Attention-based encoder-decoder end-to-end neural diarization with embedding enhancer. IEEE\/ACM Transactions on Audio, Speech, and Language Processing, 32, 1636\u20131649. https:\/\/doi.org\/10.1109\/TASLP.2024.3366756","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"10140_CR16","doi-asserted-by":"publisher","unstructured":"Conneau, A., Baevski, A., Collobert, R., Mohamed, A., & Auli, M. (2021). Unsupervised cross-lingual representation learning for speech recognition. In Proceedings of the Interspeech (pp. 2426\u20132430). https:\/\/doi.org\/10.21437\/Interspeech.2021-329","DOI":"10.21437\/Interspeech.2021-329"},{"key":"10140_CR17","doi-asserted-by":"publisher","unstructured":"Cornell, S., Omologo, M., Squartini, S., & Vincent, E. (2020). Detecting and counting overlapping speakers in distant speech scenarios. In Proceedings of the Interspeech (pp. 3107\u20133111). https:\/\/doi.org\/10.21437\/Interspeech.2020-2671","DOI":"10.21437\/Interspeech.2020-2671"},{"key":"10140_CR18","doi-asserted-by":"publisher","first-page":"1551","DOI":"10.1109\/LSP.2022.3185955","volume":"29","author":"Z Fan","year":"2022","unstructured":"Fan, Z., Dong, L., Cai, M., Ma, Z., & Xu, B. (2022). Sequence-level speaker change detection with difference-based continuous integrate-and-fire. IEEE Signal Processing Letters, 29, 1551\u20131554. https:\/\/doi.org\/10.1109\/LSP.2022.3185955","journal-title":"IEEE Signal Processing Letters"},{"key":"10140_CR19","doi-asserted-by":"publisher","unstructured":"Han, E., Lee, C., & Stolcke, A. (2021). BW-EDA-EEND: Streaming END-TO-END neural speaker diarization for a variable number of speakers. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP) (pp. 7193\u20137197). https:\/\/doi.org\/10.1109\/ICASSP39728.2021.9414371","DOI":"10.1109\/ICASSP39728.2021.9414371"},{"key":"10140_CR20","doi-asserted-by":"publisher","unstructured":"Hogg, A. O. T., Evers, C., & Naylor, P. A. (2019). Speaker change detection using fundamental frequency with application to multi-talker segmentation. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP) (pp. 5826\u20135830). https:\/\/doi.org\/10.1109\/ICASSP.2019.8682924","DOI":"10.1109\/ICASSP.2019.8682924"},{"key":"10140_CR21","doi-asserted-by":"publisher","first-page":"226","DOI":"10.1007\/978-3-319-99579-3_24","volume":"11096","author":"M Hr\u00faz","year":"2018","unstructured":"Hr\u00faz, M., & Hlav\u00e1\u010d, M. (2018). LSTM neural network for speaker change detection in telephone conversations. Speech and Computer SPECOM 2018 Lecture Notes in Computer Science, 11096, 226\u2013233. https:\/\/doi.org\/10.1007\/978-3-319-99579-3_24","journal-title":"Speech and Computer SPECOM 2018 Lecture Notes in Computer Science"},{"key":"10140_CR22","doi-asserted-by":"publisher","unstructured":"Hr\u00faz, M., & Zaj\u00edc, Z. (2017). Convolutional neural network for speaker change detection in telephone speaker diarization system. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP) (pp. 4945\u20134949). https:\/\/doi.org\/10.1109\/ICASSP.2017.7953097","DOI":"10.1109\/ICASSP.2017.7953097"},{"key":"10140_CR23","doi-asserted-by":"publisher","unstructured":"Jung, J. W., Seo, S., Heo, H. S., Kim, G., Kim, Y. J., Kwon, Y. K., Lee, M., & Lee, B. J. (2023). Encoder-decoder multimodal speaker change detection. In Proceedings of the Interspeech (pp. 5311\u20135315). https:\/\/doi.org\/10.21437\/Interspeech.2023-2289","DOI":"10.21437\/Interspeech.2023-2289"},{"key":"10140_CR24","doi-asserted-by":"publisher","unstructured":"Kazimirova, E., & Belyaev, A. (2018). Automatic detection of multi-speaker fragments with high time resolution. In Proceedings of the Interspeech (pp. 1388\u20131392). https:\/\/doi.org\/10.21437\/Interspeech.2018-1878","DOI":"10.21437\/Interspeech.2018-1878"},{"key":"10140_CR25","doi-asserted-by":"publisher","first-page":"377","DOI":"10.1007\/978-3-031-16270-1_31","volume":"13502","author":"M Kune\u0161ov\u00e1","year":"2022","unstructured":"Kune\u0161ov\u00e1, M., & \u0158ez\u00e1\u010dkov\u00e1, M. (2022). Detection of prosodic boundaries in speech using wav2vec 2.0. Text, Speech, and Dialogue TSD 2022 Lecture Notes in Computer Science, 13502, 377\u2013388. https:\/\/doi.org\/10.1007\/978-3-031-16270-1_31","journal-title":"Text, Speech, and Dialogue TSD 2022 Lecture Notes in Computer Science"},{"key":"10140_CR4","doi-asserted-by":"publisher","unstructured":"Kune\u0161ov\u00e1, M., & Zaj\u00edc, Z. (2023). Multitask detection of speaker changes, overlapping speech and voice activity using wav2vec 2.0. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP) (pp. 1\u20135). https:\/\/doi.org\/10.1109\/ICASSP49357.2023.10094972","DOI":"10.1109\/ICASSP49357.2023.10094972"},{"key":"10140_CR2","doi-asserted-by":"publisher","unstructured":"Kune\u0161ov\u00e1, M., Hr\u00faz, M., Zaj\u00edc, Z., & Radov\u00e1, V. (2019). Detection of overlapping speech for the purposes of speaker diarization. In Speech and computer SPECOM 2019 lecture notes in computer science (Vol. 11658, pp. 247\u2013257). https:\/\/doi.org\/10.1007\/978-3-030-26061-3_26","DOI":"10.1007\/978-3-030-26061-3_26"},{"key":"10140_CR26","doi-asserted-by":"publisher","first-page":"5095","DOI":"10.21437\/Interspeech.2022-10451","volume":"2022","author":"F Landini","year":"2022","unstructured":"Landini, F., Lozano-Diez, A., Diez, M., & Burget, L. (2022). From simulated mixtures to simulated conversations as training data for end-to-end neural diarization. In Proceedings of the Interspeech, 2022 (pp. 5095\u20135099). https:\/\/doi.org\/10.21437\/Interspeech.2022-10451","journal-title":"In: Proc. Interspeech"},{"key":"10140_CR27","doi-asserted-by":"publisher","unstructured":"Lehe\u010dka, J., \u0160vec, J., Pra\u017e\u00e1k, A., & Psutka, J. V. (2022). Exploring capabilities of monolingual audio transformers using large datasets in automatic speech recognition of Czech. In Proceedings of the Interspeech (pp. 1831\u20131835). https:\/\/doi.org\/10.21437\/INTERSPEECH.2022-10439","DOI":"10.21437\/INTERSPEECH.2022-10439"},{"key":"10140_CR28","doi-asserted-by":"publisher","unstructured":"Liu, A. T., Li, S. W. W., & Lee, Hy Y. (2021). TERA: Self-supervised learning of transformer encoder representation for speech. IEEE\/ACM Transactions on Audio Speech and Language Processing 29:2351\u20132366. https:\/\/doi.org\/10.1109\/TASLP.2021.3095662","DOI":"10.1109\/TASLP.2021.3095662"},{"key":"10140_CR29","doi-asserted-by":"publisher","unstructured":"Mariotte, T., Larcher, A., Montr\u00e9sor S, & Thomas, J. H. (2024). Channel-combination algorithms for robust distant voice activity and overlapped speech detection. IEEE\/ACM Transactions on Audio, Speech, and Language Processing, 32, 1859\u20131872. https:\/\/doi.org\/10.1109\/TASLP.2024.3369531","DOI":"10.1109\/TASLP.2024.3369531"},{"key":"10140_CR30","doi-asserted-by":"publisher","unstructured":"Mateju, L., Kynych, F., Cerva, P., Malek, J., & Zd\u00e1nsk\u00fd, J. (2022). Overlapped speech detection in broadcast streams using x-vectors. In Proceedings of  the Interspeech (pp. 4606\u20134610). https:\/\/doi.org\/10.21437\/Interspeech.2022-81","DOI":"10.21437\/Interspeech.2022-81"},{"key":"10140_CR31","doi-asserted-by":"publisher","first-page":"2818","DOI":"10.21437\/Interspeech.2018-2304","volume":"2018","author":"VA Miasato Filho","year":"2018","unstructured":"Miasato Filho, V. A., Silva, D. A., & Depra Cuozzo, L. G. (2018). Joint discriminative embedding learning, speech activity and overlap detection for the DIHARD speaker diarization challenge. In Proceedings of  the Interspeech, 2018 (pp. 2818\u20132822). https:\/\/doi.org\/10.21437\/Interspeech.2018-2304","journal-title":"In: Proc. Interspeech"},{"key":"10140_CR32","doi-asserted-by":"publisher","unstructured":"Panayotov, V., Chen, G., Povey, D., Khudanpur, S. (2015). LibriSpeech: An ASR corpus based on public domain audio books. In Proceedings of  the international conference on acoustics, speech, and signal processing (ICASSP) (pp. 5206\u20135210). https:\/\/doi.org\/10.1109\/ICASSP.2015.7178964","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"10140_CR33","doi-asserted-by":"publisher","unstructured":"Ramirez, J., G\u00f3rriz, J. M., & Segura, J. C. (2007). Voice activity detection. Fundamentals and speech recognition system robustness. In Robust speech recognition and understanding (pp. 1\u201322). https:\/\/doi.org\/10.5772\/4740","DOI":"10.5772\/4740"},{"key":"10140_CR34","doi-asserted-by":"publisher","unstructured":"Rouvier, M., Dupuy, G., Gay, P., Khoury, E., Merlin, T., & Meignier, S. (2013). An open-source state-of-the-art toolbox for broadcast news diarization. In Proceedings of  the Interspeech (pp. 1477\u20131481). https:\/\/doi.org\/10.21437\/INTERSPEECH.2013-383","DOI":"10.21437\/INTERSPEECH.2013-383"},{"key":"10140_CR35","unstructured":"Ryant, N., Church, K., Cieri, C., Cristia, A., Du, J., Ganapathy, S., & Liberman, M. (2018). First DIHARD challenge evaluation plan. Tech. rep., Linguistic Data Consortium, https:\/\/catalog.ldc.upenn.edu\/docs\/LDC2019S09\/first_dihard_eval_plan_v1.3.pdf"},{"key":"10140_CR36","doi-asserted-by":"publisher","unstructured":"Ryant, N., Church, K., Cieri, C., Cristia, A., Du, J., Ganapathy, S., & Liberman, M. (2019). The second DIHARD diarization challenge: Dataset, task, and baselines. In Proceedings of interspeech (pp. 978\u2013982). https:\/\/doi.org\/10.21437\/Interspeech.2019-1268","DOI":"10.21437\/Interspeech.2019-1268"},{"key":"10140_CR37","doi-asserted-by":"publisher","unstructured":"Su, H., Zhao, D., Dang, L., Li, M., Wu, X., Liu, X., & Meng, H. (2022). A multitask learning framework for speaker change detection with content information from unsupervised speech decomposition. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP) (pp. 8087\u20138091). https:\/\/doi.org\/10.1109\/ICASSP43922.2022.9746116","DOI":"10.1109\/ICASSP43922.2022.9746116"},{"key":"10140_CR38","doi-asserted-by":"publisher","unstructured":"Tong, S., Gu, H., & Yu, K. (2016). A comparative study of robustness of deep learning approaches for VAD. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP) (pp. 5695\u20135699). https:\/\/doi.org\/10.1109\/ICASSP.2016.7472768","DOI":"10.1109\/ICASSP.2016.7472768"},{"key":"10140_CR39","doi-asserted-by":"publisher","unstructured":"Vaessen, N., & Van Leeuwen, D. A. (2022). Fine-tuning wav2vec2 for speaker recognition. In 2022 IEEE international conference on acoustics, speech and signal processing (ICASSP 2022) (pp. 7967\u20137971). https:\/\/doi.org\/10.1109\/ICASSP43922.2022.9746952","DOI":"10.1109\/ICASSP43922.2022.9746952"},{"key":"10140_CR40","unstructured":"Vaswani, A. (2017). Attention is all you need. In Proceedings of the 31st international conference on neural information processing systems (NIPS\u201917) (pp. 5998\u20136008). https:\/\/papers.nips.cc\/paper\/7181-attention-is-all-you-need.pdf"},{"key":"10140_CR41","doi-asserted-by":"publisher","unstructured":"Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., & Davison, J. (2020). Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: System demonstrations (pp. 38\u201345). https:\/\/doi.org\/10.18653\/v1\/2020.emnlp-demos.6","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"10140_CR42","doi-asserted-by":"publisher","unstructured":"Wu, J., Chen, Z., Hu, M., Xiao, X., & Li, J. (2023). Speaker change detection for transformer transducer ASR. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP) (pp. 1\u20135). https:\/\/doi.org\/10.1109\/ICASSP49357.2023.10096361, arXiv: 2302.08549","DOI":"10.1109\/ICASSP49357.2023.10096361"},{"key":"10140_CR43","doi-asserted-by":"publisher","unstructured":"Yang, S. W., Chi, P. H., Chuang, Y. S., Lai, C. I. J., Lakhotia, K., Lin, Y. Y., Liu, A. T., Shi, J., Chang, X., Lin, G. T., & Huang, T. H. (2021). SUPERB: Speech processing Universal PERformance Benchmark. In Proceedings of the Interspeech (pp. 1194\u20131198). https:\/\/doi.org\/10.21437\/Interspeech.2021-1775","DOI":"10.21437\/Interspeech.2021-1775"},{"key":"10140_CR44","doi-asserted-by":"publisher","unstructured":"Yin, R., Bredin, H., & Barras, C. (2017). Speaker change detection in broadcast TV using bidirectional long short-term memory networks. In Proceedings of the Interspeech 2017 (pp. 3827\u20133831). https:\/\/doi.org\/10.21437\/Interspeech.2017-65","DOI":"10.21437\/Interspeech.2017-65"},{"key":"10140_CR3","doi-asserted-by":"crossref","unstructured":"Zaj\u00edc, Z., & Kune\u0161ov\u00e1, M. (2023). Comparison of wav2vec 2.0 transformer models for speaker change detection. In M. Abbas, A. A. Freihat (Eds.), Proceedings of the 6th international conference on natural language and speech processing (ICNLSP 2023) (pp. 233\u2013238). https:\/\/aclanthology.org\/2023.icnlsp-1.23","DOI":"10.1109\/ICASSP49357.2023.10094972"},{"key":"10140_CR45","doi-asserted-by":"publisher","unstructured":"Zaj\u00edc, Z., Kune\u0161ov\u00e1, M., Zelinka, J., & Hr\u00faz, M. (2018). ZCU-NTIS speaker diarization system for the DIHARD 2018 challenge. In Proceedings of the Interspeech (pp. 2788\u20132792). https:\/\/doi.org\/10.21437\/Interspeech.2018-1252","DOI":"10.21437\/Interspeech.2018-1252"},{"key":"10140_CR46","doi-asserted-by":"publisher","first-page":"342","DOI":"10.1007\/978-3-030-00794-2_37","volume":"11107","author":"Z Zaj\u00edc","year":"2018","unstructured":"Zaj\u00edc, Z., Soutner, D., Hr\u00faz, M., M\u00fcller, L., & Radov\u00e1, V. (2018). Recurrent neural network based speaker change detection from text transcription applied in telephone speaker diarization system. Text, Speech, and Dialogue TSD 2018 Lecture Notes in Computer Science, 11107, 342\u2013350. https:\/\/doi.org\/10.1007\/978-3-030-00794-2_37","journal-title":"Text, Speech, and Dialogue TSD 2018 Lecture Notes in Computer Science"}],"container-title":["International Journal of Speech Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10772-024-10140-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10772-024-10140-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10772-024-10140-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,16]],"date-time":"2024-12-16T10:07:17Z","timestamp":1734343637000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10772-024-10140-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,10]]},"references-count":46,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,12]]}},"alternative-id":["10140"],"URL":"https:\/\/doi.org\/10.1007\/s10772-024-10140-6","relation":{},"ISSN":["1381-2416","1572-8110"],"issn-type":[{"value":"1381-2416","type":"print"},{"value":"1572-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,10]]},"assertion":[{"value":"6 August 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 September 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 October 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}