{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,29]],"date-time":"2025-11-29T07:59:47Z","timestamp":1764403187313,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":32,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,11,7]],"date-time":"2022-11-07T00:00:00Z","timestamp":1667779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100013348","name":"Innosuisse - Schweizerische Agentur f\u00fcr Innovationsf\u00f6rderung","doi-asserted-by":"publisher","award":["38843.1 IP-ICT"],"award-info":[{"award-number":["38843.1 IP-ICT"]}],"id":[{"id":"10.13039\/501100013348","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003475","name":"Hasler Stiftung","doi-asserted-by":"publisher","award":["FLOSS 16036"],"award-info":[{"award-number":["FLOSS 16036"]}],"id":[{"id":"10.13039\/501100003475","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,11,7]]},"DOI":"10.1145\/3536220.3563689","type":"proceedings-article","created":{"date-parts":[[2022,11,4]],"date-time":"2022-11-04T22:11:40Z","timestamp":1667599900000},"page":"7-11","source":"Crossref","is-referenced-by-count":2,"title":["Towards Automatic Prediction of Non-Expert Perceived Speech Fluency Ratings"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4307-8604","authenticated-orcid":false,"given":"S. Pavankumar","family":"Dubagunta","sequence":"first","affiliation":[{"name":"Uniphore Software Systems, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8714-1409","authenticated-orcid":false,"given":"Edoardo","family":"Moneta","sequence":"additional","affiliation":[{"name":"speak and lunch S.A., Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eleni","family":"Theocharopoulos","sequence":"additional","affiliation":[{"name":"speak and lunch S.A., Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mathew","family":"Magimai Doss","sequence":"additional","affiliation":[{"name":"Idiap Research Institute, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,11,7]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Mart\u00edn Abadi 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. http:\/\/tensorflow.org\/.  Mart\u00edn Abadi 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. http:\/\/tensorflow.org\/."},{"key":"e_1_3_2_1_2_1","unstructured":"[\n  2\n  ]  A. Baevski 2020. https:\/\/github.com\/pytorch\/fairseq\/tree\/master\/examples\/wav2vec.  [2] A. Baevski 2020. https:\/\/github.com\/pytorch\/fairseq\/tree\/master\/examples\/wav2vec."},{"volume-title":"Proc. NeurIPS. Curran Associates Inc., 12449\u201312460","author":"Baevski A.","key":"e_1_3_2_1_3_1","unstructured":"A. Baevski , H. Zhou , A. Mohamed , and M. Auli . 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations . In Proc. NeurIPS. Curran Associates Inc., 12449\u201312460 . A. Baevski, H. Zhou, A. Mohamed, and M. Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proc. NeurIPS. Curran Associates Inc., 12449\u201312460."},{"key":"e_1_3_2_1_4_1","unstructured":"F. Chollet 2015. Keras. https:\/\/github.com\/fchollet\/keras.  F. Chollet 2015. Keras. https:\/\/github.com\/fchollet\/keras."},{"key":"e_1_3_2_1_5_1","volume-title":"Dementia Recognition. In Proceedings of Interspeech. 2182\u20132186","author":"Cummins N.","year":"2020","unstructured":"N. Cummins 2020 . A Comparison of Acoustic and Linguistics Methodologies for Alzheimer\u2019s Dementia Recognition. In Proceedings of Interspeech. 2182\u20132186 . N. Cummins 2020. A Comparison of Acoustic and Linguistics Methodologies for Alzheimer\u2019s Dementia Recognition. In Proceedings of Interspeech. 2182\u20132186."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7953080"},{"key":"e_1_3_2_1_7_1","volume-title":"-Doss","author":"Dubagunta P.","year":"2019","unstructured":"S.\u00a0 P. Dubagunta and M. Magimai . -Doss . 2019 . Using Speech Production Knowledge for Raw Waveform Modelling based Styrian Dialect Identification. In Proceedings of Interspeech . S.\u00a0P. Dubagunta and M. Magimai.-Doss. 2019. Using Speech Production Knowledge for Raw Waveform Modelling based Styrian Dialect Identification. In Proceedings of Interspeech."},{"volume-title":"Proceedings of ICASSP. http:\/\/publications.idiap.ch\/downloads\/papers\/2019\/Dubagunta_ICASSP-2_2019","author":"Dubagunta P.","key":"e_1_3_2_1_8_1","unstructured":"S.\u00a0 P. Dubagunta , B. Vlasenko , and M. Magimai . -Doss. 2019. Learning voice source related information for depression detection . In Proceedings of ICASSP. http:\/\/publications.idiap.ch\/downloads\/papers\/2019\/Dubagunta_ICASSP-2_2019 .pdf S.\u00a0P. Dubagunta, B. Vlasenko, and M. Magimai.-Doss. 2019. Learning voice source related information for depression detection. In Proceedings of ICASSP. http:\/\/publications.idiap.ch\/downloads\/papers\/2019\/Dubagunta_ICASSP-2_2019.pdf"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1177\/0265532217712553"},{"volume-title":"Acoustic Features and Modelling","author":"Eyben Florian","key":"e_1_3_2_1_10_1","unstructured":"Florian Eyben . 2016. Acoustic Features and Modelling . Springer Int. Publishing , Cham , 9\u2013122. https:\/\/doi.org\/10.1007\/978-3-319-27299-3_2 10.1007\/978-3-319-27299-3_2 Florian Eyben. 2016. Acoustic Features and Modelling. Springer Int. Publishing, Cham, 9\u2013122. https:\/\/doi.org\/10.1007\/978-3-319-27299-3_2"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2015.2457417"},{"volume-title":"Proc. ACM Multimedia. 1459\u20131462","author":"Eyben F.","key":"e_1_3_2_1_12_1","unstructured":"F. Eyben , M. W\u00f6llmer , and B. Schuller . 2010. Opensmile: the munich versatile and fast open-source audio feature extractor . In Proc. ACM Multimedia. 1459\u20131462 . F. Eyben, M. W\u00f6llmer, and B. Schuller. 2010. Opensmile: the munich versatile and fast open-source audio feature extractor. In Proc. ACM Multimedia. 1459\u20131462."},{"key":"e_1_3_2_1_13_1","first-page":"2544","article-title":"Automatically Measuring L2 Speech Fluency without the Need of ASR","volume":"2018","author":"Fontan L.","year":"2018","unstructured":"L. Fontan , M. Le Coz , and S. Detey . 2018 . Automatically Measuring L2 Speech Fluency without the Need of ASR : A Proof-of-concept Study with Japanese Learners of French. In Proceedings of Interspeech 2018. 2544 \u2013 2548 . https:\/\/doi.org\/10.21437\/Interspeech.2018-1336 10.21437\/Interspeech.2018-1336 L. Fontan, M. Le Coz, and S. Detey. 2018. Automatically Measuring L2 Speech Fluency without the Need of ASR: A Proof-of-concept Study with Japanese Learners of French. In Proceedings of Interspeech 2018. 2544\u20132548. https:\/\/doi.org\/10.21437\/Interspeech.2018-1336","journal-title":"A Proof-of-concept Study with Japanese Learners of French. In Proceedings of Interspeech"},{"key":"e_1_3_2_1_14_1","volume-title":"-Doss","author":"Fritsch J.","year":"2020","unstructured":"J. Fritsch , S.\u00a0 P. Dubagunta , and M. Magimai . -Doss . 2020 . Estimating The Degree of Sleepiness by Integrating Articulatory Feature Knowledge In Raw Waveform Based CNNs. In Proceedings of ICASSP. J. Fritsch, S.\u00a0P. Dubagunta, and M. Magimai.-Doss. 2020. Estimating The Degree of Sleepiness by Integrating Articulatory Feature Knowledge In Raw Waveform Based CNNs. In Proceedings of ICASSP."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2019.2955022"},{"key":"e_1_3_2_1_16_1","volume-title":"Proceedings of Interspeech.","author":"Kelly C.","year":"2020","unstructured":"A.\u00a0 C. Kelly 2020 . SoapBox Labs Fluency Assessment Platform for Child Speech . In Proceedings of Interspeech. A.\u00a0C. Kelly 2020. SoapBox Labs Fluency Assessment Platform for Child Speech. In Proceedings of Interspeech."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-2889"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8682187"},{"volume-title":"Proceedings of ICASSP. 4884\u20134888","author":"Muckenhirn H.","key":"e_1_3_2_1_19_1","unstructured":"H. Muckenhirn , M. Magimai .-Doss, and S. Marcel . 2018. Towards directly modeling raw speech signal for speaker verification using CNNs . In Proceedings of ICASSP. 4884\u20134888 . H. Muckenhirn, M. Magimai.-Doss, and S. Marcel. 2018. Towards directly modeling raw speech signal for speaker verification using CNNs. In Proceedings of ICASSP. 4884\u20134888."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2008.2004526"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-917"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2013-438"},{"volume-title":"Proceedings of ICASSP. 5206\u20135210","author":"Panayotov V.","key":"e_1_3_2_1_23_1","unstructured":"V. Panayotov , G. Chen , D. Povey , and S. Khudanpur . 2015. Librispeech: an ASR corpus based on public domain audio books . In Proceedings of ICASSP. 5206\u20135210 . V. Panayotov, G. Chen, D. Povey, and S. Khudanpur. 2015. Librispeech: an ASR corpus based on public domain audio books. In Proceedings of ICASSP. 5206\u20135210."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_3_2_1_25_1","first-page":"1","article-title":"openXBOW \u2013 Introducing the Passau Open-Source Crossmodal Bag-of-Words Toolkit","volume":"18","author":"Schmitt M.","year":"2017","unstructured":"M. Schmitt and B. Schuller . 2017 . openXBOW \u2013 Introducing the Passau Open-Source Crossmodal Bag-of-Words Toolkit . Journal of Machine Learning Research 18 , 96 (2017), 1 \u2013 5 . http:\/\/jmlr.org\/papers\/v18\/17-113.html M. Schmitt and B. Schuller. 2017. openXBOW \u2013 Introducing the Passau Open-Source Crossmodal Bag-of-Words Toolkit. Journal of Machine Learning Research 18, 96 (2017), 1\u20135. http:\/\/jmlr.org\/papers\/v18\/17-113.html","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472669"},{"key":"e_1_3_2_1_27_1","first-page":"2018","volume-title":"Proceedings of Interspeech","author":"Wagner J.","year":"2018","unstructured":"J. Wagner , D. Schiller , A. Seiderer , and E. Andr\u00e9 . 2018. Deep Learning in Paralinguistic Recognition Tasks: Are Hand-crafted Features Still Relevant? . In Proceedings of Interspeech 2018 . https:\/\/doi.org\/10.21437\/Interspeech. 2018 - 1238 10.21437\/Interspeech.2018-1238 J. Wagner, D. Schiller, A. Seiderer, and E. Andr\u00e9. 2018. Deep Learning in Paralinguistic Recognition Tasks: Are Hand-crafted Features Still Relevant?. In Proceedings of Interspeech 2018. https:\/\/doi.org\/10.21437\/Interspeech.2018-1238"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-64"},{"key":"e_1_3_2_1_29_1","first-page":"2019","volume-title":"Proc. SLaTE. 48\u201352","author":"Xue W.","unstructured":"W. Xue , C. Cucchiarini , R. van Hout , and H. Strik . 2019. Acoustic correlates of speech intelligibility: the usability of the eGeMAPS feature set for atypical speech . In Proc. SLaTE. 48\u201352 . https:\/\/doi.org\/10.21437\/SLaTE. 2019 - 2019 10.21437\/SLaTE.2019-9 W. Xue, C. Cucchiarini, R. van Hout, and H. Strik. 2019. Acoustic correlates of speech intelligibility: the usability of the eGeMAPS feature set for atypical speech. In Proc. SLaTE. 48\u201352. https:\/\/doi.org\/10.21437\/SLaTE.2019-9"},{"key":"e_1_3_2_1_30_1","volume-title":"Proceedings of Interspeech. 968\u2013969","author":"Yarra C.","year":"2019","unstructured":"C. Yarra , A. Srinivasan , S. Gottimukkala , and P.\u00a0 K. Ghosh . 2019 . SPIRE-fluent: A Self-Learning App for Tutoring Oral Fluency to Second Language English Learners . In Proceedings of Interspeech. 968\u2013969 . C. Yarra, A. Srinivasan, S. Gottimukkala, and P.\u00a0K. Ghosh. 2019. SPIRE-fluent: A Self-Learning App for Tutoring Oral Fluency to Second Language English Learners. In Proceedings of Interspeech. 968\u2013969."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12046-011-0046-0"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2016-268"}],"event":{"name":"ICMI '22: INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"],"location":"Bengaluru India","acronym":"ICMI '22"},"container-title":["INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3536220.3563689","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3536220.3563689","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:52Z","timestamp":1750182532000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3536220.3563689"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,7]]},"references-count":32,"alternative-id":["10.1145\/3536220.3563689","10.1145\/3536220"],"URL":"https:\/\/doi.org\/10.1145\/3536220.3563689","relation":{},"subject":[],"published":{"date-parts":[[2022,11,7]]}}}