{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:12:52Z","timestamp":1750219972714,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":30,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3551606","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:42:46Z","timestamp":1665416566000},"page":"7195-7199","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Audio Features from the Wav2Vec 2.0 Embeddings for the ACM Multimedia 2022 Stuttering Challenge"],"prefix":"10.1145","author":[{"given":"Claude","family":"Montaci\u00e9","sequence":"first","affiliation":[{"name":"Sorbonne University, Paris, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marie-Jos\u00e9","family":"Caraty","sequence":"additional","affiliation":[{"name":"Paris University, Paris, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nikola","family":"Lackovic","sequence":"additional","affiliation":[{"name":"Malakof Humanis, Paris, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","article-title":"Overview of the diagnosis and treatment of stuttering","volume":"2012","author":"Maguire Gerald A.","year":"2012","unstructured":"Gerald A. Maguire , Christopher Y. Yeh , and Brandon S. Ito . 2012 . Overview of the diagnosis and treatment of stuttering . Journal of Experimental & Clinical Medicine , 2012 , Vol. 4, no 2, 92--97. Gerald A. Maguire, Christopher Y. Yeh, and Brandon S. Ito. 2012. Overview of the diagnosis and treatment of stuttering. Journal of Experimental & Clinical Medicine, 2012, Vol. 4, no 2, 92--97.","journal-title":"Journal of Experimental & Clinical Medicine"},{"key":"e_1_3_2_2_2_1","volume-title":"Diagnostic and Statistical Manual of Mental Disorders","author":"American Psychiatric Association","unstructured":"American Psychiatric Association . 2013. Diagnostic and Statistical Manual of Mental Disorders ( 5 th ed.), DSM-V. VA , Arlington . American Psychiatric Association. 2013. Diagnostic and Statistical Manual of Mental Disorders (5th ed.), DSM-V. VA, Arlington.","edition":"5"},{"key":"e_1_3_2_2_3_1","first-page":"229","article-title":"Literature survey and review of techniques used for automatic assessment of Stuttered Speech","volume":"9","author":"Gupta Sakshi","year":"2019","unstructured":"Sakshi Gupta , Ravi S. Shukla , and Rajesh K. Shukla . 2019 . Literature survey and review of techniques used for automatic assessment of Stuttered Speech . International Journal of Management, Technology and Engineering , Vol. 9 , 229 -- 240 . Sakshi Gupta, Ravi S. Shukla, and Rajesh K. Shukla. 2019. Literature survey and review of techniques used for automatic assessment of Stuttered Speech. International Journal of Management, Technology and Engineering, Vol. 9, 229--240.","journal-title":"International Journal of Management, Technology and Engineering"},{"key":"e_1_3_2_2_4_1","unstructured":"Shakeel A. Sheikh Md Sahidullah Fabrice Hirsch and Slim Ouni. 2021. Machine learning for stuttering identification: Review challenges & future directions. arXiv:2107.04057. Retrieved from https:\/\/arxiv.org\/pdf\/2107.04057  Shakeel A. Sheikh Md Sahidullah Fabrice Hirsch and Slim Ouni. 2021. Machine learning for stuttering identification: Review challenges & future directions. arXiv:2107.04057. Retrieved from https:\/\/arxiv.org\/pdf\/2107.04057"},{"volume-title":"Systematic review of machine learning approaches for detecting developmental stuttering.\u00a0IEEE\/ACM Transactions on Audio, Speech, and Language Processing","author":"Barrett Liam","key":"e_1_3_2_2_5_1","unstructured":"Liam Barrett , Junchao Hu , and Peter Howell . 2022. Systematic review of machine learning approaches for detecting developmental stuttering.\u00a0IEEE\/ACM Transactions on Audio, Speech, and Language Processing , Vol. 30 , 1160--1172. Liam Barrett, Junchao Hu, and Peter Howell. 2022. Systematic review of machine learning approaches for detecting developmental stuttering.\u00a0IEEE\/ACM Transactions on Audio, Speech, and Language Processing, Vol. 30, 1160--1172."},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICIAS.2007.4658401"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-81-322-2538-6_63"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2012.05.003"},{"volume-title":"Proceedings of the 3th international Conference on Cryptography, Security and Privacy, 93--98","author":"Manjula G.","key":"e_1_3_2_2_9_1","unstructured":"G. Manjula , M. Shivakumar , and Yelimeli V. Geetha . 2019. Adaptive optimization based neural network for classification of stuttered speech . In Proceedings of the 3th international Conference on Cryptography, Security and Privacy, 93--98 . G. Manjula, M. Shivakumar, and Yelimeli V. Geetha. 2019. Adaptive optimization based neural network for classification of stuttered speech. In Proceedings of the 3th international Conference on Cryptography, Security and Privacy, 93--98."},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053893"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746638"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2019.101052"},{"key":"e_1_3_2_2_13_1","unstructured":"Rachid Riad Anne-Catherine Bachoud-L\u00e9vi Frank Rudzicz Emmanuel Dupoux. 2020. Identification of primary and collateral tracks in stuttered speech. arXiv:2003.01018. Retrieved from https:\/\/arxiv.org\/pdf\/2003.01018  Rachid Riad Anne-Catherine Bachoud-L\u00e9vi Frank Rudzicz Emmanuel Dupoux. 2020. Identification of primary and collateral tracks in stuttered speech. arXiv:2003.01018. Retrieved from https:\/\/arxiv.org\/pdf\/2003.01018"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.23919\/EUSIPCO54536.2021.9616063"},{"key":"e_1_3_2_2_15_1","unstructured":"Shakeel A. Sheikh Md Sahidullah Fabrice Hirsch and Slim Ouni. 2022. Introducing ECAPA-TDNN and Wav2Vec2.0 Embeddings to Stuttering Detection. arXiv:2204.01564. Retrieved from https:\/\/arxiv.org\/pdf\/2204.01564  Shakeel A. Sheikh Md Sahidullah Fabrice Hirsch and Slim Ouni. 2022. Introducing ECAPA-TDNN and Wav2Vec2.0 Embeddings to Stuttering Detection. arXiv:2204.01564. Retrieved from https:\/\/arxiv.org\/pdf\/2204.01564"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1044\/1092-4388(2009\/07-0129)"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jfludis.2018.03.002"},{"volume-title":"Proceedings of the IEEE-ICASSP International Conference on Acoustics, Speech and Signal Processing, 6798--6802","author":"Colin Lea","key":"e_1_3_2_2_18_1","unstructured":"Lea Colin , Mitra Vikramjit , Joshi Aparna , Kajarekar Sachin , and Jeffrey P. Bigham . 2021. Sep-28k: A dataset for stuttering event detection from podcasts with people who stutter . In Proceedings of the IEEE-ICASSP International Conference on Acoustics, Speech and Signal Processing, 6798--6802 . Lea Colin, Mitra Vikramjit, Joshi Aparna, Kajarekar Sachin, and Jeffrey P. Bigham. 2021. Sep-28k: A dataset for stuttering event detection from podcasts with people who stutter. In Proceedings of the IEEE-ICASSP International Conference on Acoustics, Speech and Signal Processing, 6798--6802."},{"key":"e_1_3_2_2_19_1","volume-title":"Florian H\u00f6nig, Elmar N\u00f6th, and Korbinian Riedhammer.","author":"Bayerl Sebastian P.","year":"2022","unstructured":"Sebastian P. Bayerl , Alexander Wolff von Gudenberg , Florian H\u00f6nig, Elmar N\u00f6th, and Korbinian Riedhammer. 2022 . KSoF: The Kassel State of Fluency Dataset - A Therapy Centered Dataset of Stuttering . arXiv:2203.05383. Retrieved from https:\/\/arxiv.org\/pdf\/2203.05383 Sebastian P. Bayerl, Alexander Wolff von Gudenberg, Florian H\u00f6nig, Elmar N\u00f6th, and Korbinian Riedhammer. 2022. KSoF: The Kassel State of Fluency Dataset - A Therapy Centered Dataset of Stuttering. arXiv:2203.05383. Retrieved from https:\/\/arxiv.org\/pdf\/2203.05383"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3551591"},{"key":"e_1_3_2_2_21_1","first-page":"12449","article-title":"wav2vec 2.0: A framework for self-supervised learning of speech representations","volume":"33","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski , Yuhao Zhou , Abdelrahman Mohamed , and Michael Auli . 2020 . wav2vec 2.0: A framework for self-supervised learning of speech representations . Advances in Neural Information Processing Systems , Vol. 33 , 12449 -- 12460 . Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, Vol. 33, 12449--12460.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2502081.2502224"},{"key":"e_1_3_2_2_23_1","unstructured":"?https:\/\/huggingface.co\/docs\/transformers\/model_doc\/wav2vec2  ?https:\/\/huggingface.co\/docs\/transformers\/model_doc\/wav2vec2"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Leonardo Pepino Pablo Riera and Luciana Ferrer. 2021. Emotion recognition from speech using wav2vec 2.0 embeddings. arXiv:2104.03502. Retrieved from https:\/\/arxiv.org\/pdf\/2104.03502  Leonardo Pepino Pablo Riera and Luciana Ferrer. 2021. Emotion recognition from speech using wav2vec 2.0 embeddings. arXiv:2104.03502. Retrieved from https:\/\/arxiv.org\/pdf\/2104.03502","DOI":"10.21437\/Interspeech.2021-703"},{"key":"e_1_3_2_2_25_1","volume-title":"Proc. AAAI SAS","author":"Dumpala Sri Harsha","year":"2022","unstructured":"Sri Harsha Dumpala , Sebastian Rodriguez , Sheri Rempel , Mehri Sajjadian , Rudolf Uher , and Sageev Oore . 2022 . Detecting Depression with a Temporal Context Of Speaker Embeddings [J] . Proc. AAAI SAS , 2022. Sri Harsha Dumpala, Sebastian Rodriguez, Sheri Rempel, Mehri Sajjadian, Rudolf Uher, and Sageev Oore. 2022. Detecting Depression with a Temporal Context Of Speaker Embeddings [J]. Proc. AAAI SAS, 2022."},{"key":"e_1_3_2_2_26_1","volume-title":"Hansen","author":"Yang Mu","year":"2022","unstructured":"Mu Yang , Kevin Hirschi , Stephen D. Looney , Okim Kang , and John H. L . Hansen . 2022 . Improving Mispronunciation Detection with Wav2vec2-based Momentum Pseudo-Labeling for Accentedness and Intelligibility Assessment. arXiv:2203.15937. Retrieved from https:\/\/arxiv.org\/pdf\/2203.15937 Mu Yang, Kevin Hirschi, Stephen D. Looney, Okim Kang, and John H. L. Hansen. 2022. Improving Mispronunciation Detection with Wav2vec2-based Momentum Pseudo-Labeling for Accentedness and Intelligibility Assessment. arXiv:2203.15937. Retrieved from https:\/\/arxiv.org\/pdf\/2203.15937"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"crossref","unstructured":"Sebastian P. Bayerl Dominik Wagner Elmar N\u00f6th and Korbinian Riedhammer. 2022. Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0. arXiv:2204.03417. Retrieved from https:\/\/arxiv.org\/pdf\/2204.03417  Sebastian P. Bayerl Dominik Wagner Elmar N\u00f6th and Korbinian Riedhammer. 2022. Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0. arXiv:2204.03417. Retrieved from https:\/\/arxiv.org\/pdf\/2204.03417","DOI":"10.21437\/Interspeech.2022-10908"},{"key":"e_1_3_2_2_28_1","article-title":"Scikit-learn: Machine learning in Python","author":"Pedregosa Fabian","year":"2011","unstructured":"Fabian Pedregosa , Ga\u00ebl Varoquaux , Alexandre Gramfort , Vincent Michel , Bertrand Thirion , Olivier Grisel , and Jake Vanderplas . 2011 . Scikit-learn: Machine learning in Python \", Journal of machine learning research, 2825--2830. Fabian Pedregosa, Ga\u00ebl Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, and Jake Vanderplas. 2011. Scikit-learn: Machine learning in Python\", Journal of machine learning research, 2825--2830.","journal-title":"Journal of machine learning research, 2825--2830."},{"key":"e_1_3_2_2_29_1","unstructured":"Li-Wei Chen and Alexander Rudnicky. 2021. Exploring Wav2vec 2.0 fine-tuning for improved speech emotion recognition. arXiv:2110.06309. Retrieved from https:\/\/arxiv.org\/pdf\/2110.06309  Li-Wei Chen and Alexander Rudnicky. 2021. Exploring Wav2vec 2.0 fine-tuning for improved speech emotion recognition. arXiv:2110.06309. Retrieved from https:\/\/arxiv.org\/pdf\/2110.06309"},{"volume-title":"Proceedings of the IEEE-ICASSP International Conference on Acoustics, Speech and Signal Processing, 7967--7971","author":"Vaessen Nik","key":"e_1_3_2_2_30_1","unstructured":"Nik Vaessen and David A . Van Leeuwen. 2022. Fine-tuning wav2vec2 for speaker recognition . In Proceedings of the IEEE-ICASSP International Conference on Acoustics, Speech and Signal Processing, 7967--7971 . Nik Vaessen and David A. Van Leeuwen. 2022. Fine-tuning wav2vec2 for speaker recognition. In Proceedings of the IEEE-ICASSP International Conference on Acoustics, Speech and Signal Processing, 7967--7971."}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Lisboa Portugal","acronym":"MM '22"},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3551606","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3551606","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:18Z","timestamp":1750182558000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3551606"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":30,"alternative-id":["10.1145\/3503161.3551606","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3551606","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}