{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T03:23:33Z","timestamp":1784085813637,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":20,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100002341","name":"Academy of Finland","doi-asserted-by":"publisher","award":["345790"],"award-info":[{"award-number":["345790"]}],"id":[{"id":"10.13039\/501100002341","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004785","name":"NordForsk","doi-asserted-by":"publisher","award":["103893"],"award-info":[{"award-number":["103893"]}],"id":[{"id":"10.13039\/501100004785","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3551572","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:43:12Z","timestamp":1665416592000},"page":"7026-7029","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":25,"title":["Wav2vec2-based Paralinguistic Systems to Recognise Vocalised Emotions and Stuttering"],"prefix":"10.1145","author":[{"given":"Tam\u00e1s","family":"Gr\u00f3sz","sequence":"first","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dejan","family":"Porjazovski","sequence":"additional","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yaroslav","family":"Getman","sequence":"additional","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sudarsana","family":"Kadiri","sequence":"additional","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mikko","family":"Kurimo","sequence":"additional","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Youtube-8m: A large-scale video classification benchmark. arXiv preprint arXiv:1609.08675","author":"Abu-El-Haija Sami","year":"2016","unstructured":"Sami Abu-El-Haija , Nisarg Kothari , Joonseok Lee , Paul Natsev , George Toderici , Balakrishnan Varadarajan , and Sudheendra Vijayanarasimhan . 2016. Youtube-8m: A large-scale video classification benchmark. arXiv preprint arXiv:1609.08675 ( 2016 ). Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. 2016. Youtube-8m: A large-scale video classification benchmark. arXiv preprint arXiv:1609.08675 (2016)."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2021-1710"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-434"},{"key":"e_1_3_2_2_4_1","volume-title":"Lin (Eds.)","volume":"33","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski , Yuhao Zhou , Abdelrahman Mohamed , and Michael Auli . 2020 . wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H . Lin (Eds.) , Vol. 33 . Curran Associates, Inc., 12449--12460. https:\/\/proceedings.neurips.cc\/paper\/ 2020\/file\/ 92d1e1eb1cd6f9fba3227870bb6d7f07-Paper.pdf Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 12449--12460. https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/ 92d1e1eb1cd6f9fba3227870bb6d7f07-Paper.pdf"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58323-1_42"},{"key":"e_1_3_2_2_6_1","volume-title":"Florian H\u00f6nig, Elmar N\u00f6th, and Korbinian Riedhammer.","author":"Bayerl Sebastian P","year":"2022","unstructured":"Sebastian P Bayerl , Alexander Wolff von Gudenberg , Florian H\u00f6nig, Elmar N\u00f6th, and Korbinian Riedhammer. 2022 . KSoF: The Kassel State of Fluency Dataset--A Therapy Centered Dataset of Stuttering . arXiv preprint arXiv:2203.05383 (2022). Sebastian P Bayerl, Alexander Wolff von Gudenberg, Florian H\u00f6nig, Elmar N\u00f6th, and Korbinian Riedhammer. 2022. KSoF: The Kassel State of Fluency Dataset--A Therapy Centered Dataset of Stuttering. arXiv preprint arXiv:2203.05383 (2022)."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"crossref","unstructured":"Gianmarco Cerutti Rahul Prasad Alessio Brutti and Elisabetta Farella. 2019. Neural Network Distillation on IoT Platforms for Sound Event Detection.. In Interspeech. 3609--3613.  Gianmarco Cerutti Rahul Prasad Alessio Brutti and Elisabetta Farella. 2019. Neural Network Distillation on IoT Platforms for Sound Event Detection.. In Interspeech. 3609--3613.","DOI":"10.21437\/Interspeech.2019-2394"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1136\/bmjinnov-2021-000668"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2021-1280"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2022-10103"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8682553"},{"key":"e_1_3_2_2_12_1","volume-title":"Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al.","author":"Hershey Shawn","year":"2017","unstructured":"Shawn Hershey , Sourish Chaudhuri , Daniel PW Ellis , Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al. 2017 . CNN architectures for large-scale audio classification. In 2017 ieee international conference on acoustics, speech and signal processing (icassp). IEEE , 131--135. Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al. 2017. CNN architectures for large-scale audio classification. In 2017 ieee international conference on acoustics, speech and signal processing (icassp). IEEE, 131--135."},{"key":"e_1_3_2_2_13_1","volume-title":"The paradoxical role of emotional intensity in the perception of vocal affect. Scientific reports 11, 1","author":"Holz Natalie","year":"2021","unstructured":"Natalie Holz , Pauline Larrouy-Maestri , and David Poeppel . 2021. The paradoxical role of emotional intensity in the perception of vocal affect. Scientific reports 11, 1 ( 2021 ), 1--10. Natalie Holz, Pauline Larrouy-Maestri, and David Poeppel. 2021. The paradoxical role of emotional intensity in the perception of vocal affect. Scientific reports 11, 1 (2021), 1--10."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1037\/emo0001048"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2021-703"},{"key":"e_1_3_2_2_16_1","volume-title":"Proceedings ACM Multimedia","author":"Schuller Bj\u00f6rn W.","year":"2022","unstructured":"Bj\u00f6rn W. Schuller , Anton Batliner , Shahin Amiriparian , Christian Bergler , Maurice Gerczuk , Natalie Holz , Pauline Larrouy-Maestri , Sebastian P. Bayerl , Korbinian Riedhammer , Adria Mallol-Ragolta , Maria Pateraki , Harry Coppock , Ivan Kiskin , Marianne Sinka , and Stephen Roberts . 2022 . The ACM Multimedia 2022 Computa- tional Paralinguistics Challenge: Vocalisations, Stuttering, Activity, & Mosquitos . In Proceedings ACM Multimedia 2022. ISCA, Lisbon, Portugal. to appear. Bj\u00f6rn W. Schuller, Anton Batliner, Shahin Amiriparian, Christian Bergler, Maurice Gerczuk, Natalie Holz, Pauline Larrouy-Maestri, Sebastian P. Bayerl, Korbinian Riedhammer, Adria Mallol-Ragolta, Maria Pateraki, Harry Coppock, Ivan Kiskin, Marianne Sinka, and Stephen Roberts. 2022. The ACM Multimedia 2022 Computa- tional Paralinguistics Challenge: Vocalisations, Stuttering, Activity, & Mosquitos. In Proceedings ACM Multimedia 2022. ISCA, Lisbon, Portugal. to appear."},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-2857"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3065234"},{"key":"e_1_3_2_2_19_1","volume-title":"Applying wav2vec2.0 to Speech Recognition in various low-resource languages. ArXiv abs\/2012.12121","author":"Yi Cheng","year":"2020","unstructured":"Cheng Yi , Jianzhong Wang , Ning Cheng , Shiyu Zhou , and Bo Xu. 2020. Applying wav2vec2.0 to Speech Recognition in various low-resource languages. ArXiv abs\/2012.12121 ( 2020 ). Cheng Yi, Jianzhong Wang, Ning Cheng, Shiyu Zhou, and Bo Xu. 2020. Applying wav2vec2.0 to Speech Recognition in various low-resource languages. ArXiv abs\/2012.12121 (2020)."},{"key":"e_1_3_2_2_20_1","volume-title":"How Does Pre-trained Wav2Vec2. 0 Perform on Domain Shifted ASR? An Extensive Benchmark on Air Traffic Control Communications. arXiv preprint arXiv:2203.16822","author":"Zuluaga-Gomez Juan","year":"2022","unstructured":"Juan Zuluaga-Gomez , Amrutha Prasad , Iuliia Nigmatulina , Saeed Sarfjoo , Petr Motlicek , Matthias Kleinert , Hartmut Helmke , Oliver Ohneiser , and Qingran Zhan . 2022. How Does Pre-trained Wav2Vec2. 0 Perform on Domain Shifted ASR? An Extensive Benchmark on Air Traffic Control Communications. arXiv preprint arXiv:2203.16822 ( 2022 ). Juan Zuluaga-Gomez, Amrutha Prasad, Iuliia Nigmatulina, Saeed Sarfjoo, Petr Motlicek, Matthias Kleinert, Hartmut Helmke, Oliver Ohneiser, and Qingran Zhan. 2022. How Does Pre-trained Wav2Vec2. 0 Perform on Domain Shifted ASR? An Extensive Benchmark on Air Traffic Control Communications. arXiv preprint arXiv:2203.16822 (2022)."}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3551572","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3551572","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:18Z","timestamp":1750182558000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3551572"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":20,"alternative-id":["10.1145\/3503161.3551572","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3551572","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}