{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:15:55Z","timestamp":1750220155917,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":64,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Swiss National Science Foundation","award":["EMIL"],"award-info":[{"award-number":["EMIL"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3551876.3554812","type":"proceedings-article","created":{"date-parts":[[2022,9,28]],"date-time":"2022-09-28T22:17:21Z","timestamp":1664403441000},"page":"37-45","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Comparing Biosignal and Acoustic feature Representation for Continuous Emotion Recognition"],"prefix":"10.1145","author":[{"given":"Sarthak","family":"Yadav","sequence":"first","affiliation":[{"name":"Idiap Research Institute, Martigny, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tilak","family":"Purohit","sequence":"additional","affiliation":[{"name":"Idiap Research Institute &amp; \u00c9cole polytechnique f\u00e9d\u00e9rale de Lausanne (EPFL), Martigny, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zohreh","family":"Mostaani","sequence":"additional","affiliation":[{"name":"Idiap Research Institute &amp; \u00c9cole polytechnique f\u00e9d\u00e9rale de Lausanne (EPFL), Martigny, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bogdan","family":"Vlasenko","sequence":"additional","affiliation":[{"name":"Idiap Research Institute, Martigny, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mathew","family":"Magimai.-Doss","sequence":"additional","affiliation":[{"name":"Idiap Research Institute, Martigny, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2017.2764438"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414866"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3129340"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3423327.3423673"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2020.01.011"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-3131"},{"key":"e_1_3_2_2_7_1","first-page":"384","article-title":"Multi-Modal Embeddings Using Multi-Task Learning for Emotion Recognition","volume":"2020","author":"Khare Aparna","year":"2020","unstructured":"Aparna Khare , Srinivas Parthasarathy , and Shiva Sundaram . 2020 . Multi-Modal Embeddings Using Multi-Task Learning for Emotion Recognition . In Proc. Interspeech 2020 , 384 -- 388 . doi: 10.21437\/Interspeech.2020--1827. 10.21437\/Interspeech.2020--1827 Aparna Khare, Srinivas Parthasarathy, and Shiva Sundaram. 2020. Multi-Modal Embeddings Using Multi-Task Learning for Emotion Recognition. In Proc. Interspeech 2020, 384--388. doi: 10.21437\/Interspeech.2020--1827.","journal-title":"Proc. Interspeech"},{"volume-title":"Proceedings of the 2020 International Conference on Multimodal Interaction. Association for Computing Machinery","year":"2020","key":"e_1_3_2_2_8_1","unstructured":"2020. Emotiw 2020 : driver gaze, group emotion, student engagement and physiological signal based challenges . Proceedings of the 2020 International Conference on Multimodal Interaction. Association for Computing Machinery , New York, NY, USA, 784--789. isbn: 9781450375818. https: \/ \/ doi . org \/10 .1145 \/ 3382507.3417973. 2020. Emotiw 2020: driver gaze, group emotion, student engagement and physiological signal based challenges. Proceedings of the 2020 International Conference on Multimodal Interaction. Association for Computing Machinery, New York, NY, USA, 784--789. isbn: 9781450375818. https: \/ \/ doi . org \/10 .1145 \/ 3382507.3417973."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1037\/h0077714"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3475957.3484446"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU.2005.1566530"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/MMSP.2019.8901758"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/0167-8760(94)90027-2"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2011.6012003"},{"key":"e_1_3_2_2_15_1","volume-title":"Articulation constrained learning with application to speech emotion recognition. EURASIP journal on audio, speech, and music processing","author":"Shah Mohit","year":"2019","unstructured":"Mohit Shah , Ming Tu , Visar Berisha , Chaitali Chakrabarti , and Andreas Spanias . 2019. Articulation constrained learning with application to speech emotion recognition. EURASIP journal on audio, speech, and music processing , 2019 , 1, 1--17. Mohit Shah, Ming Tu, Visar Berisha, Chaitali Chakrabarti, and Andreas Spanias. 2019. Articulation constrained learning with application to speech emotion recognition. EURASIP journal on audio, speech, and music processing, 2019, 1, 1--17."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2008.4607689"},{"key":"e_1_3_2_2_17_1","unstructured":"Jiahong Yuan Xingyu Cai Renjie Zheng Liang Huang and Kenneth Church. 2021. The role of phonetic units in speech emotion recognition. arXiv preprint arXiv:2108.01132. Jiahong Yuan Xingyu Cai Renjie Zheng Liang Huang and Kenneth Church. 2021. The role of phonetic units in speech emotion recognition. arXiv preprint arXiv:2108.01132."},{"key":"e_1_3_2_2_18_1","volume-title":"A waveform-feature dual branch acoustic embedding network for emotion recognition. Frontiers in Computer Science, 2. issn: 2624--9898. doi: 10.3389\/fcomp","author":"Li Jeng-Lin","year":"2020","unstructured":"Jeng-Lin Li , Tzu-Yun Huang , Chun-Min Chang , and ChiChun Lee . 2020. A waveform-feature dual branch acoustic embedding network for emotion recognition. Frontiers in Computer Science, 2. issn: 2624--9898. doi: 10.3389\/fcomp . 2020 .00013. https:\/\/www.frontiersin.org\/article\/10.3389\/ fcomp.2020.00013. 10.3389\/fcomp Jeng-Lin Li, Tzu-Yun Huang, Chun-Min Chang, and ChiChun Lee. 2020. A waveform-feature dual branch acoustic embedding network for emotion recognition. Frontiers in Computer Science, 2. issn: 2624--9898. doi: 10.3389\/fcomp. 2020.00013. https:\/\/www.frontiersin.org\/article\/10.3389\/ fcomp.2020.00013."},{"key":"e_1_3_2_2_19_1","first-page":"2182","article-title":"A comparison of acoustic and linguistics methodologies for alzheimer's dementia recognition","volume":"2020","author":"N Cummins","year":"2020","unstructured":"N Cummins et al. 2020 . A comparison of acoustic and linguistics methodologies for alzheimer's dementia recognition . Proceedngs of Interspeech 2020 , 2182 -- 2186 . N Cummins et al. 2020. A comparison of acoustic and linguistics methodologies for alzheimer's dementia recognition. Proceedngs of Interspeech 2020, 2182--2186.","journal-title":"Proceedngs of Interspeech"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"crossref","unstructured":"Carlos Busso Murtaza Bulut Chi-Chun Lee Abe Kazemzadeh Emily Mower Samuel Kim Jeannette N Chang Sungbok Lee and Shrikanth S Narayanan. 2008. Iemocap: interactive emotional dyadic motion capture database. Language resources and evaluation 42 4 335--359. Carlos Busso Murtaza Bulut Chi-Chun Lee Abe Kazemzadeh Emily Mower Samuel Kim Jeannette N Chang Sungbok Lee and Shrikanth S Narayanan. 2008. Iemocap: interactive emotional dyadic motion capture database. Language resources and evaluation 42 4 335--359.","DOI":"10.1007\/s10579-008-9076-6"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2008.4517932"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSLP.1996.607793"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2010.09.020"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0301-0511(98)00025-8"},{"key":"e_1_3_2_2_25_1","unstructured":"Qiang Zhang Xianxiang Chen Qingyuan Zhan Ting Yang and Shanhong Xia. 2017. Respiration-based emotion recognition with deep learning. Computers in Industry 92--93 84--90. issn: 0166--3615. doi: https:\/\/doi.org\/10.1016\/j.compind.2017. 04.005. https:\/\/www.sciencedirect.com\/science\/article\/pii\/ S0166361516303104. 10.1016\/j.compind.2017 Qiang Zhang Xianxiang Chen Qingyuan Zhan Ting Yang and Shanhong Xia. 2017. Respiration-based emotion recognition with deep learning. Computers in Industry 92--93 84--90. issn: 0166--3615. doi: https:\/\/doi.org\/10.1016\/j.compind.2017. 04.005. https:\/\/www.sciencedirect.com\/science\/article\/pii\/ S0166361516303104."},{"key":"e_1_3_2_2_26_1","volume-title":"Hewitt","author":"MacLarnon Ann","year":"1999","unstructured":"Ann MacLarnon and Gwen P . Hewitt . 1999 . The evolution of human speech: the role of enhanced breathing control. American journal of physical anthropology, 109, 341--63. doi: 10.1002\/(SICI)1096--8644(199907)109:3<341::AID-AJPA5>3. 0.CO;2--2. 10.1002\/(SICI)1096--8644(199907)109:3<341::AID-AJPA5>3 Ann MacLarnon and Gwen P. Hewitt. 1999. The evolution of human speech: the role of enhanced breathing control. American journal of physical anthropology, 109, 341--63. doi: 10.1002\/(SICI)1096--8644(199907)109:3<341::AID-AJPA5>3. 0.CO;2--2."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053753"},{"key":"e_1_3_2_2_28_1","volume-title":"The INTERSPEECH 2020 Computational Paralinguistics Challenge: Elderly Emotion, Breathing & Masks. In Proc. Interspeech 2020","author":"Schuller Bj\u00f6rn W.","year":"2020","unstructured":"Bj\u00f6rn W. Schuller , Anton Batliner , Christian Bergler , EvaMaria Messner , Antonia Hamilton , Shahin Amiriparian , Alice Baird , Georgios Rizos , Maximilian Schmitt , Lukas Stappen , Harald Baumeister , Alexis Deighton MacIntyre , and Simone Hantke . 2020 . The INTERSPEECH 2020 Computational Paralinguistics Challenge: Elderly Emotion, Breathing & Masks. In Proc. Interspeech 2020 , 2042--2046. doi: 10.21437\/ Interspeech.2020--32. Bj\u00f6rn W. Schuller, Anton Batliner, Christian Bergler, EvaMaria Messner, Antonia Hamilton, Shahin Amiriparian, Alice Baird, Georgios Rizos, Maximilian Schmitt, Lukas Stappen, Harald Baumeister, Alexis Deighton MacIntyre, and Simone Hantke. 2020. The INTERSPEECH 2020 Computational Paralinguistics Challenge: Elderly Emotion, Breathing & Masks. In Proc. Interspeech 2020, 2042--2046. doi: 10.21437\/ Interspeech.2020--32."},{"key":"e_1_3_2_2_29_1","first-page":"2072","article-title":"Ensembling End-to-End Deep Models for Computational Paralinguistics Tasks: ComParE 2020 Mask and Breathing Sub-Challenges","volume":"2020","author":"Markitantov Maxim","year":"2020","unstructured":"Maxim Markitantov , Denis Dresvyanskiy , Danila Mamontov , Heysem Kaya , Wolfgang Minker , and Alexey Karpov . 2020 . Ensembling End-to-End Deep Models for Computational Paralinguistics Tasks: ComParE 2020 Mask and Breathing Sub-Challenges . In Proc. Interspeech 2020 , 2072 -- 2076 . doi: 10.21437\/Interspeech.2020--2666. 10.21437\/Interspeech.2020--2666 Maxim Markitantov, Denis Dresvyanskiy, Danila Mamontov, Heysem Kaya, Wolfgang Minker, and Alexey Karpov. 2020. Ensembling End-to-End Deep Models for Computational Paralinguistics Tasks: ComParE 2020 Mask and Breathing Sub-Challenges. In Proc. Interspeech 2020, 2072--2076. doi: 10.21437\/Interspeech.2020--2666.","journal-title":"Proc. Interspeech"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2021.03.029"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414756"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746271"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2022-10271"},{"key":"e_1_3_2_2_34_1","unstructured":"Aaron van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748. Aaron van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748."},{"key":"e_1_3_2_2_35_1","first-page":"12449","article-title":"Wav2vec 2.0: a framework for self-supervised learning of speech representations","volume":"33","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski , Yuhao Zhou , Abdelrahman Mohamed , and Michael Auli . 2020 . Wav2vec 2.0: a framework for self-supervised learning of speech representations . In Advances in Neural Information Processing Systems. Volume 33 , 12449 -- 12460 . Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. Wav2vec 2.0: a framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems. Volume 33, 12449--12460.","journal-title":"Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2020.2985586"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN52387.2021.9534474"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Luyu Wang Pauline Luc Yan Wu Adria Recasens Lucas Smaira Andrew Brock Andrew Jaegle Jean-Baptiste Alayrac Sander Dieleman Joao Carreira etal 2021. Towards learning universal audio representations. arXiv preprint arXiv:2111.12124. Luyu Wang Pauline Luc Yan Wu Adria Recasens Lucas Smaira Andrew Brock Andrew Jaegle Jean-Baptiste Alayrac Sander Dieleman Joao Carreira et al. 2021. Towards learning universal audio representations. arXiv preprint arXiv:2111.12124.","DOI":"10.1109\/ICASSP43922.2022.9746790"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413528"},{"key":"e_1_3_2_2_41_1","unstructured":"Daisuke Niizumi Daiki Takeuchi Yasunori Ohishi Noboru Harada and Kunio Kashino. 2022. Masked spectrogram modeling using masked autoencoders for learning general-purpose audio representation. arXiv preprint arXiv:2204.12260. Daisuke Niizumi Daiki Takeuchi Yasunori Ohishi Noboru Harada and Kunio Kashino. 2022. Masked spectrogram modeling using masked autoencoders for learning general-purpose audio representation. arXiv preprint arXiv:2204.12260."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Dading Chong Helin Wang Peilin Zhou and Qingcheng Zeng. 2022. Masked spectrogram prediction for self-supervised audio pre-training. arXiv preprint arXiv:2204.12768. Dading Chong Helin Wang Peilin Zhou and Qingcheng Zeng. 2022. Masked spectrogram prediction for self-supervised audio pre-training. arXiv preprint arXiv:2204.12768.","DOI":"10.1109\/ICASSP49357.2023.10095691"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"crossref","unstructured":"Sarthak Yadav and Neil Zeghidour. 2022. Learning neural audio features without supervision. arXiv preprint arXiv:2203.15519. Sarthak Yadav and Neil Zeghidour. 2022. Learning neural audio features without supervision. arXiv preprint arXiv:2203.15519.","DOI":"10.21437\/Interspeech.2022-10834"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"e_1_3_2_2_45_1","volume-title":"International conference on machine learning. PMLR, 6105--6114","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le . 2019 . Efficientnet: rethinking model scaling for convolutional neural networks . In International conference on machine learning. PMLR, 6105--6114 . Mingxing Tan and Quoc Le. 2019. Efficientnet: rethinking model scaling for convolutional neural networks. In International conference on machine learning. PMLR, 6105--6114."},{"key":"e_1_3_2_2_46_1","volume-title":"audax. Version 0.x.x. (February","author":"Yadav Sarthak","year":"2022","unstructured":"Sarthak Yadav . 2022. audax. Version 0.x.x. (February 2022 ). https:\/\/github.com\/SarthakYadav\/audax. Sarthak Yadav. 2022. audax. Version 0.x.x. (February 2022). https:\/\/github.com\/SarthakYadav\/audax."},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"crossref","unstructured":"Sanyuan Chen et al. 2021. Wavlm: large-scale self-supervised pre-training for full stack speech processing. ArXiv. Sanyuan Chen et al. 2021. Wavlm: large-scale self-supervised pre-training for full stack speech processing. ArXiv.","DOI":"10.1109\/JSTSP.2022.3188113"},{"key":"e_1_3_2_2_49_1","unstructured":"Shu-wen Yang et al. 2021. Superb: speech processing universal performance benchmark. arXiv preprint arXiv:2105.01051. Shu-wen Yang et al. 2021. Superb: speech processing universal performance benchmark. arXiv preprint arXiv:2105.01051."},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"crossref","unstructured":"Dimitri Palaz Ronan Collobert and Mathew Magimai-Doss. 2013. Estimating phoneme class conditional probabilities from raw speech signal using convolutional neural networks. In INTERSPEECH. Dimitri Palaz Ronan Collobert and Mathew Magimai-Doss. 2013. Estimating phoneme class conditional probabilities from raw speech signal using convolutional neural networks. In INTERSPEECH.","DOI":"10.21437\/Interspeech.2013-438"},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"crossref","unstructured":"Tara Sainath Ron J Weiss Kevin Wilson Andrew W Senior and Oriol Vinyals. 2015. Learning the speech front-end with raw waveform cldnns. Tara Sainath Ron J Weiss Kevin Wilson Andrew W Senior and Oriol Vinyals. 2015. Learning the speech front-end with raw waveform cldnns.","DOI":"10.21437\/Interspeech.2015-1"},{"key":"e_1_3_2_2_52_1","unstructured":"Ronan Collobert Christian Puhrsch and Gabriel Synnaeve. 2016. Wav2letter: an end-to-end convnet-based speech recognition system. arXiv preprint arXiv:1609.03193. Ronan Collobert Christian Puhrsch and Gabriel Synnaeve. 2016. Wav2letter: an end-to-end convnet-based speech recognition system. arXiv preprint arXiv:1609.03193."},{"key":"e_1_3_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8462165"},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/SLT.2018.8639585"},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"crossref","unstructured":"Selen Hande Kabil Hannah Muckenhirn and Mathew MagimaiDoss. 2018. On learning to identify genders from raw speech signal using cnns. In Interspeech 287--291 Selen Hande Kabil Hannah Muckenhirn and Mathew MagimaiDoss. 2018. On learning to identify genders from raw speech signal using cnns. In Interspeech 287--291","DOI":"10.21437\/Interspeech.2018-1240"},{"key":"e_1_3_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683498"},{"key":"e_1_3_2_2_57_1","volume-title":"International Conference on Learning Representations. https : \/ \/ openreview. net \/ forum ?id = jM76BCb6F9m.","author":"Zeghidour Neil","year":"2021","unstructured":"Neil Zeghidour , Olivier Teboul , F\u00e9lix de Chaumont Quitry , and Marco Tagliasacchi . 2021 . {leaf}: a learnable frontend for audio classification . In International Conference on Learning Representations. https : \/ \/ openreview. net \/ forum ?id = jM76BCb6F9m. Neil Zeghidour, Olivier Teboul, F\u00e9lix de Chaumont Quitry, and Marco Tagliasacchi. 2021. {leaf}: a learnable frontend for audio classification. In International Conference on Learning Representations. https : \/ \/ openreview. net \/ forum ?id = jM76BCb6F9m."},{"key":"e_1_3_2_2_58_1","volume-title":"Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15, 1","author":"Srivastava Nitish","year":"1929","unstructured":"Nitish Srivastava , Geoffrey Hinton , Alex Krizhevsky , Ilya Sutskever , and Ruslan Salakhutdinov . 2014. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15, 1 , 1929 --1958. Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15, 1, 1929--1958."},{"key":"e_1_3_2_2_59_1","volume-title":"Proceedings of the 3rd Multimodal Sentiment Analysis Challenge. Workshop held at ACM Multimedia","author":"Christ Lukas","year":"2022","unstructured":"Lukas Christ , Shahin Amiriparian , Alice Baird , Panagiotis Tzirakis , Alexander Kathan , Niklas M\u00fcller , Lukas Stappen , Eva-Maria Me\u00dfner , Andreas K\u00f6nig , Alan Cowen , Erik Cambria , and Bj\u00f6rn W. Schuller . 2022. The muse 2022 multimodal sentiment analysis challenge: humor, emotional reactions, and stress . In Proceedings of the 3rd Multimodal Sentiment Analysis Challenge. Workshop held at ACM Multimedia 2022 , to appear. Association for Computing Machinery, Lisbon, Portugal. Lukas Christ, Shahin Amiriparian, Alice Baird, Panagiotis Tzirakis, Alexander Kathan, Niklas M\u00fcller, Lukas Stappen, Eva-Maria Me\u00dfner, Andreas K\u00f6nig, Alan Cowen, Erik Cambria, and Bj\u00f6rn W. Schuller. 2022. The muse 2022 multimodal sentiment analysis challenge: humor, emotional reactions, and stress. In Proceedings of the 3rd Multimodal Sentiment Analysis Challenge. Workshop held at ACM Multimedia 2022, to appear. Association for Computing Machinery, Lisbon, Portugal."},{"key":"e_1_3_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1159\/000119004"},{"key":"e_1_3_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-434"},{"key":"e_1_3_2_2_62_1","volume-title":"International Conference on Learning Representations. https : \/ \/ openreview. net \/ forum ?id = Bkg6RiCqY7.","author":"Loshchilov Ilya","year":"2019","unstructured":"Ilya Loshchilov and Frank Hutter . 2019 . Decoupled weight decay regularization . In International Conference on Learning Representations. https : \/ \/ openreview. net \/ forum ?id = Bkg6RiCqY7. Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations. https : \/ \/ openreview. net \/ forum ?id = Bkg6RiCqY7."},{"key":"e_1_3_2_2_63_1","volume-title":"Proceedings of the ICML Expressive Vocalizations Workshop. Workshop held in conjunction with the 39th International Conference on Machine Learning.","author":"Purohit Tilak","year":"2022","unstructured":"Tilak Purohit , Imen Ben Mahmoud , Bogdan Vlasenko , and Mathew Magimai Doss . 2022 . Comparing supervised and selfsupervised embedding for exvo multi-task learning track . Proceedings of the ICML Expressive Vocalizations Workshop. Workshop held in conjunction with the 39th International Conference on Machine Learning. Tilak Purohit, Imen Ben Mahmoud, Bogdan Vlasenko, and Mathew Magimai Doss. 2022. Comparing supervised and selfsupervised embedding for exvo multi-task learning track. Proceedings of the ICML Expressive Vocalizations Workshop. Workshop held in conjunction with the 39th International Conference on Machine Learning."},{"key":"e_1_3_2_2_64_1","doi-asserted-by":"crossref","unstructured":"Alice Baird et al. 2022. The icml 2022 expressive vocalizations workshop and competition: recognizing generating and personalizing vocal bursts. (2022). doi: 10.48550 \/ARXIV. 2205.01780 Alice Baird et al. 2022. The icml 2022 expressive vocalizations workshop and competition: recognizing generating and personalizing vocal bursts. (2022). doi: 10.48550 \/ARXIV. 2205.01780","DOI":"10.1109\/ACIIW57231.2022.10086002"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Lisboa Portugal","acronym":"MM '22"},"container-title":["Proceedings of the 3rd International on Multimodal Sentiment Analysis Workshop and Challenge"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3551876.3554812","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3551876.3554812","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:17Z","timestamp":1750186817000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3551876.3554812"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":64,"alternative-id":["10.1145\/3551876.3554812","10.1145\/3551876"],"URL":"https:\/\/doi.org\/10.1145\/3551876.3554812","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}