{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,11]],"date-time":"2025-12-11T07:43:25Z","timestamp":1765439005196,"version":"3.44.0"},"publisher-location":"New York, NY, USA","reference-count":26,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,10,28]],"date-time":"2024-10-28T00:00:00Z","timestamp":1730073600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100006374","name":"Business Finland","doi-asserted-by":"publisher","award":["7817\/31\/2022"],"award-info":[{"award-number":["7817\/31\/2022"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Aalto ScienceIT"},{"name":"Finnish Center for Artificial Intelligence (FCAI)"},{"DOI":"10.13039\/501100006374","name":"Academy of Finland","doi-asserted-by":"publisher","award":["345790"],"award-info":[{"award-number":["345790"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,10,28]]},"DOI":"10.1145\/3689062.3689376","type":"proceedings-article","created":{"date-parts":[[2024,10,23]],"date-time":"2024-10-23T18:29:19Z","timestamp":1729708159000},"page":"60-64","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Multimodal Humor Detection and Social Perception Prediction"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-9378-3117","authenticated-orcid":false,"given":"Mehedi Hasan","family":"Bijoy","sequence":"first","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7219-9042","authenticated-orcid":false,"given":"Dejan","family":"Porjazovski","sequence":"additional","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2040-9834","authenticated-orcid":false,"given":"Nhan","family":"Phan","sequence":"additional","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5880-7956","authenticated-orcid":false,"given":"Guangpu","family":"Huang","sequence":"additional","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7918-9579","authenticated-orcid":false,"given":"Tam\u00e1s","family":"Gr\u00f3sz","sequence":"additional","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5278-7974","authenticated-orcid":false,"given":"Mikko","family":"Kurimo","sequence":"additional","affiliation":[{"name":"Aalto University, Espoo, Finland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,10,28]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","unstructured":"Shahin Amiriparian Lukas Christ Alexander Kathan Maurice Gerczuk Niklas M\u00fcller Steffen Klug Lukas Stappen Andreas K\u00f6nig Erik Cambria Bj\u00f6rn Schuller et al. 2024. The MuSe 2024 Multimodal Sentiment Analysis Challenge: Social Perception and Humor Recognition. arXiv preprint arXiv:2406.07753 (2024). https:\/\/doi.org\/10.48550\/arXiv.2406.07753","DOI":"10.48550\/arXiv.2406.07753"},{"key":"e_1_3_2_1_2_1","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: a framework for self-supervised learning of speech representations. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS '20). Curran Associates Inc., Red Hook, NY, USA, Article 1044, 12 pages."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-809633-8.20349-X"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.3390\/asi5040080"},{"key":"e_1_3_2_1_5_1","volume-title":"Attention-based models for speech recognition. Advances in neural information processing systems","author":"Chorowski Jan K","year":"2015","unstructured":"Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015. Attention-based models for speech recognition. Advances in neural information processing systems, Vol. 28 (2015)."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3606039.3613114"},{"key":"e_1_3_2_1_7_1","volume-title":"Schuller","author":"Christ Lukas","year":"2024","unstructured":"Lukas Christ, Shahin Amiriparian, Alexander Kathan, Niklas M\u00fcller, Andreas K\u00f6nig, and Bj\u00f6rn W. Schuller. 2024. Towards Multimodal Prediction of Spontaneous Humour: A Novel Dataset and First Results. arxiv: 2209.14272 [cs.LG] https:\/\/arxiv.org\/abs\/2209.14272"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR56361.2022.9956592"},{"key":"e_1_3_2_1_9_1","volume-title":"Unsupervised Cross-lingual Representation Learning at Scale. CoRR","author":"Conneau Alexis","year":"2019","unstructured":"Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm\u00e1n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised Cross-lingual Representation Learning at Scale. CoRR, Vol. abs\/1911.02116 (2019). showeprint[arXiv]1911.02116 http:\/\/arxiv.org\/abs\/1911.02116"},{"key":"e_1_3_2_1_10_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. CoRR","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. CoRR, Vol. abs\/1810.04805 (2018). arxiv: 1810.04805 http:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1142\/S0218488598000094"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053762"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"crossref","unstructured":"DN Krishna and Ankita Patil. 2020. Multimodal Emotion Recognition Using Cross-Modal Attention and 1D Convolutional Neural Networks.. In Interspeech. 4243--4247.","DOI":"10.21437\/Interspeech.2020-1190"},{"key":"e_1_3_2_1_14_1","unstructured":"Trent W Lewis and David MW Powers. 2004. Sensor fusion weighting measures in audio-visual speech recognition. In ACSC. Citeseer 305--314."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU57964.2023.10389642"},{"key":"e_1_3_2_1_16_1","volume-title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR, Vol. abs\/1907.11692 (2019). arxiv: 1907.11692 http:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3271019"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2022.105667"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Michail Mitsios Georgios Vamvoukakis Georgia Maniati Nikolaos Ellinas Georgios Dimitriou Konstantinos Markopoulos Panos Kakoulidis Alexandra Vioni Myrsini Christidou Junkwang Oh et al. 2024. Improved Text Emotion Prediction Using Combined Valence and Arousal Ordinal Classification. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers). Association for Computational Linguistics Mexico City Mexico 808--813. https:\/\/aclanthology.org\/2024.naacl-short.72","DOI":"10.18653\/v1\/2024.naacl-short.72"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007468413059"},{"key":"e_1_3_2_1_21_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever et al. 2019. Language models are unsupervised multitask learners. OpenAI blog Vol. 1 8 (2019) 9."},{"key":"e_1_3_2_1_22_1","volume-title":"International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Z1Qlm11uOM","author":"Shi Bowen","year":"2022","unstructured":"Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed. 2022. Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Z1Qlm11uOM"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.359"},{"key":"e_1_3_2_1_24_1","volume-title":"Proceedings of the conference. Association for computational linguistics. Meeting","volume":"2019","author":"Hubert Tsai Yao-Hung","year":"2019","unstructured":"Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the conference. Association for computational linguistics. Meeting, Vol. 2019. NIH Public Access, 6558."},{"key":"e_1_3_2_1_25_1","volume-title":"Attention is all you need. Advances in neural information processing systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, Vol. 30 (2017)."},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2021--1775"}],"event":{"name":"MM '24: The 32nd ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Melbourne VIC Australia","acronym":"MM '24"},"container-title":["Proceedings of the 5th on Multimodal Sentiment Analysis Challenge and Workshop: Social Perception and Humor"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3689062.3689376","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3689062.3689376","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T18:23:53Z","timestamp":1755973433000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3689062.3689376"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,28]]},"references-count":26,"alternative-id":["10.1145\/3689062.3689376","10.1145\/3689062"],"URL":"https:\/\/doi.org\/10.1145\/3689062.3689376","relation":{},"subject":[],"published":{"date-parts":[[2024,10,28]]},"assertion":[{"value":"2024-10-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}