{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T02:38:45Z","timestamp":1783564725609,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":47,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,10,21]],"date-time":"2023-10-21T00:00:00Z","timestamp":1697846400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"NSFC","award":["62206193,62076176,62076175"],"award-info":[{"award-number":["62206193,62076176,62076175"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,10,21]]},"DOI":"10.1145\/3583780.3615024","type":"proceedings-article","created":{"date-parts":[[2023,10,21]],"date-time":"2023-10-21T07:45:42Z","timestamp":1697874342000},"page":"1045-1055","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Real-time Emotion Pre-Recognition in Conversations with Contrastive Multi-modal Dialogue Pre-training"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-6183-8096","authenticated-orcid":false,"given":"Xincheng","family":"Ju","sequence":"first","affiliation":[{"name":"Soochow University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8948-2856","authenticated-orcid":false,"given":"Dong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Soochow University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1529-4278","authenticated-orcid":false,"given":"Suyang","family":"Zhu","sequence":"additional","affiliation":[{"name":"Soochow University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7829-6348","authenticated-orcid":false,"given":"Junhui","family":"Li","sequence":"additional","affiliation":[{"name":"Soochow University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1000-3278","authenticated-orcid":false,"given":"Shoushan","family":"Li","sequence":"additional","affiliation":[{"name":"Soochow University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7887-5099","authenticated-orcid":false,"given":"Guodong","family":"Zhou","sequence":"additional","affiliation":[{"name":"Soochow University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,10,21]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3049732"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-008-9076-6"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.55"},{"key":"e_1_3_2_1_4_1","first-page":"1064","article-title":"Learning What and When to Drop","volume":"2021","author":"Chen Feiyu","year":"2021","unstructured":"Feiyu Chen , Zhengxiao Sun , Deqiang Ouyang , Xueliang Liu , and Jie Shao . 2021 . Learning What and When to Drop : Adaptive Multimodal and Contextual Dynamics for Emotion Recognition in Conversation. In Proceedings of ACM MM 2021. 1064 -- 1073 . https:\/\/doi.org\/10.1145\/3474085.3475661 10.1145\/3474085.3475661 Feiyu Chen, Zhengxiao Sun, Deqiang Ouyang, Xueliang Liu, and Jie Shao. 2021. Learning What and When to Drop: Adaptive Multimodal and Contextual Dynamics for Emotion Recognition in Conversation. In Proceedings of ACM MM 2021. 1064--1073. https:\/\/doi.org\/10.1145\/3474085.3475661","journal-title":"In Proceedings of ACM MM"},{"key":"e_1_3_2_1_5_1","first-page":"375","volume-title":"Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis. In Findings of EMNLP 2022","author":"Chu Iek-Heng","year":"2022","unstructured":"Iek-Heng Chu , Ziyi Chen , Xinlu Yu , Mei Han , Jing Xiao , and Peng Chang . 2022 . Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis. In Findings of EMNLP 2022 . Association for Computational Linguistics, 5105--5114. https:\/\/aclanthology.org\/ 2022.findings-emnlp. 375 Iek-Heng Chu, Ziyi Chen, Xinlu Yu, Mei Han, Jing Xiao, and Peng Chang. 2022. Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis. In Findings of EMNLP 2022. Association for Computational Linguistics, 5105--5114. https:\/\/aclanthology.org\/2022.findings-emnlp.375"},{"key":"e_1_3_2_1_6_1","volume-title":"Proceedings of NAACL-HLT 2019","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . In Proceedings of NAACL-HLT 2019 . Association for Computational Linguistics, 4171--4186. https:\/\/doi.org\/10. 18653\/v1\/n19--1423 10.18653\/v1 Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT 2019. Association for Computational Linguistics, 4171--4186. https:\/\/doi.org\/10.18653\/v1\/n19--1423"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.224"},{"key":"e_1_3_2_1_8_1","volume-title":"Proceedings of EMNLP-IJCNLP 2019","author":"Ghosal Deepanway","year":"1865","unstructured":"Deepanway Ghosal , Navonil Majumder , Soujanya Poria , Niyati Chhaya , and Alexander F. Gelbukh . 2019. DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation . In Proceedings of EMNLP-IJCNLP 2019 . Association for Computational Linguistics, 154--164. https:\/\/doi.org\/10. 1865 3\/v1\/D19--1015 10.18653\/v1 Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, and Alexander F. Gelbukh. 2019. DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation. In Proceedings of EMNLP-IJCNLP 2019. Association for Computational Linguistics, 154--164. https:\/\/doi.org\/10.18653\/v1\/D19--1015"},{"key":"e_1_3_2_1_9_1","volume-title":"Proceedings of ACL","author":"Hasegawa Takayuki","year":"2013","unstructured":"Takayuki Hasegawa , Nobuhiro Kaji , Naoki Yoshinaga , and Masashi Toyoda . 2013 . Predicting and Eliciting Addressee's Emotion in Online Dialogue . In Proceedings of ACL 2013. The Association for Computer Linguistics, 964--972. https:\/\/aclanthology.org\/P13--1095\/ Takayuki Hasegawa, Nobuhiro Kaji, Naoki Yoshinaga, and Masashi Toyoda. 2013. Predicting and Eliciting Addressee's Emotion in Online Dialogue. In Proceedings of ACL 2013. The Association for Computer Linguistics, 964--972. https:\/\/aclanthology.org\/P13--1095\/"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1280"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_3_2_1_12_1","volume-title":"MM-DFN: Multimodal Dynamic Fusion Network for Emotion Recognition in Conversations. CoRR","author":"Hu Dou","year":"2022","unstructured":"Dou Hu , Xiaolong Hou , Lingwei Wei , Lian-Xin Jiang , and Yang Mo. 2022. MM-DFN: Multimodal Dynamic Fusion Network for Emotion Recognition in Conversations. CoRR , Vol. abs\/ 2203 .02385 ( 2022 ). https:\/\/doi.org\/10.48550\/arXiv.2203.02385 showeprint[arXiv]2203.02385 10.48550\/arXiv.2203.02385 Dou Hu, Xiaolong Hou, Lingwei Wei, Lian-Xin Jiang, and Yang Mo. 2022. MM-DFN: Multimodal Dynamic Fusion Network for Emotion Recognition in Conversations. CoRR, Vol. abs\/2203.02385 (2022). https:\/\/doi.org\/10.48550\/arXiv.2203.02385 showeprint[arXiv]2203.02385"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.440"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.435"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6309"},{"key":"e_1_3_2_1_16_1","volume-title":"Proceedings of ICLR","author":"Diederik","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization . In Proceedings of ICLR 2015 . http:\/\/arxiv.org\/abs\/1412.6980 Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In Proceedings of ICLR 2015. http:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.416"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2021.3079263"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240575"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.749"},{"key":"e_1_3_2_1_22_1","volume-title":"Proceedings of LREC 2016","author":"Lison Pierre","year":"2016","unstructured":"Pierre Lison and J\u00f6 rg Tiedemann . 2016 . OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles . In Proceedings of LREC 2016 . European Language Resources Association (ELRA). http:\/\/www.lrec-conf.org\/proceedings\/lrec 2016\/summaries\/947.html Pierre Lison and J\u00f6 rg Tiedemann. 2016. OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. In Proceedings of LREC 2016. European Language Resources Association (ELRA). http:\/\/www.lrec-conf.org\/proceedings\/lrec2016\/summaries\/947.html"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11955"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2019.2900910"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33016818"},{"key":"e_1_3_2_1_26_1","volume-title":"Proceedings of NIPS 2013","author":"Mikolov Tom\u00e1","year":"2013","unstructured":"Tom\u00e1 s Mikolov , Ilya Sutskever , Kai Chen , Gregory S. Corrado , and Jeffrey Dean . 2013 . Distributed Representations of Words and Phrases and their Compositionality . In Proceedings of NIPS 2013 . 3111--3119. https:\/\/proceedings.neurips.cc\/paper\/2013\/hash\/9aa42b31882ec039965f3c4923ce901b-Abstract.html Tom\u00e1 s Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. In Proceedings of NIPS 2013. 3111--3119. https:\/\/proceedings.neurips.cc\/paper\/2013\/hash\/9aa42b31882ec039965f3c4923ce901b-Abstract.html"},{"key":"e_1_3_2_1_27_1","volume-title":"Proceedings of CogSci 2020","author":"Ong Desmond C.","year":"2020","unstructured":"Desmond C. Ong , Marie Therese Quieta , and Basura Fernando . 2020 . Intention Inference in a Dynamic Multi-Goal Environment . In Proceedings of CogSci 2020 . cognitivesciencesociety.org. https:\/\/cogsci.mindmodeling.org\/2020\/papers\/0392\/index.html Desmond C. Ong, Marie Therese Quieta, and Basura Fernando. 2020. Intention Inference in a Dynamic Multi-Goal Environment. In Proceedings of CogSci 2020. cognitivesciencesociety.org. https:\/\/cogsci.mindmodeling.org\/2020\/papers\/0392\/index.html"},{"key":"e_1_3_2_1_28_1","volume-title":"Proceedings of ICML 2013 (JMLR Workshop and Conference Proceedings","volume":"1318","author":"Pascanu Razvan","year":"2013","unstructured":"Razvan Pascanu , Tom\u00e1 s Mikolov , and Yoshua Bengio . 2013 . On the difficulty of training recurrent neural networks . In Proceedings of ICML 2013 (JMLR Workshop and Conference Proceedings , Vol. 28). JMLR.org, 1310-- 1318 . http:\/\/proceedings.mlr.press\/v28\/pascanu13.html Razvan Pascanu, Tom\u00e1 s Mikolov, and Yoshua Bengio. 2013. On the difficulty of training recurrent neural networks. In Proceedings of ICML 2013 (JMLR Workshop and Conference Proceedings, Vol. 28). JMLR.org, 1310--1318. http:\/\/proceedings.mlr.press\/v28\/pascanu13.html"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1081"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1050"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3053400"},{"key":"e_1_3_2_1_32_1","volume-title":"Proceedings of ICML 2021 (Proceedings of Machine Learning Research","volume":"8763","author":"Radford Alec","year":"2021","unstructured":"Alec Radford , Jong Wook Kim , Chris Hallacy , Aditya Ramesh , Gabriel Goh , Sandhini Agarwal , Girish Sastry , Amanda Askell , Pamela Mishkin , Jack Clark , Gretchen Krueger , and Ilya Sutskever . 2021 . Learning Transferable Visual Models From Natural Language Supervision . In Proceedings of ICML 2021 (Proceedings of Machine Learning Research , Vol. 139). PMLR, 8748-- 8763 . http:\/\/proceedings.mlr.press\/v139\/radford21a.html Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of ICML 2021 (Proceedings of Machine Learning Research, Vol. 139). PMLR, 8748--8763. http:\/\/proceedings.mlr.press\/v139\/radford21a.html"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3536221.3556601"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.semeval-1.133"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i15.17625"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.398"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-short.38"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2670313"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.221"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1046"},{"key":"e_1_3_2_1_41_1","volume-title":"Proceedings of NeurIPS 2019","author":"Yang Zhilin","year":"2019","unstructured":"Zhilin Yang , Zihang Dai , Yiming Yang , Jaime G. Carbonell , Ruslan Salakhutdinov , and Quoc V. Le . 2019. XLNet: Generalized Autoregressive Pretraining for Language Understanding . In Proceedings of NeurIPS 2019 ,. 5754--5764. https:\/\/proceedings.neurips.cc\/paper\/ 2019 \/hash\/dc6a7e655d7e5840e66733e9ee67cc69-Abstract.html Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. XLNet: Generalized Autoregressive Pretraining for Language Understanding. In Proceedings of NeurIPS 2019,. 5754--5764. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/dc6a7e655d7e5840e66733e9ee67cc69-Abstract.html"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413949"},{"key":"e_1_3_2_1_43_1","volume-title":"MEmoBERT: Pre-training Model with Prompt-based Learning for Multimodal Emotion Recognition. CoRR","author":"Zhao Jinming","year":"2021","unstructured":"Jinming Zhao , Ruichen Li , Qin Jin , Xinchao Wang , and Haizhou Li. 2021. MEmoBERT: Pre-training Model with Prompt-based Learning for Multimodal Emotion Recognition. CoRR , Vol. abs\/ 2111 .00865 ( 2021 ). showeprint[arXiv]2111.00865 https:\/\/arxiv.org\/abs\/2111.00865 Jinming Zhao, Ruichen Li, Qin Jin, Xinchao Wang, and Haizhou Li. 2021. MEmoBERT: Pre-training Model with Prompt-based Learning for Multimodal Emotion Recognition. CoRR, Vol. abs\/2111.00865 (2021). showeprint[arXiv]2111.00865 https:\/\/arxiv.org\/abs\/2111.00865"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.391"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2022\/628"},{"key":"e_1_3_2_1_46_1","volume-title":"Proceedings of COLING 2022","author":"Zhao Weixiang","year":"2022","unstructured":"Weixiang Zhao , Yanyan Zhao , and Bing Qin . 2022 c. MuCDN: Mutual Conversational Detachment Network for Emotion Recognition in Multi-Party Conversations . In Proceedings of COLING 2022 . International Committee on Computational Linguistics, 7020--7030. https:\/\/aclanthology.org\/ 2022.coling-1.612 Weixiang Zhao, Yanyan Zhao, and Bing Qin. 2022c. MuCDN: Mutual Conversational Detachment Network for Emotion Recognition in Multi-Party Conversations. In Proceedings of COLING 2022. International Committee on Computational Linguistics, 7020--7030. https:\/\/aclanthology.org\/2022.coling-1.612"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1103"}],"event":{"name":"CIKM '23: The 32nd ACM International Conference on Information and Knowledge Management","location":"Birmingham United Kingdom","acronym":"CIKM '23","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","SIGIR ACM Special Interest Group on Information Retrieval"]},"container-title":["Proceedings of the 32nd ACM International Conference on Information and Knowledge Management"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3583780.3615024","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3583780.3615024","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:36:55Z","timestamp":1750178215000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3583780.3615024"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,21]]},"references-count":47,"alternative-id":["10.1145\/3583780.3615024","10.1145\/3583780"],"URL":"https:\/\/doi.org\/10.1145\/3583780.3615024","relation":{},"subject":[],"published":{"date-parts":[[2023,10,21]]},"assertion":[{"value":"2023-10-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}