{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T20:30:31Z","timestamp":1777926631875,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":35,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,20]],"date-time":"2021-10-20T00:00:00Z","timestamp":1634688000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,24]]},"DOI":"10.1145\/3475957.3484454","type":"proceedings-article","created":{"date-parts":[[2021,10,15]],"date-time":"2021-10-15T23:34:16Z","timestamp":1634340856000},"page":"61-67","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":28,"title":["Multimodal Sentiment Analysis based on Recurrent Neural Network and Multimodal Attention"],"prefix":"10.1145","author":[{"given":"Cong","family":"Cai","sequence":"first","affiliation":[{"name":"University of Chinese Academy of Sciences &amp; Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu","family":"He","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences &amp; Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Licai","family":"Sun","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences &amp; Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheng","family":"Lian","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bin","family":"Liu","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianhua","family":"Tao","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences &amp; University of Chinese Academy of Sciences &amp; Chinese Academy of Sciences Center for Excellence in Brain Science and Intelligence Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mingyu","family":"Xu","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences &amp; Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kexin","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences &amp; Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,20]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240578"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"crossref","unstructured":"Shahin Amiriparian Maurice Gerczuk Sandra Ottl Nicholas Cummins Michael Freitag Sergey Pugachevskiy Alice Baird and Bj\u00f6rn Schuller. 2017. Snore sound classification using image-based deep spectrum features. (2017).  Shahin Amiriparian Maurice Gerczuk Sandra Ottl Nicholas Cummins Michael Freitag Sergey Pugachevskiy Alice Baird and Bj\u00f6rn Schuller. 2017. Snore sound classification using image-based deep spectrum features. (2017).","DOI":"10.21437\/Interspeech.2017-434"},{"key":"e_1_3_2_2_3_1","volume-title":"Jamie Ryan Kiros, and Geoffrey E. Hinton","author":"Ba Jimmy Lei","year":"2016","unstructured":"Jimmy Lei Ba , Jamie Ryan Kiros, and Geoffrey E. Hinton . 2016 . Layer Normalization. arXiv: Machine Learning ( 2016). Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer Normalization. arXiv: Machine Learning (2016)."},{"key":"e_1_3_2_2_4_1","volume-title":"An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv preprint arXiv:1803.01271","author":"Bai Shaojie","year":"2018","unstructured":"Shaojie Bai , J. Zico Kolter , and Vladlen Koltun . 2018. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv preprint arXiv:1803.01271 ( 2018 ). Shaojie Bai, J. Zico Kolter, and Vladlen Koltun. 2018. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv preprint arXiv:1803.01271 (2018)."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/MMSP.2019.8901785"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2798607"},{"key":"e_1_3_2_2_7_1","volume-title":"2016 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1--10","author":"Tadas Baltruvs","year":"2016","unstructured":"Tadas Baltruvs aitis, Peter Robinson , and Louis-Philippe Morency . 2016 . Openface: an open source facial behavior analysis toolkit . In 2016 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1--10 . Tadas Baltruvs aitis, Peter Robinson, and Louis-Philippe Morency. 2016. Openface: an open source facial behavior analysis toolkit. In 2016 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1--10."},{"key":"e_1_3_2_2_8_1","first-page":"2511","article-title":"TDCA-Net: Time-Domain Channel Attention Network for Depression Detection","volume":"2021","author":"Cai Cong","year":"2021","unstructured":"Cong Cai , Mingyue Niu , Bin Liu , Jianhua Tao , and Xuefei Liu . 2021 . TDCA-Net: Time-Domain Channel Attention Network for Depression Detection . Proc. Interspeech 2021 (2021), 2511 -- 2515 . Cong Cai, Mingyue Niu, Bin Liu, Jianhua Tao, and Xuefei Liu. 2021. TDCA-Net: Time-Domain Channel Attention Network for Depression Detection. Proc. Interspeech 2021 (2021), 2511--2515.","journal-title":"Proc. Interspeech"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/FG.2018.00020"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661806.2661811"},{"key":"e_1_3_2_2_11_1","volume-title":"Transformer Encoder with Multi-modal Multi-head Attention for Continuous Affect Recognition","author":"Chen Haifeng","year":"2020","unstructured":"Haifeng Chen , Dongmei Jiang , and Hichem Sahli . 2020. Transformer Encoder with Multi-modal Multi-head Attention for Continuous Affect Recognition . IEEE Transactions on Multimedia ( 2020 ). Haifeng Chen, Dongmei Jiang, and Hichem Sahli. 2020. Transformer Encoder with Multi-modal Multi-head Attention for Continuous Affect Recognition. IEEE Transactions on Multimedia (2020)."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967286"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133944.3133949"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/79.911197"},{"key":"e_1_3_2_2_15_1","volume-title":"et almbox","author":"Eyben Florian","year":"2015","unstructured":"Florian Eyben , Klaus R Scherer , Bj\u00f6rn W Schuller , Johan Sundberg , Elisabeth Andr\u00e9 , Carlos Busso , Laurence Y Devillers , Julien Epps , Petri Laukka , Shrikanth S Narayanan , et almbox . 2015 . The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing. IEEE transactions on affective computing , Vol. 7 , 2 (2015), 190--202. Florian Eyben, Klaus R Scherer, Bj\u00f6rn W Schuller, Johan Sundberg, Elisabeth Andr\u00e9, Carlos Busso, Laurence Y Devillers, Julien Epps, Petri Laukka, Shrikanth S Narayanan, et almbox. 2015. The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing. IEEE transactions on affective computing , Vol. 7, 2 (2015), 190--202."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874246"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2808196.2811641"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133944.3133946"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053762"},{"key":"e_1_3_2_2_20_1","volume-title":"Adam: A Method for Stochastic Optimization. In ICLR 2015 : International Conference on Learning Representations 2015 .","author":"Diederik","unstructured":"Diederik P. Kingma and Jimmy Lei Ba. 2015 . Adam: A Method for Stochastic Optimization. In ICLR 2015 : International Conference on Learning Representations 2015 . Diederik P. Kingma and Jimmy Lei Ba. 2015. Adam: A Method for Stochastic Optimization. In ICLR 2015 : International Conference on Learning Representations 2015 ."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.2307\/2532051"},{"key":"e_1_3_2_2_22_1","volume-title":"NeuroKit2: A Python toolbox for neurophysiological signal processing. Behavior Research Methods 51","author":"MakosKi D.","year":"2021","unstructured":"D. MakosKi , T. Pham , Z. J. Lau , J. C. Brammer , and Sha Chen . 2021. NeuroKit2: A Python toolbox for neurophysiological signal processing. Behavior Research Methods 51 ( 2021 ). D. MakosKi, T. Pham, Z. J. Lau, J. C. Brammer, and Sha Chen. 2021. NeuroKit2: A Python toolbox for neurophysiological signal processing. Behavior Research Methods 51 (2021)."},{"key":"e_1_3_2_2_23_1","unstructured":"Nikhil Mishra Mostafa Rohaninejad Xi Chen and Pieter Abbeel. 2017. Meta-Learning with Temporal Convolutions. (2017).  Nikhil Mishra Mostafa Rohaninejad Xi Chen and Pieter Abbeel. 2017. Meta-Learning with Temporal Convolutions. (2017)."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2512530.2512534"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2009-103"},{"key":"e_1_3_2_2_26_1","volume-title":"The INTERSPEECH 2010 paralinguistic challenge. In Eleventh Annual Conference of the International Speech Communication Association .","author":"Schuller Bj\u00f6rn","year":"2010","unstructured":"Bj\u00f6rn Schuller , Stefan Steidl , Anton Batliner , Felix Burkhardt , Laurence Devillers , Christian M\u00fcller , and Shrikanth S Narayanan . 2010 . The INTERSPEECH 2010 paralinguistic challenge. In Eleventh Annual Conference of the International Speech Communication Association . Bj\u00f6rn Schuller, Stefan Steidl, Anton Batliner, Felix Burkhardt, Laurence Devillers, Christian M\u00fcller, and Shrikanth S Narayanan. 2010. The INTERSPEECH 2010 paralinguistic challenge. In Eleventh Annual Conference of the International Speech Communication Association ."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3475957.3484450"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3423327.3423673"},{"key":"e_1_3_2_2_29_1","volume-title":"2020 b. Cross-lingual zero-and few-shot hate speech detection utilising frozen transformer language models and AXEL. arXiv preprint arXiv:2004.13850","author":"Stappen Lukas","year":"2020","unstructured":"Lukas Stappen , Fabian Brunn , and Bj\u00f6rn Schuller . 2020 b. Cross-lingual zero-and few-shot hate speech detection utilising frozen transformer language models and AXEL. arXiv preprint arXiv:2004.13850 ( 2020 ). Lukas Stappen, Fabian Brunn, and Bj\u00f6rn Schuller. 2020 b. Cross-lingual zero-and few-shot hate speech detection utilising frozen transformer language models and AXEL. arXiv preprint arXiv:2004.13850 (2020)."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3423327.3423672"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-020-00280-0"},{"key":"e_1_3_2_2_32_1","volume-title":"Proceedings of the conference. Association for Computational Linguistics. Meeting","volume":"2019","author":"Hubert Tsai Yao-Hung","year":"2019","unstructured":"Yao-Hung Hubert Tsai , Shaojie Bai , Paul Pu Liang , J Zico Kolter , Louis-Philippe Morency , and Ruslan Salakhutdinov . 2019 . Multimodal transformer for unaligned multimodal language sequences . In Proceedings of the conference. Association for Computational Linguistics. Meeting , Vol. 2019 . NIH Public Access, 6558. Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the conference. Association for Computational Linguistics. Meeting , Vol. 2019. NIH Public Access, 6558."},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2988257.2988258"},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3266302.3266313"}],"event":{"name":"MM '21: ACM Multimedia Conference","location":"Virtual Event China","acronym":"MM '21","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 2nd on Multimodal Sentiment Analysis Challenge"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3475957.3484454","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3475957.3484454","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:28:33Z","timestamp":1750195713000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3475957.3484454"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,20]]},"references-count":35,"alternative-id":["10.1145\/3475957.3484454","10.1145\/3475957"],"URL":"https:\/\/doi.org\/10.1145\/3475957.3484454","relation":{},"subject":[],"published":{"date-parts":[[2021,10,20]]},"assertion":[{"value":"2021-10-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}