{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,25]],"date-time":"2026-08-25T21:08:39Z","timestamp":1787692119997,"version":"build-2784847793"},"publisher-location":"New York, NY, USA","reference-count":42,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3548025","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:42:46Z","timestamp":1665416566000},"page":"3722-3729","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":149,"title":["CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression Estimation"],"prefix":"10.1145","author":[{"given":"Hao","family":"Sun","sequence":"first","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongyi","family":"Wang","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiaqing","family":"Liu","sequence":"additional","affiliation":[{"name":"Ritsumeikan University, Shiga, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yen-Wei","family":"Chen","sequence":"additional","affiliation":[{"name":"Ritsumeikan University, Shiga, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lanfen","family":"Lin","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.93"},{"key":"e_1_3_2_2_3_1","volume-title":"A Transformer-based joint-encoding for Emotion Recognition and Sentiment Analysis. arXiv preprint arXiv:2006.15955","author":"Delbrouck Jean-Benoit","year":"2020","unstructured":"Jean-Benoit Delbrouck , No\u00e9 Tits , Mathilde Brousmiche , and St\u00e9phane Dupont . 2020b. A Transformer-based joint-encoding for Emotion Recognition and Sentiment Analysis. arXiv preprint arXiv:2006.15955 ( 2020 ). Jean-Benoit Delbrouck, No\u00e9 Tits, Mathilde Brousmiche, and St\u00e9phane Dupont. 2020b. A Transformer-based joint-encoding for Emotion Recognition and Sentiment Analysis. arXiv preprint arXiv:2006.15955 (2020)."},{"key":"e_1_3_2_2_4_1","volume-title":"Modulated fusion using transformer for linguistic-acoustic emotion recognition. arXiv preprint arXiv:2010.02057","author":"Delbrouck Jean-Benoit","year":"2020","unstructured":"Jean-Benoit Delbrouck , No\u00e9 Tits , and St\u00e9phane Dupont . 2020a. Modulated fusion using transformer for linguistic-acoustic emotion recognition. arXiv preprint arXiv:2010.02057 ( 2020 ). Jean-Benoit Delbrouck, No\u00e9 Tits, and St\u00e9phane Dupont. 2020a. Modulated fusion using transformer for linguistic-acoustic emotion recognition. arXiv preprint arXiv:2010.02057 (2020)."},{"key":"e_1_3_2_2_5_1","volume-title":"Dense Fusion Network with Multimodal Residual for Sentiment Classification. In 2021 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1--6.","author":"Deng Huan","year":"2021","unstructured":"Huan Deng , Peipei Kang , Zhenguo Yang , Tianyong Hao , Qing Li , and Wenyin Liu . 2021 . Dense Fusion Network with Multimodal Residual for Sentiment Classification. In 2021 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1--6. Huan Deng, Peipei Kang, Zhenguo Yang, Tianyong Hao, Qing Li, and Wenyin Liu. 2021. Dense Fusion Network with Multimodal Residual for Sentiment Classification. In 2021 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1--6."},{"key":"e_1_3_2_2_6_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_2_7_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly etal 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020).  Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)."},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3347320.3357695"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1207"},{"key":"e_1_3_2_2_10_1","volume-title":"Hire-mlp: Vision mlp via hierarchical rearrangement. arXiv preprint arXiv:2108.13341","author":"Guo Jianyuan","year":"2021","unstructured":"Jianyuan Guo , Yehui Tang , Kai Han , Xinghao Chen , Han Wu , Chao Xu , Chang Xu , and Yunhe Wang . 2021 . Hire-mlp: Vision mlp via hierarchical rearrangement. arXiv preprint arXiv:2108.13341 (2021). Jianyuan Guo, Yehui Tang, Kai Han, Xinghao Chen, Han Wu, Chao Xu, Chang Xu, and Yunhe Wang. 2021. Hire-mlp: Vision mlp via hierarchical rearrangement. arXiv preprint arXiv:2108.13341 (2021)."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3462244.3479919"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413678"},{"key":"e_1_3_2_2_13_1","volume-title":"Long short-term memory. Neural computation","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber . 1997. Long short-term memory. Neural computation , Vol. 9 , 8 ( 1997 ), 1735--1780. Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural computation, Vol. 9, 8 (1997), 1735--1780."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12193-013-0123-2"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3347320.3357691"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"crossref","unstructured":"Kurt Kroenke and Robert L Spitzer. 2002. The PHQ-9: a new depression diagnostic and severity measure. 509--515 pages.  Kurt Kroenke and Robert L Spitzer. 2002. The PHQ-9: a new depression diagnostic and severity measure. 509--515 pages.","DOI":"10.3928\/0048-5713-20020901-06"},{"key":"e_1_3_2_2_18_1","volume-title":"A concordance correlation coefficient to evaluate reproducibility. Biometrics","author":"Lawrence I","year":"1989","unstructured":"I Lawrence and Kuei Lin . 1989. A concordance correlation coefficient to evaluate reproducibility. Biometrics ( 1989 ), 255--268. I Lawrence and Kuei Lin. 1989. A concordance correlation coefficient to evaluate reproducibility. Biometrics (1989), 255--268."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-63031-7_26"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3068598"},{"key":"e_1_3_2_2_22_1","volume-title":"Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , 2019 . Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems , Vol. 32 (2019). Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems , Vol. 32 (2019)."},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCI.2020.2998234"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3347320.3357697"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"crossref","unstructured":"Fabien Ringeval Bj\u00f6rn Schuller Michel Valstar Nicholas Cummins Roddy Cowie Leili Tavabi Maximilian Schmitt Sina Alisamir Shahin Amiriparian Eva-Maria Messner etal 2019. AVEC 2019 workshop and challenge: state-of-mind detecting depression with AI and cross-cultural affect recognition. In Proceedings of the 9th International on Audio\/visual Emotion Challenge and Workshop. 3--12.  Fabien Ringeval Bj\u00f6rn Schuller Michel Valstar Nicholas Cummins Roddy Cowie Leili Tavabi Maximilian Schmitt Sina Alisamir Shahin Amiriparian Eva-Maria Messner et al. 2019. AVEC 2019 workshop and challenge: state-of-mind detecting depression with AI and cross-cultural affect recognition. In Proceedings of the 9th International on Audio\/visual Emotion Challenge and Workshop. 3--12.","DOI":"10.1145\/3347320.3357688"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3347320.3357694"},{"key":"e_1_3_2_2_27_1","volume-title":"Low Rank Fusion based Transformers for Multimodal Sequences. arXiv preprint arXiv:2007.02038","author":"Sahay Saurav","year":"2020","unstructured":"Saurav Sahay , Eda Okur , Shachi H Kumar , and Lama Nachman . 2020. Low Rank Fusion based Transformers for Multimodal Sequences. arXiv preprint arXiv:2007.02038 ( 2020 ). Saurav Sahay, Eda Okur, Shachi H Kumar, and Lama Nachman. 2020. Low Rank Fusion based Transformers for Multimodal Sequences. arXiv preprint arXiv:2007.02038 (2020)."},{"key":"e_1_3_2_2_28_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.3390\/s21144764"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6431"},{"key":"e_1_3_2_2_31_1","volume-title":"Advances in Neural Information Processing Systems","volume":"34","author":"Tolstikhin Ilya O","year":"2021","unstructured":"Ilya O Tolstikhin , Neil Houlsby , Alexander Kolesnikov , Lucas Beyer , Xiaohua Zhai , Thomas Unterthiner , Jessica Yung , Andreas Steiner , Daniel Keysers , Jakob Uszkoreit , 2021 . Mlp-mixer: An all-mlp architecture for vision . Advances in Neural Information Processing Systems , Vol. 34 (2021). Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al. 2021. Mlp-mixer: An all-mlp architecture for vision. Advances in Neural Information Processing Systems , Vol. 34 (2021)."},{"key":"e_1_3_2_2_32_1","volume-title":"Resmlp: Feedforward networks for image classification with data-efficient training. arXiv preprint arXiv:2105.03404","author":"Touvron Hugo","year":"2021","unstructured":"Hugo Touvron , Piotr Bojanowski , Mathilde Caron , Matthieu Cord , Alaaeldin El-Nouby , Edouard Grave , Gautier Izacard , Armand Joulin , Gabriel Synnaeve , Jakob Verbeek , 2021 . Resmlp: Feedforward networks for image classification with data-efficient training. arXiv preprint arXiv:2105.03404 (2021). Hugo Touvron, Piotr Bojanowski, Mathilde Caron, Matthieu Cord, Alaaeldin El-Nouby, Edouard Grave, Gautier Izacard, Armand Joulin, Gabriel Synnaeve, Jakob Verbeek, et al. 2021. Resmlp: Feedforward networks for image classification with data-efficient training. arXiv preprint arXiv:2105.03404 (2021)."},{"key":"e_1_3_2_2_33_1","volume-title":"Proceedings of the conference. Association for Computational Linguistics. Meeting","volume":"2019","author":"Hubert Tsai Yao-Hung","year":"2019","unstructured":"Yao-Hung Hubert Tsai , Shaojie Bai , Paul Pu Liang , J Zico Kolter , Louis-Philippe Morency , and Ruslan Salakhutdinov . 2019 . Multimodal transformer for unaligned multimodal language sequences . In Proceedings of the conference. Association for Computational Linguistics. Meeting , Vol. 2019 . NIH Public Access, 6558. Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the conference. Association for Computational Linguistics. Meeting, Vol. 2019. NIH Public Access, 6558."},{"key":"e_1_3_2_2_34_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380000"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3347320.3357696"},{"key":"e_1_3_2_2_37_1","volume-title":"Tensor fusion network for multimodal sentiment analysis. arXiv preprint arXiv:1707.07250","author":"Zadeh Amir","year":"2017","unstructured":"Amir Zadeh , Minghai Chen , Soujanya Poria , Erik Cambria , and Louis-Philippe Morency . 2017. Tensor fusion network for multimodal sentiment analysis. arXiv preprint arXiv:1707.07250 ( 2017 ). Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. arXiv preprint arXiv:1707.07250 (2017)."},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12021"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12024"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2016.94"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1208"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Ziping Zhao Qifei Li Nicholas Cummins Bin Liu Haishuai Wang Jianhua Tao and Bj\u00f6rn W Schuller. 2020. Hybrid Network Feature Extraction for Depression Assessment from Speech.. In INTERSPEECH. 4956--4960.io  Ziping Zhao Qifei Li Nicholas Cummins Bin Liu Haishuai Wang Jianhua Tao and Bj\u00f6rn W Schuller. 2020. Hybrid Network Feature Extraction for Depression Assessment from Speech.. In INTERSPEECH. 4956--4960.io","DOI":"10.21437\/Interspeech.2020-2396"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548025","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3548025","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:29Z","timestamp":1750186949000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548025"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":42,"alternative-id":["10.1145\/3503161.3548025","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3548025","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}