{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,7]],"date-time":"2026-06-07T10:18:36Z","timestamp":1780827516276,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":32,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,10,15]],"date-time":"2019-10-15T00:00:00Z","timestamp":1571097600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,10,15]]},"DOI":"10.1145\/3347320.3357690","type":"proceedings-article","created":{"date-parts":[[2019,10,24]],"date-time":"2019-10-24T19:04:48Z","timestamp":1571943888000},"page":"19-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":27,"title":["Efficient Spatial Temporal Convolutional Features for Audiovisual Continuous Affect Recognition"],"prefix":"10.1145","author":[{"given":"Haifeng","family":"Chen","sequence":"first","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yifan","family":"Deng","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shiwen","family":"Cheng","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yixuan","family":"Wang","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dongmei","family":"Jiang","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University &amp; PengCheng Laboratory, Xi'an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hichem","family":"Sahli","sequence":"additional","affiliation":[{"name":"Vrije University Brussel &amp; Interuniversity Microelectronics Centre, Brussels, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,10,15]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"ISPRS Journal of Photogrammetry and Remote Sensing","volume":"150","author":"Ai Tinghua","year":"2019","unstructured":"Tinghua Ai and Xiongfeng Yan . 2019 . a graph convolution neuroal network . ISPRS Journal of Photogrammetry and Remote Sensing , Vol. 150 (03 2019). https:\/\/doi.org\/10.1016\/j.isprsjprs.2019.02.010 10.1016\/j.isprsjprs.2019.02.010 Tinghua Ai and Xiongfeng Yan. 2019. a graph convolution neuroal network. ISPRS Journal of Photogrammetry and Remote Sensing, Vol. 150 (03 2019). https:\/\/doi.org\/10.1016\/j.isprsjprs.2019.02.010"},{"key":"e_1_3_2_1_2_1","volume-title":"Valstar","author":"Almaev Timur R.","year":"2013","unstructured":"Timur R. Almaev and Michel F . Valstar . 2013 . Local Gabor Binary Patterns from Three Orthogonal Planes for Automatic Facial Expression Recognition. In Affective Computing and Intelligent Interaction . Timur R. Almaev and Michel F. Valstar. 2013. Local Gabor Binary Patterns from Three Orthogonal Planes for Automatic Facial Expression Recognition. In Affective Computing and Intelligent Interaction."},{"key":"e_1_3_2_1_4_1","volume-title":"Soundnet: Learning sound representations from unlabeled video. In Advances in Neural Information Processing Systems.","author":"Aytar Yusuf","year":"2016","unstructured":"Yusuf Aytar , Carl Vondrick , and Antonio Torralba . 2016 . Soundnet: Learning sound representations from unlabeled video. In Advances in Neural Information Processing Systems. Yusuf Aytar, Carl Vondrick, and Antonio Torralba. 2016. Soundnet: Learning sound representations from unlabeled video. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_1_5_1","volume-title":"An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. CoRR","author":"Bai Shaojie","year":"2018","unstructured":"Shaojie Bai , J. Zico Kolter , and Vladlen Koltun . 2018. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. CoRR , Vol. abs\/ 1803 .01271 ( 2018 ). arxiv: 1803.01271 http:\/\/arxiv.org\/abs\/1803.01271 Shaojie Bai, J. Zico Kolter, and Vladlen Koltun. 2018. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. CoRR, Vol. abs\/1803.01271 (2018). arxiv: 1803.01271 http:\/\/arxiv.org\/abs\/1803.01271"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2015.09.009"},{"key":"e_1_3_2_1_7_1","volume-title":"Training Deep Networks for Facial Expression Recognition with Crowd-Sourced Label Distribution. CoRR","author":"Barsoum Emad","year":"2016","unstructured":"Emad Barsoum , Cha Zhang , Cristian Canton-Ferrer , and Zhengyou Zhang . 2016. Training Deep Networks for Facial Expression Recognition with Crowd-Sourced Label Distribution. CoRR , Vol. abs\/ 1608 .01041 ( 2016 ). arxiv: 1608.01041 http:\/\/arxiv.org\/abs\/1608.01041 Emad Barsoum, Cha Zhang, Cristian Canton-Ferrer, and Zhengyou Zhang. 2016. Training Deep Networks for Facial Expression Recognition with Crowd-Sourced Label Distribution. CoRR, Vol. abs\/1608.01041 (2016). arxiv: 1608.01041 http:\/\/arxiv.org\/abs\/1608.01041"},{"key":"e_1_3_2_1_8_1","volume-title":"Proceedings of the 6th International Workshop on Audio\/Visual Emotion Challenge (AVEC '16)","author":"Brady Kevin","unstructured":"Kevin Brady , Youngjune Gwon , Pooya Khorrami , Elizabeth Godoy , William Campbell , Charlie Dagli , and Thomas S. Huang . 2016. Multi-Modal Audio, Video and Physiological Sensor Learning for Continuous Emotion Prediction . In Proceedings of the 6th International Workshop on Audio\/Visual Emotion Challenge (AVEC '16) . ACM, New York, NY, USA, 97--104. https:\/\/doi.org\/10.1145\/2988257.2988264 10.1145\/2988257.2988264 Kevin Brady, Youngjune Gwon, Pooya Khorrami, Elizabeth Godoy, William Campbell, Charlie Dagli, and Thomas S. Huang. 2016. Multi-Modal Audio, Video and Physiological Sensor Learning for Continuous Emotion Prediction. In Proceedings of the 6th International Workshop on Audio\/Visual Emotion Challenge (AVEC '16). ACM, New York, NY, USA, 97--104. https:\/\/doi.org\/10.1145\/2988257.2988264"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133944.3133949"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2015.2457417"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Fabien Ringeval and Bj\u00f6rn Schuller and Michel Valstar and Nicholas Cummins and Roddy Cowie and Leili Tavabi and Maximilian Schmitt and Sina Alisamir and Shahin Amiriparian and Eva-Maria Messner and Siyang Song and Shuo Lui and Ziping Zhao and Adria Mallol-Ragolta and Zhao Ren and Maja Pantic. 2019. AVEC 2019 Workshop and Challenge: State-of-Mind Depression with AI and Cross-Cultural Affect Recognition. In Proceedings of the 9th International Workshop on Audio\/Visual Emotion Challenge AVEC'19 co-located with the 27th ACM International Conference on Multimedia MM 2019 Fabien Ringeval Bj\u00f6rn Schuller Michel Valstar Nicholas Cummins Roddy Cowie and Maja Pantic (Eds.). ACM Nice France.  Fabien Ringeval and Bj\u00f6rn Schuller and Michel Valstar and Nicholas Cummins and Roddy Cowie and Leili Tavabi and Maximilian Schmitt and Sina Alisamir and Shahin Amiriparian and Eva-Maria Messner and Siyang Song and Shuo Lui and Ziping Zhao and Adria Mallol-Ragolta and Zhao Ren and Maja Pantic. 2019. AVEC 2019 Workshop and Challenge: State-of-Mind Depression with AI and Cross-Cultural Affect Recognition. In Proceedings of the 9th International Workshop on Audio\/Visual Emotion Challenge AVEC'19 co-located with the 27th ACM International Conference on Multimedia MM 2019 Fabien Ringeval Bj\u00f6rn Schuller Michel Valstar Nicholas Cummins Roddy Cowie and Maja Pantic (Eds.). ACM Nice France.","DOI":"10.1145\/3347320.3357688"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.3390\/e21050479"},{"key":"e_1_3_2_1_13_1","volume-title":"The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops.","author":"Mohammad Mahoor Behzad H.","year":"2017","unstructured":"Behzad H. Hasani; Mohammad Mahoor . 2017 . Facial Expression Recognition Using Enhanced Deep 3D Convolutional Neural Networks . In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops. Behzad H. Hasani; Mohammad Mahoor. 2017. Facial Expression Recognition Using Enhanced Deep 3D Convolutional Neural Networks. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops."},{"key":"e_1_3_2_1_14_1","volume-title":"Deep Residual Learning for Image Recognition. CoRR","author":"He Kaiming","year":"2015","unstructured":"Kaiming He , Xiangyu Zhang , Shaoqing Ren , and Jian Sun . 2015. Deep Residual Learning for Image Recognition. CoRR , Vol. abs\/ 1512 .03385 ( 2015 ). arxiv: 1512.03385 http:\/\/arxiv.org\/abs\/1512.03385 Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. CoRR, Vol. abs\/1512.03385 (2015). arxiv: 1512.03385 http:\/\/arxiv.org\/abs\/1512.03385"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"crossref","unstructured":"Shawn Hershey Sourish Chaudhuri Daniel P. W. Ellis Jort F. Gemmeke and Kevin Wilson. 2016. CNN architectures for large-scale audio classification. (2016).  Shawn Hershey Sourish Chaudhuri Daniel P. W. Ellis Jort F. Gemmeke and Kevin Wilson. 2016. CNN architectures for large-scale audio classification. (2016).","DOI":"10.1109\/ICASSP.2017.7952132"},{"key":"e_1_3_2_1_16_1","volume-title":"Kam Star, Elnar Hajiyev, and Maja Pantic.","author":"Kossaifi Jean","year":"2019","unstructured":"Jean Kossaifi , Robert Walecki , Yannis Panagakis , Jie Shen , Maximilian Schmitt , Fabien Ringeval , Jing Han , Vedhas Pandit , Bj\u00f6 rn W. Schuller , Kam Star, Elnar Hajiyev, and Maja Pantic. 2019 . SEWA DB : A Rich Database for Audio-Visual Emotion and Sentiment Research in the Wild. CoRR , Vol. abs\/ 1901 .02839 (2019). arxiv: 1901.02839 http:\/\/arxiv.org\/abs\/1901.02839 Jean Kossaifi, Robert Walecki, Yannis Panagakis, Jie Shen, Maximilian Schmitt, Fabien Ringeval, Jing Han, Vedhas Pandit, Bj\u00f6 rn W. Schuller, Kam Star, Elnar Hajiyev, and Maja Pantic. 2019. SEWA DB: A Rich Database for Audio-Visual Emotion and Sentiment Research in the Wild. CoRR, Vol. abs\/1901.02839 (2019). arxiv: 1901.02839 http:\/\/arxiv.org\/abs\/1901.02839"},{"key":"e_1_3_2_1_17_1","volume-title":"Multimodal Affective Dimension Prediction Using Deep Bidirectional Long Short-Term Memory Recurrent Neural Networks. In International Workshop on Audio\/visual Emotion Challenge. ACM.","author":"Lang He","year":"2015","unstructured":"He Lang and Jiang Dongmei . 2015 . Multimodal Affective Dimension Prediction Using Deep Bidirectional Long Short-Term Memory Recurrent Neural Networks. In International Workshop on Audio\/visual Emotion Challenge. ACM. He Lang and Jiang Dongmei. 2015. Multimodal Affective Dimension Prediction Using Deep Bidirectional Long Short-Term Memory Recurrent Neural Networks. In International Workshop on Audio\/visual Emotion Challenge. ACM."},{"key":"e_1_3_2_1_18_1","volume-title":"Spatio-Temporal Graph Convolution for Skeleton Based Action Recognition. CoRR","author":"Li Chaolong","year":"2018","unstructured":"Chaolong Li , Zhen Cui , Wenming Zheng , Chunyan Xu , and Jian Yang . 2018. Spatio-Temporal Graph Convolution for Skeleton Based Action Recognition. CoRR , Vol. abs\/ 1802 .09834 ( 2018 ). arxiv: 1802.09834 http:\/\/arxiv.org\/abs\/1802.09834 Chaolong Li, Zhen Cui, Wenming Zheng, Chunyan Xu, and Jian Yang. 2018. Spatio-Temporal Graph Convolution for Skeleton Based Action Recognition. CoRR, Vol. abs\/1802.09834 (2018). arxiv: 1802.09834 http:\/\/arxiv.org\/abs\/1802.09834"},{"key":"e_1_3_2_1_19_1","volume-title":"Efficient Estimation of Word Representations in Vector Space. Computer Science","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov , Kai Chen , Greg Corrado , and Jeffrey Dean . 2013. Efficient Estimation of Word Representations in Vector Space. Computer Science ( 2013 ). Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. Computer Science (2013)."},{"key":"e_1_3_2_1_20_1","volume-title":"Huang","author":"Pantic Maja","year":"2007","unstructured":"Maja Pantic , Alex Pentland , Anton Nijholt , and Thomas S . Huang . 2007 . Human Computing and Machine Understanding of Human Behavior: A Survey . 47--71. https:\/\/doi.org\/10.1007\/978--3--540--72348--6_3 10.1007\/978--3--540--72348--6_3 Maja Pantic, Alex Pentland, Anton Nijholt, and Thomas S. Huang. 2007. Human Computing and Machine Understanding of Human Behavior: A Survey. 47--71. https:\/\/doi.org\/10.1007\/978--3--540--72348--6_3"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2003.817122"},{"key":"e_1_3_2_1_22_1","volume-title":"Glove: Global Vectors for Word Representation. In Conference on Empirical Methods in Natural Language Processing.","author":"Pennington Jeffrey","year":"2014","unstructured":"Jeffrey Pennington , Richard Socher , and Christopher Manning . 2014 . Glove: Global Vectors for Word Representation. In Conference on Empirical Methods in Natural Language Processing. Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. Glove: Global Vectors for Word Representation. In Conference on Empirical Methods in Natural Language Processing."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2806408"},{"key":"e_1_3_2_1_24_1","volume-title":"AVEC","author":"Schuller Bjorn","year":"2011","unstructured":"Bjorn Schuller , Michel Valstar , Florian Eyben , Gary Mckeown , Roddy Cowie , and Maja Pantic . 2011 . AVEC 2011, The First International Audio\/Visual Emotion Challenge. 415--424. https:\/\/doi.org\/10.1007\/978--3--642--24571--8_53 10.1007\/978--3--642--24571--8_53 Bjorn Schuller, Michel Valstar, Florian Eyben, Gary Mckeown, Roddy Cowie, and Maja Pantic. 2011. AVEC 2011, The First International Audio\/Visual Emotion Challenge. 415--424. https:\/\/doi.org\/10.1007\/978--3--642--24571--8_53"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"crossref","unstructured":"Du Tran Heng Wang Lorenzo Torresani Jamie Ray Yann Lecun and Manohar Paluri. 2017. A Closer Look at Spatiotemporal Convolutions for Action Recognition. (2017).  Du Tran Heng Wang Lorenzo Torresani Jamie Ray Yann Lecun and Manohar Paluri. 2017. A Closer Look at Spatiotemporal Convolutions for Action Recognition. (2017).","DOI":"10.1109\/CVPR.2018.00675"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2017.2764438"},{"key":"e_1_3_2_1_28_1","volume-title":"AVEC 2016 - Depression, Mood, and Emotion Recognition Workshop and Challenge.","author":"Valstar Michel","year":"2016","unstructured":"Michel Valstar , Jonathan Gratch , Bjorn Schuller , Fabien Ringeval , Denis Lalanne , Mercedes Torres Torres , Stefan Scherer , Guiota Stratou , Roddy Cowie , and Maja Pantic . 2016 . AVEC 2016 - Depression, Mood, and Emotion Recognition Workshop and Challenge. (2016). Michel Valstar, Jonathan Gratch, Bjorn Schuller, Fabien Ringeval, Denis Lalanne, Mercedes Torres Torres, Stefan Scherer, Guiota Stratou, Roddy Cowie, and Maja Pantic. 2016. AVEC 2016 - Depression, Mood, and Emotion Recognition Workshop and Challenge. (2016)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Michel F. Valstar Bihan Jiang Marc Mehu Maja Pantic and Klaus Scherer. 2011. The first facial expression recognition and analysis challenge. (2011).  Michel F. Valstar Bihan Jiang Marc Mehu Maja Pantic and Klaus Scherer. 2011. The first facial expression recognition and analysis challenge. (2011).","DOI":"10.1109\/FG.2011.5771374"},{"key":"#cr-split#-e_1_3_2_1_30_1.1","doi-asserted-by":"crossref","unstructured":"L. Yang D. Jiang and H. Sahli. 2018. Integrating Deep and Shallow Models for Multi-Modal Depression Analysis Hybrid Architectures. IEEE Transactions on Affective Computing (2018) 1--1. https:\/\/doi.org\/10.1109\/TAFFC.2018.2870398 10.1109\/TAFFC.2018.2870398","DOI":"10.1109\/TAFFC.2018.2870398"},{"key":"#cr-split#-e_1_3_2_1_30_1.2","doi-asserted-by":"crossref","unstructured":"L. Yang D. Jiang and H. Sahli. 2018. Integrating Deep and Shallow Models for Multi-Modal Depression Analysis Hybrid Architectures. IEEE Transactions on Affective Computing (2018) 1--1. https:\/\/doi.org\/10.1109\/TAFFC.2018.2870398","DOI":"10.1109\/TAFFC.2018.2870398"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2017.2719043"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3266302.3266313"}],"event":{"name":"MM '19: The 27th ACM International Conference on Multimedia","location":"Nice France","acronym":"MM '19","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 9th International on Audio\/Visual Emotion Challenge and Workshop"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3347320.3357690","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3347320.3357690","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T19:05:46Z","timestamp":1750273546000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3347320.3357690"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,15]]},"references-count":32,"alternative-id":["10.1145\/3347320.3357690","10.1145\/3347320"],"URL":"https:\/\/doi.org\/10.1145\/3347320.3357690","relation":{},"subject":[],"published":{"date-parts":[[2019,10,15]]},"assertion":[{"value":"2019-10-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}