{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,7]],"date-time":"2026-08-07T10:52:38Z","timestamp":1786099958072,"version":"3.56.0"},"publisher-location":"New York, NY, USA","reference-count":29,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,10,15]],"date-time":"2019-10-15T00:00:00Z","timestamp":1571097600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012659","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U180120050,61702192,U1636218"],"award-info":[{"award-number":["U180120050,61702192,U1636218"]}],"id":[{"id":"10.13039\/501100012659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,10,15]]},"DOI":"10.1145\/3347320.3357695","type":"proceedings-article","created":{"date-parts":[[2019,10,24]],"date-time":"2019-10-24T19:04:48Z","timestamp":1571943888000},"page":"73-80","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":44,"title":["Multi-modality Depression Detection via Multi-scale Temporal Dilated CNNs"],"prefix":"10.1145","author":[{"given":"Weiquan","family":"Fan","sequence":"first","affiliation":[{"name":"South China University of Technology, GuangZhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhiwei","family":"He","sequence":"additional","affiliation":[{"name":"South China University of Technology, GuangZhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaofen","family":"Xing","sequence":"additional","affiliation":[{"name":"South China University of Technology, GuangZhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bolun","family":"Cai","sequence":"additional","affiliation":[{"name":"Tencent Wechat AI, GuangZhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weirui","family":"Lu","sequence":"additional","affiliation":[{"name":"South China University of Technology, GuangZhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,10,15]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1056\/NEJMra073096"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.143"},{"key":"e_1_3_2_1_3_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXivpreprint arXiv:1810.04805(2018).","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXivpreprint arXiv:1810.04805(2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXivpreprint arXiv:1810.04805(2018)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3266302.3268997"},{"key":"e_1_3_2_1_5_1","first-page":"190","article-title":"The Geneva minimalistic acoustic parameter set (GeMAPS)for voice research and affective computing","volume":"2","author":"Eyben Florian","year":"2015","unstructured":"Florian Eyben , Klaus R Scherer , Bj\u00f6rn W Schuller , Johan Sundberg , Elisabeth Andr\u00e9 , Carlos Busso , Laurence Y Devillers , Julien Epps , Petri Laukka , Shrikanth S Narayanan , 2015 . The Geneva minimalistic acoustic parameter set (GeMAPS)for voice research and affective computing . IEEE Transactions on Affective Com-puting7 , 2 (2015), 190 -- 202 . Florian Eyben, Klaus R Scherer, Bj\u00f6rn W Schuller, Johan Sundberg, Elisabeth Andr\u00e9, Carlos Busso, Laurence Y Devillers, Julien Epps, Petri Laukka, Shrikanth S Narayanan, et al. 2015. The Geneva minimalistic acoustic parameter set (GeMAPS)for voice research and affective computing. IEEE Transactions on Affective Com-puting7, 2 (2015), 190--202.","journal-title":"IEEE Transactions on Affective Com-puting7"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Fabien Ringeval and Bj\u00f6rn Schuller and Michel Valstar and Nicholas Cumminsand Roddy Cowie and Mohammad Soleymani and Maximilian Schmitt and ShahinAmiriparian and Eva-Maria Messner and Leili Tavabi and Siyang Song and Sina Alisamir and Shuo Lui and Ziping Zhao and Maja Pantic. 2019. AVEC 2019 Work-shop and Challenge: State-of-Mind Depression with AI and Cross-Cultural Affect Recognition. In Proceedings of the 9th International Workshop on Audio\/Visual Emotion Challenge AVEC'19 co-located with the 27th ACM International Conference on Multimedia MM 2019 Fabien Ringeval Bj\u00f6rn Schuller Michel Valstar Nicholas Cummins Roddy Cowie and Maja Pantic (Eds.). ACM Nice France.  Fabien Ringeval and Bj\u00f6rn Schuller and Michel Valstar and Nicholas Cumminsand Roddy Cowie and Mohammad Soleymani and Maximilian Schmitt and ShahinAmiriparian and Eva-Maria Messner and Leili Tavabi and Siyang Song and Sina Alisamir and Shuo Lui and Ziping Zhao and Maja Pantic. 2019. AVEC 2019 Work-shop and Challenge: State-of-Mind Depression with AI and Cross-Cultural Affect Recognition. In Proceedings of the 9th International Workshop on Audio\/Visual Emotion Challenge AVEC'19 co-located with the 27th ACM International Conference on Multimedia MM 2019 Fabien Ringeval Bj\u00f6rn Schuller Michel Valstar Nicholas Cummins Roddy Cowie and Maja Pantic (Eds.). ACM Nice France.","DOI":"10.1145\/3347320.3357688"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133944.3133945"},{"key":"e_1_3_2_1_8_1","volume-title":"Measuring Depression Symptom Severity from Spoken Language and 3D Facial Expressions.CoRRabs\/1811.08592","author":"Haque Albert","year":"2018","unstructured":"Albert Haque , Michelle Guo , Adam S. Miner , and Li Fei-Fei . 2018. Measuring Depression Symptom Severity from Spoken Language and 3D Facial Expressions.CoRRabs\/1811.08592 ( 2018 ). arXiv:1811.08592 http:\/\/arxiv.org\/abs\/1811.08592 Albert Haque, Michelle Guo, Adam S. Miner, and Li Fei-Fei. 2018. Measuring Depression Symptom Severity from Spoken Language and 3D Facial Expressions.CoRRabs\/1811.08592 (2018). arXiv:1811.08592 http:\/\/arxiv.org\/abs\/1811.08592"},{"key":"e_1_3_2_1_9_1","volume-title":"Joyce T Berry, and Ali H Mokdad.","author":"Kroenke Kurt","year":"2009","unstructured":"Kurt Kroenke , Tara W Strine , Robert L Spitzer , Janet BW Williams , Joyce T Berry, and Ali H Mokdad. 2009 . The PHQ- 8 as a measure of current depression in thegeneral population.Journal of affective disorders 114, 1--3 (2009), 163--173. Kurt Kroenke, Tara W Strine, Robert L Spitzer, Janet BW Williams, Joyce T Berry, and Ali H Mokdad. 2009. The PHQ-8 as a measure of current depression in thegeneral population.Journal of affective disorders 114, 1--3 (2009), 163--173."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.3115\/1118108.1118117"},{"key":"e_1_3_2_1_11_1","unstructured":"Lindasalwa Muda Mumtaj Begam and Irraivan Elamvazuthi. 2010. Voice recognition algorithms using mel frequency cepstral coefficient (MFCC) and dynamictime warping (DTW) techniques. arXiv preprint arXiv:1003.4083(2010).  Lindasalwa Muda Mumtaj Begam and Irraivan Elamvazuthi. 2010. Voice recognition algorithms using mel frequency cepstral coefficient (MFCC) and dynamictime warping (DTW) techniques. arXiv preprint arXiv:1003.4083(2010)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Matthew E Peters Mark Neumann Mohit Iyyer Matt Gardner Christopher Clark Kenton Lee and Luke Zettlemoyer. 2018. Deep contextualized word representations. arXiv preprint arXiv:1802.05365(2018).  Matthew E Peters Mark Neumann Mohit Iyyer Matt Gardner Christopher Clark Kenton Lee and Luke Zettlemoyer. 2018. Deep contextualized word representations. arXiv preprint arXiv:1802.05365(2018).","DOI":"10.18653\/v1\/N18-1202"},{"key":"e_1_3_2_1_13_1","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever.2018. Improving language understanding by generative pre-training. https:\/\/s3-us-west-2.amazonaws.com\/openai-assets\/research-covers\/languageunsupervised\/language understanding paper. pdf(2018).  Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever.2018. Improving language understanding by generative pre-training. https:\/\/s3-us-west-2.amazonaws.com\/openai-assets\/research-covers\/languageunsupervised\/language understanding paper. pdf(2018)."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133944.3133953"},{"key":"e_1_3_2_1_15_1","volume-title":"Fifteenth Annual Conference of the International Speech Communication Association.","author":"Schuller Bj\u00f6rn","year":"2014","unstructured":"Bj\u00f6rn Schuller , Stefan Steidl , Anton Batliner , Julien Epps , Florian Eyben , Fabien Ringeval , Erik Marchi , and Yue Zhang . 2014 . I (Special Session)******** The INTER-SPEECH 2014 Computational Paralinguistics Challenge: Cognitive & Physical Load . In Fifteenth Annual Conference of the International Speech Communication Association. Bj\u00f6rn Schuller, Stefan Steidl, Anton Batliner, Julien Epps, Florian Eyben, Fabien Ringeval, Erik Marchi, and Yue Zhang. 2014. I (Special Session)******** The INTER-SPEECH 2014 Computational Paralinguistics Challenge: Cognitive & Physical Load. In Fifteenth Annual Conference of the International Speech Communication Association."},{"key":"e_1_3_2_1_16_1","volume-title":"Proceedings of the 4th International Workshop on Audio\/Visual Emotion Challenge. ACM, 57--63","author":"Senoussaoui Mohammed","unstructured":"Mohammed Senoussaoui , Milton Sarria-Paja , Jo\u00e3o F Santos , and Tiago H Falk .2014. Model fusion for multimodal depression classification and level detection . In Proceedings of the 4th International Workshop on Audio\/Visual Emotion Challenge. ACM, 57--63 . Mohammed Senoussaoui, Milton Sarria-Paja, Jo\u00e3o F Santos, and Tiago H Falk.2014. Model fusion for multimodal depression classification and level detection. In Proceedings of the 4th International Workshop on Audio\/Visual Emotion Challenge. ACM, 57--63."},{"key":"e_1_3_2_1_17_1","volume-title":"Proceedings of the 2013 conferenceon empirical methods in natural language processing. 1631--1642","author":"Socher Richard","year":"2013","unstructured":"Richard Socher , Alex Perelygin , Jean Wu , Jason Chuang , Christopher D Manning , Andrew Ng , and Christopher Potts . 2013 . Recursive deep models for semantic compositionality over a sentiment treebank . In Proceedings of the 2013 conferenceon empirical methods in natural language processing. 1631--1642 . Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning,Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conferenceon empirical methods in natural language processing. 1631--1642."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133944.3133951"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2988257.2988258"},{"key":"e_1_3_2_1_21_1","volume-title":"ukasz Kaiser, and Illia Polosukhin","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , ukasz Kaiser, and Illia Polosukhin . 2017 . Attention is all you need. In Advances in neural information processing systems. 5998--6008. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2988257.2988263"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661806.2661809"},{"key":"e_1_3_2_1_24_1","volume-title":"Proceedings of the 2018 on Audio\/Visual Emotion Challenge and Workshop. ACM, 31--37","author":"Xing Xiaofen","unstructured":"Xiaofen Xing , Bolun Cai , Yinhu Zhao , Shuzhen Li , Zhiwei He , and Weiquan Fan .2018. Multi-modality Hierarchical Recall based on GBDTs for Bipolar Disorder Classification . In Proceedings of the 2018 on Audio\/Visual Emotion Challenge and Workshop. ACM, 31--37 . Xiaofen Xing, Bolun Cai, Yinhu Zhao, Shuzhen Li, Zhiwei He, and Weiquan Fan.2018. Multi-modality Hierarchical Recall based on GBDTs for Bipolar Disorder Classification. In Proceedings of the 2018 on Audio\/Visual Emotion Challenge and Workshop. ACM, 31--37."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2988257.2988269"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133944.3133948"},{"key":"e_1_3_2_1_27_1","volume-title":"Bipolar Disorder Recognition with Histogram Features of Arousal and Body Gestures. InProceedings of the 2018 on Audio\/Visual Emotion Challenge and Workshop. ACM, 15--21","author":"Yang Le","year":"2018","unstructured":"Le Yang , Yan Li , Haifeng Chen , Dongmei Jiang , Meshia C\u00e9dric Oveneke , and Hichem Sahli . 2018 . Bipolar Disorder Recognition with Histogram Features of Arousal and Body Gestures. InProceedings of the 2018 on Audio\/Visual Emotion Challenge and Workshop. ACM, 15--21 . Le Yang, Yan Li, Haifeng Chen, Dongmei Jiang, Meshia C\u00e9dric Oveneke, and Hichem Sahli. 2018. Bipolar Disorder Recognition with Histogram Features of Arousal and Body Gestures. InProceedings of the 2018 on Audio\/Visual Emotion Challenge and Workshop. ACM, 15--21."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133944.3133950"},{"key":"e_1_3_2_1_29_1","unstructured":"Fisher Yu and Vladlen Koltun. 2015. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122(2015).  Fisher Yu and Vladlen Koltun. 2015. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122(2015)."}],"event":{"name":"MM '19: The 27th ACM International Conference on Multimedia","location":"Nice France","acronym":"MM '19","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 9th International on Audio\/Visual Emotion Challenge and Workshop"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3347320.3357695","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3347320.3357695","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T19:05:46Z","timestamp":1750273546000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3347320.3357695"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,15]]},"references-count":29,"alternative-id":["10.1145\/3347320.3357695","10.1145\/3347320"],"URL":"https:\/\/doi.org\/10.1145\/3347320.3357695","relation":{},"subject":[],"published":{"date-parts":[[2019,10,15]]},"assertion":[{"value":"2019-10-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}