{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T04:19:18Z","timestamp":1760242758224,"version":"3.41.0"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2015,6,30]],"date-time":"2015-06-30T00:00:00Z","timestamp":1435622400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"NISHA project"},{"name":"SNSF UBImpressed project"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Interact. Intell. Syst."],"published-print":{"date-parts":[[2015,7,9]]},"abstract":"<jats:p>The prevalent \u201cshare what's on your mind\u201d paradigm of social media can be examined from the perspective of mood: short-term affective states revealed by the shared data. This view takes on new relevance given the emergence of conversational social video as a popular genre among viewers looking for entertainment and among video contributors as a channel for debate, expertise sharing, and artistic expression. From the perspective of human behavior understanding, in conversational social video both verbal and nonverbal information is conveyed by speakers and decoded by viewers. We present a systematic study of classification and ranking of mood impressions in social video, using vlogs from YouTube. Our approach considers eleven natural mood categories labeled through crowdsourcing by external observers on a diverse set of conversational vlogs. We extract a comprehensive number of nonverbal and verbal behavioral cues from the audio and video channels to characterize the mood of vloggers. Then we implement and validate vlog classification and vlog ranking tasks using supervised learning methods. Following a reliability and correlation analysis of the mood impression data, our study demonstrates that, while the problem is challenging, several mood categories can be inferred with promising performance. Furthermore, multimodal features perform consistently better than single-channel features. Finally, we show that addressing mood as a ranking problem is a promising practical direction for several of the mood categories studied.<\/jats:p>","DOI":"10.1145\/2641577","type":"journal-article","created":{"date-parts":[[2015,6,30]],"date-time":"2015-06-30T15:45:37Z","timestamp":1435679137000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["In the Mood for Vlog"],"prefix":"10.1145","volume":"5","author":[{"given":"Dairazalia","family":"Sanchez-Cortes","sequence":"first","affiliation":[{"name":"Idiap Research Institute, Martigny, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shiro","family":"Kumano","sequence":"additional","affiliation":[{"name":"NTT Communication Science Laboratories, Atsugi-shi, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kazuhiro","family":"Otsuka","sequence":"additional","affiliation":[{"name":"NTT Communication Science Laboratories, Atsugi-shi, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daniel","family":"Gatica-Perez","sequence":"additional","affiliation":[{"name":"Idiap Research Institute and Ecole Polytechnique F\u00e9d\u00e9rale de Lausanne (EPFL), Martigny, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,6,30]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1037\/0033-2909.111.2.256"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2037676.2037690"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of International Conference on Weblogs and Social Media.","author":"Biel Joan-Isaac","year":"2012","unstructured":"Joan-Isaac Biel and Daniel Gatica-Perez . 2012 . The good, the bad, and the angry: Analyzing crowdsourced impressions of vloggers . In Proceedings of International Conference on Weblogs and Social Media. Joan-Isaac Biel and Daniel Gatica-Perez. 2012. The good, the bad, and the angry: Analyzing crowdsourced impressions of vloggers. In Proceedings of International Conference on Weblogs and Social Media."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2012.2225032"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2522848.2522877"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2388676.2388689"},{"key":"e_1_2_1_7_1","unstructured":"Paul Boersma. 2002. Praat a system for doing phonetics by computer. Glot International 5 9\/10 341--345.  Paul Boersma. 2002. Praat a system for doing phonetics by computer. Glot International 5 9\/10 341--345."},{"volume-title":"Learning OpenCV: Computer vision with the OpenCV library","author":"Bradski Gary","key":"e_1_2_1_8_1","unstructured":"Gary Bradski and Adrian Kaehler . 2008. Learning OpenCV: Computer vision with the OpenCV library . O'Reilly Media, Inc. , Sebastopol, CA . Gary Bradski and Adrian Kaehler. 2008. Learning OpenCV: Computer vision with the OpenCV library. O'Reilly Media, Inc., Sebastopol, CA."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1010933404324"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the AAAI International Conference on Weblogs and Social Media.","author":"Choudhury Munmun De","year":"2012","unstructured":"Munmun De Choudhury , Scott Counts , and Michael Gamon . 2012 . Not all moods are created equal! Exploring human emotional states in social media . In Proceedings of the AAAI International Conference on Weblogs and Social Media. Munmun De Choudhury, Scott Counts, and Michael Gamon. 2012. Not all moods are created equal! Exploring human emotional states in social media. In Proceedings of the AAAI International Conference on Weblogs and Social Media."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/2503308.2503355"},{"key":"e_1_2_1_12_1","volume-title":"Friesen","author":"Ekman Paul","year":"2003","unstructured":"Paul Ekman and Wallace V . Friesen . 2003 . Unmasking the face: A guide to recognizing emotions from facial clues. ISHK, San Jose, CA. Paul Ekman and Wallace V. Friesen. 2003. Unmasking the face: A guide to recognizing emotions from facial clues. ISHK, San Jose, CA."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2005.10.010"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2436256.2436274"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/945365.964285"},{"key":"e_1_2_1_16_1","unstructured":"Felix Gillete. 2014. Hollywood's Big-Money YouTube Hit Factory. Bloomberg Business Week. Aug 28.  Felix Gillete. 2014. Hollywood's Big-Money YouTube Hit Factory. Bloomberg Business Week. Aug 28."},{"volume-title":"Proceedings of the IEEE International Conference and Workshops onAutomatic Face and Gesture Recognition (FG'13)","author":"Girard Jeffrey M.","key":"e_1_2_1_17_1","unstructured":"Jeffrey M. Girard , Jeffrey F. Cohn , Mohammad H. Mahoor , Seyedmohammad Mavadati , and Dean P. Rosenwald . 2013. Social risk and depression: Evidence from manual and automatic facial expression analysis . In Proceedings of the IEEE International Conference and Workshops onAutomatic Face and Gesture Recognition (FG'13) . Jeffrey M. Girard, Jeffrey F. Cohn, Mohammad H. Mahoor, Seyedmohammad Mavadati, and Dean P. Rosenwald. 2013. Social risk and depression: Evidence from manual and automatic facial expression analysis. In Proceedings of the IEEE International Conference and Workshops onAutomatic Face and Gesture Recognition (FG'13)."},{"key":"e_1_2_1_18_1","volume-title":"Macy","author":"Golder Scott A.","year":"2011","unstructured":"Scott A. Golder and Michael W . Macy . 2011 . Diurnal and seasonal mood vary with work, sleep, and daylength across diverse cultures. Science 333, 6051, 1878--1881. Scott A. Golder and Michael W. Macy. 2011. Diurnal and seasonal mood vary with work, sleep, and daylength across diverse cultures. Science 333, 6051, 1878--1881."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2011.2163395"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2205597"},{"volume-title":"Rank correlation methods","author":"Kendall Maurice George","key":"e_1_2_1_21_1","unstructured":"Maurice George Kendall . 1975. Rank correlation methods . Griffin , London, UK . Maurice George Kendall. 1975. Rank correlation methods. Griffin, London, UK."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/NLPKE.2009.5313734"},{"volume-title":"Nonverbal Communication in Human Interaction. Wadsworth","author":"Knapp Mark","key":"e_1_2_1_23_1","unstructured":"Mark Knapp and Judith Hall . 2008. Nonverbal Communication in Human Interaction. Wadsworth , Cengage Learning , Boston MA . Mark Knapp and Judith Hall. 2008. Nonverbal Communication in Human Interaction. Wadsworth, Cengage Learning, Boston MA."},{"volume-title":"Intraclass correlation coefficient. Encyclopedia of Statistical Sciences","author":"Koch Gary G.","key":"e_1_2_1_24_1","unstructured":"Gary G. Koch . 1982. Intraclass correlation coefficient. Encyclopedia of Statistical Sciences . John Wiley & Sons, Inc. Gary G. Koch. 1982. Intraclass correlation coefficient. Encyclopedia of Statistical Sciences. John Wiley & Sons, Inc."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2212776.2223776"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2004.838534"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1125451.1125646"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/FG.2011.5771414"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1322192.1322198"},{"key":"e_1_2_1_30_1","volume-title":"Retrieved","author":"LIWC.","year":"2007","unstructured":"LIWC. 2007 . LIWC Incorporation . Retrieved May 17, 2015 from http:\/\/www.liwc.net\/index.php. LIWC. 2007. LIWC Incorporation. Retrieved May 17, 2015 from http:\/\/www.liwc.net\/index.php."},{"key":"e_1_2_1_31_1","volume-title":"Retrieved","author":"Lowry Richard","year":"1998","unstructured":"Richard Lowry . 1998 . Concepts and applications of inferential statistics . Retrieved May 17, 2015 from http:\/\/vassarstats.net\/textbook\/. Richard Lowry. 1998. Concepts and applications of inferential statistics. Retrieved May 17, 2015 from http:\/\/vassarstats.net\/textbook\/."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2011.12.003"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.5555\/1622637.1622649"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/FG.2013.6553750"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2010.5583006"},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of SIGIR, Workshop on Stylistic Analysis of Text for Information Access.","author":"Mishne Gilad","year":"2005","unstructured":"Gilad Mishne . 2005 . Experiments with mood classification in blog posts . In Proceedings of SIGIR, Workshop on Stylistic Analysis of Text for Information Access. Gilad Mishne. 2005. Experiments with mood classification in blog posts. In Proceedings of SIGIR, Workshop on Stylistic Analysis of Text for Information Access."},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the AAAI Spring Symposium on Computational Approaches to Analysing Weblogs. 145--152","author":"Mishne Gilad","year":"2006","unstructured":"Gilad Mishne and Maarten de Rijke . 2006 . Capturing global mood levels using blog posts . In Proceedings of the AAAI Spring Symposium on Computational Approaches to Analysing Weblogs. 145--152 . Gilad Mishne and Maarten de Rijke. 2006. Capturing global mood levels using blog posts. In Proceedings of the AAAI Spring Symposium on Computational Approaches to Analysing Weblogs. 145--152."},{"key":"e_1_2_1_38_1","volume-title":"Retrieved","author":"Mislove Alan","year":"2015","unstructured":"Alan Mislove , Sune Lehmann , Yong-Yeol Ahn , Jukka-Pekka Onnela , and J. Niels Rosenquist . 2010. Pulse of the nation: US mood throughout the day inferred from Twitter . Retrieved May 17, 2015 from http:\/\/www.ccs.neu.edu\/home\/amislove\/twittermood\/. Alan Mislove, Sune Lehmann, Yong-Yeol Ahn, Jukka-Pekka Onnela, and J. Niels Rosenquist. 2010. Pulse of the nation: US mood throughout the day inferred from Twitter. Retrieved May 17, 2015 from http:\/\/www.ccs.neu.edu\/home\/amislove\/twittermood\/."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2070481.2070509"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-13672-6_28"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/T-AFFC.2011.9"},{"key":"e_1_2_1_42_1","unstructured":"OMRON. 2007. OKAO Vision. http:\/\/www.omron.com\/ecb\/products\/mobile\/.  OMRON. 2007. OKAO Vision. http:\/\/www.omron.com\/ecb\/products\/mobile\/."},{"key":"e_1_2_1_43_1","volume-title":"Retrieved","author":"Dictionaries Oxford","year":"2014","unstructured":"Oxford Dictionaries . 2014 . Oxford Online Dictionary . Retrieved May 17, 2015 from http:\/\/oxforddictionaries.com\/definition\/english\/mood. Oxford Dictionaries. 2014. Oxford Online Dictionary. Retrieved May 17, 2015 from http:\/\/oxforddictionaries.com\/definition\/english\/mood."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000011"},{"volume-title":"Linguistic Inquiry and Word Count: LIWC2001","author":"Pennebaker James","key":"e_1_2_1_45_1","unstructured":"James Pennebaker , Martha E. Francis , and Roger J. Booth . 2001 . Linguistic Inquiry and Word Count: LIWC2001 . Erlbaum Publishers, Mahwah, NJ . James Pennebaker, Martha E. Francis, and Roger J. Booth. 2001. Linguistic Inquiry and Word Count: LIWC2001. Erlbaum Publishers, Mahwah, NJ ."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1037\/0022-3514.77.6.1296"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2541831.2541864"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-6393(02)00084-5"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2011.01.011"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2388676.2388776"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2006.489"},{"volume-title":"Proceedings of Conference on Empirical Methods in International Conference on Natural Language Processing. Association for Computational Linguistics.","author":"Snow Rion","key":"e_1_2_1_52_1","unstructured":"Rion Snow , Brendan O'Connor , Daniel Jurafsky , and Andrew Y. Ng . 2008. Cheap and fast\u2014but is it good?: Evaluating non-expert annotations for natural language tasks . In Proceedings of Conference on Empirical Methods in International Conference on Natural Language Processing. Association for Computational Linguistics. Rion Snow, Brendan O'Connor, Daniel Jurafsky, and Andrew Y. Ng. 2008. Cheap and fast\u2014but is it good?: Evaluating non-expert annotations for natural language tasks. In Proceedings of Conference on Empirical Methods in International Conference on Natural Language Processing. Association for Computational Linguistics."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.5555\/1621474.1621487"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/1363686.1364052"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/FG.2011.5771374"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2013.34"},{"volume-title":"Retrieved","year":"2014","key":"e_1_2_1_57_1","unstructured":"YouTube. 2014 a. YouTube Channels . Retrieved May 17, 2015 from http:\/\/www.youtube.com\/channels. YouTube. 2014a. YouTube Channels. Retrieved May 17, 2015 from http:\/\/www.youtube.com\/channels."},{"volume-title":"Retrieved","year":"2014","key":"e_1_2_1_58_1","unstructured":"YouTube. 2014 b. YouTube Statistics . Retrieved May 17, 2015 from http:\/\/www.youtube.com\/yt\/press\/statistics.html. YouTube. 2014b. YouTube Statistics. Retrieved May 17, 2015 from http:\/\/www.youtube.com\/yt\/press\/statistics.html."}],"container-title":["ACM Transactions on Interactive Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2641577","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2641577","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T06:56:18Z","timestamp":1750229778000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2641577"}},"subtitle":["Multimodal Inference in Conversational Social Video"],"short-title":[],"issued":{"date-parts":[[2015,6,30]]},"references-count":58,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2015,7,9]]}},"alternative-id":["10.1145\/2641577"],"URL":"https:\/\/doi.org\/10.1145\/2641577","relation":{},"ISSN":["2160-6455","2160-6463"],"issn-type":[{"type":"print","value":"2160-6455"},{"type":"electronic","value":"2160-6463"}],"subject":[],"published":{"date-parts":[[2015,6,30]]},"assertion":[{"value":"2014-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-03-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-06-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}