{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T13:03:51Z","timestamp":1772802231126,"version":"3.50.1"},"reference-count":67,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2021,2,3]],"date-time":"2021-02-03T00:00:00Z","timestamp":1612310400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-1815274"],"award-info":[{"award-number":["CNS-1815274"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-1704899"],"award-info":[{"award-number":["CNS-1704899"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-11943396"],"award-info":[{"award-number":["CNS-11943396"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-1837022"],"award-info":[{"award-number":["CNS-1837022"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In this paper, we present HOMER, a cloud-based system for video highlight generation which enables the automated, relevant, and flexible segmentation of videos. Our system outperforms state-of-the-art solutions by fusing internal video content-based features with the user\u2019s emotion data. While current research mainly focuses on creating video summaries without the use of affective data, our solution achieves the subjective task of detecting highlights by leveraging human emotions. In two separate experiments, including videos filmed with a dual camera setup, and home videos randomly picked from Microsoft\u2019s Video Titles in the Wild (VTW) dataset, HOMER demonstrates an improvement of up to 38% in F1-score from baseline, while not requiring any external hardware. We demonstrated both the portability and scalability of HOMER through the implementation of two smartphone applications.<\/jats:p>","DOI":"10.3390\/s21041035","type":"journal-article","created":{"date-parts":[[2021,2,3]],"date-time":"2021-02-03T20:31:51Z","timestamp":1612384311000},"page":"1035","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Intelligent Video Highlights Generation with Front-Camera Emotion Sensing"],"prefix":"10.3390","volume":"21","author":[{"given":"Hugo","family":"Meyer","sequence":"first","affiliation":[{"name":"Department of Electrical Engineering, Columbia University, New York, NY 10027, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peter","family":"Wei","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, Columbia University, New York, NY 10027, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6480-0299","authenticated-orcid":false,"given":"Xiaofan","family":"Jiang","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, Columbia University, New York, NY 10027, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,2,3]]},"reference":[{"key":"ref_1","unstructured":"(2019, August 31). Cisco Visual Networking Index: Forecast and Trends, 2017\u20132022. Technical Report. Available online: https:\/\/www.cisco.com\/c\/en\/us\/solutions\/collateral\/service-provider\/visual-networking-index-vni\/white-paper-c11-741490.html."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"539","DOI":"10.1109\/TMM.2011.2131638","article-title":"Editing by Viewing: Automatic Home Video Summarization by Viewing Behavior Analysis","volume":"13","author":"Peng","year":"2011","journal-title":"IEEE Trans. Multimed."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Zhang, S., Tian, Q., Huang, Q., Gao, W., and Li, S. (2009, January 7\u201310). Utilizing affective analysis for efficient movie browsing. Proceedings of the 2009 16th IEEE International Conference on Image Processing (ICIP), Cairo, Egypt.","DOI":"10.1109\/ICIP.2009.5413590"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1126004.1126005","article-title":"Content-based Multimedia Information Retrieval: State of the Art and Challenges","volume":"2","author":"Lew","year":"2006","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl."},{"key":"ref_5","unstructured":"Hanjalic, A. (2003, January 14\u201317). Generic approach to highlights extraction from a sport video. Proceedings of the 2003 International Conference on Image Processing (Cat. No.03CH37429), Barcelona, Spain."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1114","DOI":"10.1109\/TMM.2005.858397","article-title":"Adaptive extraction of highlights from a sport video based on excitement modeling","volume":"7","author":"Hanjalic","year":"2005","journal-title":"IEEE Trans. Multimed."},{"key":"ref_7","unstructured":"Assfalg, J., Bertini, M., Colombo, C., Bimbo, A.D., and Nunziati, W. (2003, January 14\u201317). Automatic extraction and annotation of soccer video highlights. Proceedings of the 2003 International Conference on Image Processing (Cat. No.03CH37429), Barcelona, Spain."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Chakraborty, P.R., Zhang, L., Tjondronegoro, D., and Chandran, V. (2015, January 23\u201326). Using Viewer\u2019s Facial Expression and Heart Rate for Sports Video Highlights Detection. Proceedings of the 5th ACM on International Conference on Multimedia Retrieval, Shanghai, China.","DOI":"10.1145\/2671188.2749361"},{"key":"ref_9","unstructured":"Butler, D., and Ortutay, B. (2019). Facebook Auto-Generates Videos Celebrating Extremist Images, AP News."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Joho, H., Jose, J.M., Valenti, R., and Sebe, N. (2009, January 8\u201310). Exploiting Facial Expressions for Affective Video Summarisation. Proceedings of the ACM International Conference on Image and Video Retrieval, Santorini Island, Greece.","DOI":"10.1145\/1646396.1646435"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"505","DOI":"10.1007\/s11042-010-0632-x","article-title":"Looking at the viewer: Analysing facial activity to detect personal highlights of multimedia contents","volume":"51","author":"Joho","year":"2011","journal-title":"Multimed. Tools Appl."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1186\/s13634-019-0611-y","article-title":"A bottom-up summarization algorithm for videos in the wild","volume":"2019","author":"Pan","year":"2019","journal-title":"EURASIP J. Adv. Signal Process."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Al Nahian, M., Iftekhar, A.S.M., Islam, M., Rahman, S.M.M., and Hatzinakos, D. (2017, January 11\u201313). CNN-Based Prediction of Frame-Level Shot Importance for Video Summarization. Proceedings of the 2017 International Conference on New Trends in Computing Sciences (ICTCS), Amman, Jordan.","DOI":"10.1109\/ICTCS.2017.13"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"907","DOI":"10.1109\/TMM.2005.854410","article-title":"A generic framework of user attention model and its application in video summarization","volume":"7","author":"Ma","year":"2005","journal-title":"IEEE Trans. Multimed."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhang, K., Chao, W., Sha, F., and Grauman, K. (2016). Video Summarization with Long Short-term Memory. Computer Vision\u2014ECCV 2016, Springer.","DOI":"10.1007\/978-3-319-46478-7_47"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Lai, S.H., Lepetit, V., Nishino, K., and Sato, Y. (2017). Video Summarization Using Deep Semantic Features. Computer Vision\u2014ACCV 2016, Springer International Publishing.","DOI":"10.1007\/978-3-319-54190-7"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yang, H., Wang, B., Lin, S., Wipf, D.P., Guo, M., and Guo, B. (2015, January 7\u201313). Unsupervised Extraction of Video Highlights Via Robust Recurrent Auto-encoders. Proceedings of the IEEE International Conference on Computer Vision 2015, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.526"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Mahasseni, B., Lam, M., and Todorovic, S. (2017, January 21\u201326). Unsupervised Video Summarization With Adversarial LSTM Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.318"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Fleet, D., Pajdla, T., Schiele, B., and Tuytelaars, T. (2014). Ranking Domain-Specific Highlights by Analyzing Edited Videos. Computer Vision\u2014ECCV 2014, Springer International Publishing.","DOI":"10.1007\/978-3-319-10578-9"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"410","DOI":"10.1109\/TAFFC.2015.2432791","article-title":"Video Affective Content Analysis: A Survey of State-of-the-Art Methods","volume":"6","author":"Wang","year":"2015","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1257","DOI":"10.1007\/s11042-013-1450-8","article-title":"Hybrid video emotional tagging using users\u2019 EEG and video content","volume":"72","author":"Wang","year":"2014","journal-title":"Multimed. Tools Appl."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1109\/T-AFFC.2011.37","article-title":"Multimodal Emotion Recognition in Response to Videos","volume":"3","author":"Soleymani","year":"2012","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1109\/TAFFC.2015.2436926","article-title":"Analysis of EEG Signals and Facial Expressions for Continuous Emotion Detection","volume":"7","author":"Soleymani","year":"2016","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Fleureau, J., Guillotel, P., and Orlac, I. (2013, January 3\u20135). Affective Benchmarking of Movies Based on the Physiological Responses of a Real Audience. Proceedings of the 2013 Humaine Association Conference on Affective Computing and Intelligent Interaction, Geneva, Switzerland.","DOI":"10.1109\/ACII.2013.19"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"4679","DOI":"10.1007\/s11042-013-1830-0","article-title":"Implicit video emotion tagging from audiences\u2019 facial expression","volume":"74","author":"Wang","year":"2015","journal-title":"Multimed. Tools Appl."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1016\/j.jvcir.2007.04.002","article-title":"Video summarisation: A conceptual framework and survey of the state of the art","volume":"19","author":"Money","year":"2008","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Shukla, P., Sadana, H., Bansal, A., Verma, D., Elmadjian, C., Raman, B., and Turk, M. (2018, January 18\u201322). Automatic cricket highlight generation using event-driven and excitement-based features. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00233"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"576","DOI":"10.1109\/TMM.2006.888013","article-title":"Generation of Personalized Music Sports Video Using Multimodal Cues","volume":"9","author":"Wang","year":"2007","journal-title":"IEEE Trans. Multimed."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Yao, T., Mei, T., and Rui, Y. (July, January 26). Highlight Detection with Pairwise Deep Ranking for First-Person Video Summarization. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.112"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Panda, R., Das, A., Wu, Z., Ernst, J., and Roy-Chowdhury, A.K. (2017, January 22\u201329). Weakly Supervised Summarization of Web Videos. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.395"},{"key":"ref_31","unstructured":"Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N.D., and Weinberger, K.Q. (2014). Diverse Sequential Subset Selection for Supervised Video Summarization. Advances in Neural Information Processing Systems 27, Citeseer."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Sharghi, A., Gong, B., and Shah, M. (2016, January 8\u201316). Query-Focused Extractive Video Summarization. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_1"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhang, K., Chao, W., Sha, F., and Grauman, K. (2016, January 27\u201330). Summary Transfer: Exemplar-based Subset Selection for Video Summarization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2016, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.120"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Gygli, M., Grabner, H., and Van Gool, L. (2015, January 7\u201312). Video summarization by learning submodular mixtures of objectives. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298928"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Mor\u00e8re, O., Goh, H., Veillard, A., Chandrasekhar, V., and Lin, J. (2015, January 27\u201330). Co-regularized deep representations for video summarization. Proceedings of the 2015 IEEE International Conference on Image Processing (ICIP), Quebec City, QC, Canada.","DOI":"10.1109\/ICIP.2015.7351387"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1016\/j.patrec.2010.08.004","article-title":"VSUMM: A mechanism designed to produce static video summaries and a novel evaluation method","volume":"32","author":"Lopes","year":"2011","journal-title":"Pattern Recognit. Lett."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Khosla, A., Hamid, R., Lin, C., and Sundaresan, N. (2013, January 23\u201328). Large-Scale Video Summarization Using Web-Image Priors. Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.348"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1007\/s00799-005-0129-9","article-title":"Keyframe-based video summarization using Delaunay clustering","volume":"6","author":"Mundur","year":"2006","journal-title":"Int. J. Digit. Libr."},{"key":"ref_39","unstructured":"Ngo, C.-W., Ma, Y.-T., and Zhang, H.-J. (2003, January 13\u201316). Automatic video summarization by graph modeling. Proceedings of the Ninth IEEE International Conference on Computer Vision, Nice, France."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Lu, Z., and Grauman, K. (2013, January 23\u201328). Story-Driven Summarization for Egocentric Video. Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.350"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Nie, J., Hu, Y., Wang, Y., Xia, S., and Jiang, X. (2020, January 21\u201324). SPIDERS: Low-Cost Wireless Glasses for Continuous In-Situ Bio-Signal Acquisition and Emotion Recognition. Proceedings of the 2020 IEEE\/ACM Fifth International Conference on Internet-of-Things Design and Implementation (IoTDI), Sydney, Australia.","DOI":"10.1109\/IoTDI49375.2020.00011"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Ramzan, N., van Zwol, R., Lee, J.S., Cl\u00fcver, K., and Hua, X.S. (2013). Highlight Detection in Movie Scenes Through Inter-users, Physiological Linkage. Social Media Retrieval, Springer.","DOI":"10.1007\/978-1-4471-4555-4"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Fi\u00e3o, G., Rom\u00e3o, T., Correia, N., Centieiro, P., and Dias, A.E. (2016, January 9\u201312). Automatic Generation of Sport Video Highlights Based on Fan\u2019s Emotions and Content. Proceedings of the 13th International Conference on Advances in Computer Entertainment Technology, Osaka, Japan.","DOI":"10.1145\/3001773.3001802"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Ringer, C., and Nicolaou, M.A. (2018, January 7\u201310). Deep unsupervised multi-view detection of video game stream highlights. Proceedings of the 13th International Conference on the Foundations of Digital Games, Malm\u00f6, Sweden.","DOI":"10.1145\/3235765.3235781"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1016\/j.techfore.2017.07.011","article-title":"A neuro-advertising property video recommendation system","volume":"131","author":"Kaklauskas","year":"2018","journal-title":"Technol. Forecast. Soc. Chang."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"357","DOI":"10.24846\/v28i3y201912","article-title":"INVAR Neuromarketing Method and System","volume":"28","author":"Kaklauskas","year":"2019","journal-title":"Stud. Inform. Control"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Gunawardena, P., Amila, O., Sudarshana, H., Nawaratne, R., Luhach, A.K., Alahakoon, D., Perera, A.S., Chitraranjan, C., Chilamkurti, N., and De Silva, D. (2020). Real-time automated video highlight generation with dual-stream hierarchical growing self-organizing maps. J. Real Time Image Process., 147.","DOI":"10.1007\/s11554-020-00957-0"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"376","DOI":"10.1016\/j.patrec.2018.07.030","article-title":"Unsupervised object-level video summarization with online motion auto-encoder","volume":"130","author":"Zhang","year":"2020","journal-title":"Pattern Recognit. Lett."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Moses, T.M., and Balachandran, K. (2019, January 1\u20132). A Deterministic Key-Frame Indexing and Selection for Surveillance Video Summarization. Proceedings of the 2019 International Conference on Data Science and Communication (IconDSC), Bangalore, India.","DOI":"10.1109\/IconDSC.2019.8816901"},{"key":"ref_50","unstructured":"Lien, J.J., Kanade, T., Cohn, J.F., and Ching-Chung, L. (1998, January 14\u201316). Automated facial expression recognition based on FACS action units. Proceedings of the Third IEEE International Conference on Automatic Face and Gesture Recognition, Nara, Japan."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1016\/S0921-8890(99)00103-7","article-title":"Detection, tracking, and classification of action units in facial expression","volume":"31","author":"Lien","year":"2000","journal-title":"Robot. Auton. Syst."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1007\/s12193-015-0195-2","article-title":"EmoNets: Multimodal deep learning approaches for emotion recognition in video","volume":"10","author":"Kahou","year":"2016","journal-title":"J. Multimodal User Interfaces"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1109\/TAFFC.2017.2740923","article-title":"AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild","volume":"10","author":"Mollahosseini","year":"2019","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_54","first-page":"481","article-title":"Video frames similarity function based gaussian video segmentation and summarization","volume":"10","author":"Zhang","year":"2014","journal-title":"Int. J. Innov. Comput. Inf. Control"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Cakir, E., Heittola, T., Huttunen, H., and Virtanen, T. (2015, January 12\u201316). Polyphonic sound event detection using multi label deep neural networks. Proceedings of the 2015 International Joint Conference on Neural Networks (IJCNN), Killarney, Ireland.","DOI":"10.1109\/IJCNN.2015.7280624"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Parascandolo, G., Huttunen, H., and Virtanen, T. (2016, January 20\u201325). Recurrent neural networks for polyphonic sound event detection in real life recordings. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472917"},{"key":"ref_57","unstructured":"Gorin, A., Makhazhanov, N., and Shmyrev, N. (2016, January 3). DCASE 2016 sound event detection system based on convolutional neural network. Proceedings of the IEEE AASP Challenge: Detection and Classification of Acoustic Scenes and Events, Budapest, Hungary."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Wagner, J., Schiller, D., Seiderer, A., and Andr\u00e9, E. (2018, January 2\u20136). Deep Learning in Paralinguistic Recognition Tasks: Are Hand-crafted Features Still Relevant?. Proceedings of the Interspeech, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1238"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Choi, Y., Atif, O., Lee, J., Park, D., and Chung, Y. (2018). Noise-Robust Sound-Event Classification System with Texture Analysis. Symmetry, 10.","DOI":"10.3390\/sym10090402"},{"key":"ref_60","unstructured":"Arroyo, I., Cooper, D.G., Burleson, W., Woolf, B.P., Muldner, K., and Christopherson, R. (2009, January 6\u201310). Emotion Sensors Go To School. Proceedings of the 2009 Conference on Artificial Intelligence in Education: Building Learning Systems That Care: From Knowledge Representation to Affective Modelling, Brighton, UK."},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"724","DOI":"10.1016\/j.ijhcs.2007.02.003","article-title":"Automatic prediction of frustration","volume":"65","author":"Kapoor","year":"2007","journal-title":"Int. J. Hum. Comput. Stud."},{"key":"ref_62","unstructured":"Castellano, G., Kessous, L., and Caridakis, G. (2008). Affect and Emotion in Human-Computer Interaction, Springer. Chapter Emotion Recognition Through Multiple Modalities: Face, Body Gesture, Speech."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Kang, H.B. (2003, January 2\u20138). Affective content detection using HMMs. Proceedings of the eleventh ACM international conference on Multimedia, Berkeley, CA, USA.","DOI":"10.1145\/957013.957066"},{"key":"ref_64","doi-asserted-by":"crossref","first-page":"2553","DOI":"10.1016\/j.neucom.2007.11.043","article-title":"User and context adaptive neural networks for emotion recognition","volume":"71","author":"Caridakis","year":"2008","journal-title":"Neurocomputing"},{"key":"ref_65","doi-asserted-by":"crossref","first-page":"328","DOI":"10.1177\/1555412018788161","article-title":"Watching Players: An Exploration of Media Enjoyment on Twitch","volume":"15","author":"Wulf","year":"2020","journal-title":"Games Cult."},{"key":"ref_66","doi-asserted-by":"crossref","first-page":"985","DOI":"10.1016\/j.chb.2016.10.019","article-title":"Why do people watch others play video games? An empirical study on the motivations of Twitch users","volume":"75","author":"Hamari","year":"2017","journal-title":"Comput. Hum. Behav."},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Zeng, K.H., Chen, T.H., Niebles, J.C., and Sun, M. (2016). Title Generation for User Generated Videos. arXiv.","DOI":"10.1007\/978-3-319-46475-6_38"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/4\/1035\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:19:28Z","timestamp":1760159968000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/4\/1035"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2,3]]},"references-count":67,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2021,2]]}},"alternative-id":["s21041035"],"URL":"https:\/\/doi.org\/10.3390\/s21041035","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,2,3]]}}}