{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,27]],"date-time":"2025-12-27T07:24:47Z","timestamp":1766820287355,"version":"build-2065373602"},"reference-count":80,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2024,8,31]],"date-time":"2024-08-31T00:00:00Z","timestamp":1725062400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62276240","CUC23ZDTJ010"],"award-info":[{"award-number":["62276240","CUC23ZDTJ010"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["62276240","CUC23ZDTJ010"],"award-info":[{"award-number":["62276240","CUC23ZDTJ010"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Most existing intelligent editing tools for music and video rely on the cross-modal matching technology of the affective consistency or the similarity of feature representations. However, these methods are not fully applicable to complex audiovisual matching scenarios, resulting in low matching accuracy and suboptimal audience perceptual effects due to ambiguous matching rules and associated factors. To address these limitations, this paper focuses on both the similarity and integration of affective distribution for the artistic audiovisual works of movie and television video and music. Based on the rich emotional perception elements, we propose a hybrid matching model based on feature canonical correlation analysis (CCA) and fine-grained affective similarity. The model refines KCCA fusion features by analyzing both matched and unmatched music\u2013video pairs. Subsequently, the model employs XGBoost to predict relevance and to compute similarity by considering fine-grained affective semantic distance as well as affective factor distance. Ultimately, the matching prediction values are obtained through weight allocation. Experimental results on a self-built dataset demonstrate that the proposed affective matching model balances feature parameters and affective semantic cognitions, yielding relatively high prediction accuracy and better subjective experience of audiovisual association. This paper is crucial for exploring the affective association mechanisms of audiovisual objects from a sensory perspective and improving related intelligent tools, thereby offering a novel technical approach to retrieval and matching in music\u2013video editing.<\/jats:p>","DOI":"10.3390\/s24175681","type":"journal-article","created":{"date-parts":[[2024,9,2]],"date-time":"2024-09-02T12:54:42Z","timestamp":1725281682000},"page":"5681","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["An Audiovisual Correlation Matching Method Based on Fine-Grained Emotion and Feature Fusion"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2680-4675","authenticated-orcid":false,"given":"Zhibin","family":"Su","sequence":"first","affiliation":[{"name":"State Key Laboratory of Media Convergence and Communication, Beijing 100024, China"},{"name":"Key Laboratory of Acoustic Visual Technology and Intelligent Control System, Ministry of Culture and Tourism, Beijing 100024, China"},{"name":"School of Information and Communication Engineering, Communication University of China, Beijing 100024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yiming","family":"Feng","sequence":"additional","affiliation":[{"name":"Key Laboratory of Acoustic Visual Technology and Intelligent Control System, Ministry of Culture and Tourism, Beijing 100024, China"},{"name":"School of Information and Communication Engineering, Communication University of China, Beijing 100024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinyu","family":"Liu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Acoustic Visual Technology and Intelligent Control System, Ministry of Culture and Tourism, Beijing 100024, China"},{"name":"School of Information and Communication Engineering, Communication University of China, Beijing 100024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jing","family":"Peng","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Communication University of China, Beijing 100024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Jiang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Media Convergence and Communication, Beijing 100024, China"},{"name":"Key Laboratory of Acoustic Visual Technology and Intelligent Control System, Ministry of Culture and Tourism, Beijing 100024, China"},{"name":"School of Information and Communication Engineering, Communication University of China, Beijing 100024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingyu","family":"Liu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Media Convergence and Communication, Beijing 100024, China"},{"name":"Key Laboratory of Acoustic Visual Technology and Intelligent Control System, Ministry of Culture and Tourism, Beijing 100024, China"},{"name":"School of Information and Communication Engineering, Communication University of China, Beijing 100024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,8,31]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Pang, N., Guo, S., Yan, M., and Chan, C.A. (2023). A Short Video Classification Framework Based on Cross-Modal Fusion. Sensors, 23.","DOI":"10.3390\/s23208425"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Tao, R., Zhu, M., Cao, H., and Ren, H. (2024). Fine-Grained Cross-Modal Semantic Consistency in Natural Conservation Image Data from a Multi-Task Perspective. Sensors, 24.","DOI":"10.20944\/preprints202404.0847.v1"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Chatterjee, S., Mishra, J., Sundram, F., and Roop, P. (2024). Towards Personalised Mood Prediction and Explanation for Depression from Biophysical Data. Sensors, 24.","DOI":"10.3390\/s24010164"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Hehenkamp, N., Rizzi, F.G., Grundh\u00f6fer, L., and Gewies, S. (2024). Prediction of Ground Wave Propagation Delay for MF R-Mode. Sensors, 24.","DOI":"10.3390\/s24010282"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Leung, R. (2023). Using AI\u2013ML to Augment the Capabilities of Social Media for Telehealth and Remote Patient Monitoring. Healthcare, 11.","DOI":"10.3390\/healthcare11121704"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1049\/cit2.12153","article-title":"A Semantic and Emotion-based Dual Latent Variable Generation Model for a Dialogue System","volume":"8","author":"Yan","year":"2023","journal-title":"CAAI Trans. Intell. Technol."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"765","DOI":"10.1007\/s11042-019-08192-x","article-title":"Recognition of emotion in music based on deep convolutional neural network","volume":"79","author":"Sarkar","year":"2020","journal-title":"Multimed. Tools Appl."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Mittal, T., Guhan, P., Bhattacharya, U., Chandra, B., Bera, A., and Manocha, D. (2020, January 13\u201319). Emoticon: Context-aware multimodal emotion recognition using frege\u2019s principle. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01424"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"824","DOI":"10.1109\/TPAMI.2013.225","article-title":"Multimodal Similarity-Preserving Hashing","volume":"36","author":"Masci","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Li, T., Sun, Z., Zhang, H., Sun, Z., Li, J., and Wu, Z. (2021, January 11\u201315). Deep music retrieval for fine-grained videos by exploiting cross-modal-encoded voice-overs. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), New York, NY, USA.","DOI":"10.1145\/3404835.3462993"},{"key":"ref_11","first-page":"1","article-title":"Audeosynth: Music-driven video montage","volume":"34","author":"Liao","year":"2015","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Nakatsuka, T., Hamasaki, M., and Goto, M. (2023, January 2\u20137). Content-Based Music-Image Retrieval Using Self-and Cross-Modal Feature Embedding Memory. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA.","DOI":"10.1109\/WACV56688.2023.00221"},{"key":"ref_13","unstructured":"Picard, R.W. (2000). Affective Computing, MIT Press. Technical Report."},{"key":"ref_14","unstructured":"Chen, C.H., Weng, M.F., Jeng, S.K., and Ghuang, Y.Y. (2008, January 9\u201311). Emotion-based Music Visualization using Photos. Proceedings of the Advances in Multimedia Modeling: 14th International Multimedia Modeling Conference (MMM), Kyoto, Japan."},{"key":"ref_15","first-page":"93","article-title":"An automatic music classification method based on emotion","volume":"10","author":"Su","year":"2021","journal-title":"Inf. Technol."},{"key":"ref_16","unstructured":"Zhan, C., She, D., Zhao, S., Cheng, M., and Yang, J. (November, January 27). Zero-shot emotion recognition via affective structural embedding. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"106","DOI":"10.1109\/MSP.2021.3106232","article-title":"Music Emotion Recognition: Toward new, robust standards in personalized and context-sensitive applications","volume":"38","author":"Cano","year":"2021","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Yu, X. (2021, January 9\u201311). Adaptability of Simple Classifier and Active Learning in Music Emotion Recognition. Proceedings of the 4th International Conference on Electronics, Communications and Control Engineering (ICECC), New York, NY, USA.","DOI":"10.1145\/3462676.3462679"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, J., Wen, X., Cho, A., and Whang, M. (2021). An Empathy Evaluation System Using Spectrogram Image Features of Audio. Sensors, 21.","DOI":"10.3390\/s21217111"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Al-Saadawi, H.F.T., and Das, R. (2024). TER-CA-WGNN: Trimodel Emotion Recognition Using Cumulative Attribute-Weighted Graph Neural Network. Appl. Sci., 14.","DOI":"10.3390\/app14062252"},{"key":"ref_21","first-page":"373","article-title":"Music Emotion Recognition based on Wide Deep Learning Networks","volume":"48","author":"Wang","year":"2022","journal-title":"J. East China Univ. Sci. Technol."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"695","DOI":"10.1587\/transinf.2019EDP7175","article-title":"Combining CNN and Broad Learning for Music Classification","volume":"103","author":"Tang","year":"2020","journal-title":"IEICE Trans. Inf. Syst."},{"key":"ref_23","first-page":"91","article-title":"Classification of music emotion appreciation based on forward neural network multi-feature fusion algorithm","volume":"37","author":"Ning","year":"2021","journal-title":"Microcomput. Appl."},{"key":"ref_24","first-page":"760","article-title":"Music emotion recognition using convolutional long short term memory deep neural networks","volume":"24","author":"Hizlisoy","year":"2021","journal-title":"Eng. Sci. Technol. Int. J."},{"key":"ref_25","first-page":"10","article-title":"Music emotion recognition fusion on CNN-BiLSTM and self-attention model","volume":"59","author":"Zhong","year":"2023","journal-title":"Comput. Eng. Appl."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, Y., Wu, J., Heracleous, P., Wada, S., Kimura, R., and Kurihara, S. (2020, January 25\u201329). Implicit knowledge injectable cross attention audiovisual model for group emotion recognition. Proceedings of the 2020 International Conference on Multimodal Interaction (ICMI), Utrecht, The Netherlands.","DOI":"10.1145\/3382507.3417960"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Huang, R., Chen, Z., He, J., and Chu, X. (2022). Dynamic Heterogeneous User Generated Contents-Driven Relation Assessment via Graph Representation Learning. Sensors, 22.","DOI":"10.3390\/s22041402"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Ma, Y., Xu, Y., Liu, Y., Yan, F., Zhang, Q., Li, Q., and Liu, Q. (2024). Multi-Scale Cross-Attention Fusion Network Based on Image Super-Resolution. Appl. Sci., 14.","DOI":"10.3390\/app14062634"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Xu, H., Jiang, C., Liang, X., and Li, Z. (2019, January 15\u201320). Spatial-aware graph relation network for large-scale object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00952"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Chaudhari, A., Bhatt, C., Krishna, A., and Mazzeo, P.L. (2022). ViTFER: Facial Emotion Recognition with Vision Transformers. Appl. Syst. Innov., 5.","DOI":"10.3390\/asi5040080"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Roshdy, A., Karar, A., Kork, S.A., Beyrouthy, T., and Nait-ali, A. (2024). Advancements in EEG Emotion Recognition: Leveraging Multi-Modal Database Integration. Appl. Sci., 14.","DOI":"10.3390\/app14062487"},{"key":"ref_32","first-page":"120","article-title":"Research on emotion recognition method based on audio and video feature fusion","volume":"36","author":"Tie","year":"2022","journal-title":"J. Chongqing Univ. Technol. (Nat. Sci.)"},{"key":"ref_33","first-page":"1250","article-title":"Reasearch on image sentiment analysis based on muti-visual object fusion","volume":"38","author":"Liao","year":"2021","journal-title":"Appl. Res. Comput."},{"key":"ref_34","unstructured":"Lee, J., Kim, S., Kim, S., Park, J., and Sohn, K. (November, January 27). Context-aware emotion recognition networks. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Avramidis, K., Stewart, S., and Narayanan, S. (2023, January 4\u201310). On the role of visual context in enriching music representations. Proceedings of the ICASSP 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece.","DOI":"10.1109\/ICASSP49357.2023.10094915"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"94273","DOI":"10.1109\/ACCESS.2022.3204305","article-title":"Self-Supervised Learning for Audio-Visual Relationships of Videos with Stereo Sounds","volume":"10","author":"Sato","year":"2022","journal-title":"IEEE Access"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Sur\u00eds, D., Vondrick, C., Russell, B., and Salamon, J. (2022, January 18\u201324). It\u2019s time for artistic correspondence in music and video. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01031"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Cao, Y., Long, M., Wang, J., Yang, Q., and Yu, P.S. (2016, January 13\u201317). Deep visual-semantic hashing for cross-modal retrieval. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA.","DOI":"10.1145\/2939672.2939812"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Shen, Y., Liu, L., Shao, L., and Song, J. (2017, January 22\u201329). Deep binaries: Encoding semantic-rich cues for efficient textual-visual cross retrieval. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.441"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1089","DOI":"10.1109\/TPAMI.2016.2567386","article-title":"Cross-Domain Visual Matching via Generalized Similarity Measure and Feature Learning","volume":"39","author":"Liang","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"401","DOI":"10.1109\/TCSVT.2020.2974877","article-title":"Deep multiscale fusion hashing for cross-modal retrieval","volume":"31","author":"Nie","year":"2020","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Rasiwasia, N., Costa Pereira, J., Coviello, E., Doyle, G., Lanckriet, G.R., Levy, R., and Vasconcelos, N. (2010, January 25\u201329). A new approach to cross-modal multimedia retrieval. Proceedings of the 18th ACM International Conference on Multimedia, Firenze, Italy.","DOI":"10.1145\/1873951.1873987"},{"key":"ref_43","unstructured":"Andrew, G., Arora, R., Bilmes, J., and Livescu, K. (2013, January 17\u201319). Deep canonical correlation analysis. Proceedings of the International Conference on Machine Learning, PMLR, Atlanta, GA, USA."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"618","DOI":"10.1016\/j.neucom.2016.06.047","article-title":"Deep canonical correlation analysis with progressive and hypergraph learning for cross-modal retrieval","volume":"214","author":"Shao","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_45","first-page":"1","article-title":"Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-modal Retrieval","volume":"16","author":"Zeng","year":"2020","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl. (TOMM)"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Vukoti\u0107, V., Raymond, C., and Gravier, G. (2016, January 6\u20139). Bidirectional joint representation learning with symmetrical deep neural networks for multimodal and crossmodal applications. Proceedings of the 2016 ACM on International Conference on Multimedia Retrieval, New York, NY, USA.","DOI":"10.1145\/2911996.2912064"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"649","DOI":"10.14778\/2732296.2732301","article-title":"Effective multi-modal retrieval based on stacked auto-encoders","volume":"7","author":"Wei","year":"2014","journal-title":"Proc. VLDB Endow."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"1363","DOI":"10.1109\/TMM.2016.2558463","article-title":"Cross-Modal Retrieval via Deep and Bidirectional Representation Learning","volume":"18","author":"He","year":"2016","journal-title":"IEEE Trans. Multimed."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Hua, Y., Tian, H., Cai, A., and Shi, P. (2015, January 13\u201316). Cross-modal correlation learning with deep convolutional architecture. Proceedings of the 2015 Visual Communications and Image Processing (VCIP), Singapore.","DOI":"10.1109\/VCIP.2015.7457841"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Zhang, J., Peng, Y., and Yuan, M. (2018, January 2\u20137). Unsupervised generative adversarial cross-modal hashing. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11263"},{"key":"ref_51","unstructured":"Nawaz, S., Janjua, M.K., Calefati, A., and Gallo, L. (2018). Revisiting cross modal retrieval. axXiv."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Gu, J., Cai, J., Joty, S.R., Niu, L., and Wang, G. (2018, January 18\u201323). Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00750"},{"key":"ref_53","unstructured":"Su, S., Zhong, Z., and Zhang, C. (November, January 27). Deep joint-semantics reconstructing hashing for large-scale unsupervised cross-modal retrieval. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_54","unstructured":"Li, C., Deng, C., Wang, L., Xie, D., and Liu, X. (February, January 27). Coupled cyclegan: Unsupervised hashing network for cross-modal retrieval. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Wang, P., Wang, X., Wang, Z., and Dong, Y. (2024). Learning Accurate Pseudo-Labels via Feature Similarity in the Presence of Label Noise. Appl. Sci., 14.","DOI":"10.3390\/app14072759"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Wang, P., Liu, S., and Chen, J. (2024). CCDA: A Novel Method to Explore the Cross-Correlation in Dual-Attention for Multimodal Sentiment Analysis. Appl. Sci., 14.","DOI":"10.3390\/app14051934"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Zhao, X., Li, X., Tie, Y., Hu, Z., and Qi, L. (2023, January 10\u201314). Video Background Music Recommendation Based on Multi-level Fusion Features. Proceedings of the 2023 IEEE International Conference on Multimedia and Expo Workshops (ICMEW), Brisbane, Australia.","DOI":"10.1109\/ICMEW59549.2023.00076"},{"key":"ref_58","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3592614","article-title":"Learning explicit and implicit dual common subspaces for audio-visual cross-modal retrieval","volume":"19","author":"Zeng","year":"2023","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl."},{"key":"ref_59","unstructured":"Rasiwasia, N., Mahajan, D., Mahadevan, V., and Aggarwal, G. (2014, January 22\u201325). Cluster Canonical Correlation Analysis. Proceedings of the Seventeenth International Conference on Artificial Intelligence and Statistics, PMLR, Reykjavik, Iceland."},{"key":"ref_60","first-page":"911","article-title":"A kernel method for canonical correlation analysis","volume":"18","author":"Akaho","year":"2006","journal-title":"Neural Netw."},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"2887","DOI":"10.1007\/s11042-020-08836-3","article-title":"Deep learning-based late fusion of multimodal information for emotion classification of music video","volume":"80","author":"Pandeya","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"ref_62","unstructured":"Chua, P., Makris, D., Herremans, D., Roig, G., and Agres, K. (2022). Predicting emotion from music videos: Exploring the relative contribution of visual and auditory information to affective responses. arXiv."},{"key":"ref_63","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1109\/MMUL.2012.26","article-title":"Collecting Large, Richly Annotated Facial-Expression Databases from Movies","volume":"19","author":"Dhall","year":"2012","journal-title":"IEEE MultiMedia"},{"key":"ref_64","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1109\/TAFFC.2015.2396531","article-title":"LIRIS-ACCEDE: A video database for affective content analysis","volume":"6","author":"Baveye","year":"2015","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_65","first-page":"1","article-title":"Harnessing music-related visual stereotypes for music information retrieval","volume":"8","author":"Schindler","year":"2016","journal-title":"ACM Trans. Intell. Syst. Technol. (TIST)"},{"key":"ref_66","first-page":"75","article-title":"Research on emotion space of film and television scene images based on subjective perception","volume":"26","author":"Su","year":"2019","journal-title":"J. China Univ. Posts Telecommun."},{"key":"ref_67","doi-asserted-by":"crossref","first-page":"063014","DOI":"10.1117\/1.JEI.30.6.063014","article-title":"Multidimensional sentiment recognition of film and television scene images","volume":"30","author":"Su","year":"2021","journal-title":"J. Electron. Imaging"},{"key":"ref_68","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_70","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning spatiotemporal features with 3d convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Fan, H., Xiong, B., Mangalam, K., Li, Y., Yan, Z., Malik, J., and Feichtenhofer, C. (2021, January 11\u201317). Multiscale vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00675"},{"key":"ref_72","first-page":"4","article-title":"Is space-time attention all you need for video understanding?","volume":"2","author":"Bertasius","year":"2021","journal-title":"ICML"},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Su, Z., Peng, J., Ren, H., and Zhang, Y. (2022, January 3\u20135). Fine-grained Sentiment Semantic Analysis and Matching of Music and Image. Proceedings of the 2022 IEEE 6th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC), Beijing, China.","DOI":"10.1109\/IAEAC54830.2022.9929967"},{"key":"ref_74","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1016\/0098-3004(93)90090-R","article-title":"Principal components analysis (PCA)","volume":"19","author":"Ratajczak","year":"1993","journal-title":"Comput. Geosci."},{"key":"ref_75","doi-asserted-by":"crossref","first-page":"3744","DOI":"10.1137\/090748330","article-title":"Breaking the curse of dimensionality, or how to use SVD in many dimensions","volume":"31","author":"Oseledets","year":"2009","journal-title":"SIAM J. Sci. Comput."},{"key":"ref_76","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"ref_77","first-page":"52","article-title":"Lightgbm: A highly efficient gradient boosting decision tree","volume":"30","author":"Ke","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_78","doi-asserted-by":"crossref","unstructured":"Chen, T., and Guestrin, C. (2016, January 13\u201317). Xgboost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA.","DOI":"10.1145\/2939672.2939785"},{"key":"ref_79","unstructured":"Zhao, Y. (2020). Research on Stock Price Prediction Model Based on CCA-GA-BPNN Comprehensive Technology. [Master\u2019s Thesis, South China University of Technology]. Available online: https:\/\/d.wanfangdata.com.cn\/thesis\/D02084355."},{"key":"ref_80","first-page":"558","article-title":"Research on the cross-media synesthesia matching of Chinese poetry and folk music based on emotional characteristics","volume":"59","author":"Xing","year":"2020","journal-title":"J. Fudan Univ."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/17\/5681\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:46:21Z","timestamp":1760111181000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/17\/5681"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,31]]},"references-count":80,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2024,9]]}},"alternative-id":["s24175681"],"URL":"https:\/\/doi.org\/10.3390\/s24175681","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2024,8,31]]}}}