{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,24]],"date-time":"2026-01-24T19:30:03Z","timestamp":1769283003513,"version":"3.49.0"},"reference-count":51,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2026,1,22]],"date-time":"2026-01-22T00:00:00Z","timestamp":1769040000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U23A20316"],"award-info":[{"award-number":["U23A20316"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62072346"],"award-info":[{"award-number":["62072346"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Joint Laboratory on Credit Technology"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["BDCC"],"abstract":"<jats:p>With the rapid growth of visual content, automated aesthetic evaluation has become increasingly important. However, existing research faces three key challenges: (1) the absence of datasets combining Image Aesthetic Assessment (IAA) scores and Image Aesthetic Captioning (IAC) descriptions; (2) limited integration of quantitative scores and qualitative text, hindering comprehensive modeling; (3) the subjective nature of aesthetics, which complicates consistent fine-grained evaluation. To tackle these issues, we propose a unified multimodal framework. To address the lack of data, we develop the Textual Aesthetic Sentiment Labeling Pipeline (TASLP) for automatic annotation and construct the Reddit Multimodal Sentiment Dataset (RMSD) with paired IAA and IAC labels. To improve annotation integration, we introduce the Aesthetic Category Sentiment Analysis (ACSA) task, which models fine-grained aesthetic attributes across modalities. To handle subjectivity, we design two models\u2014LAGA for IAA and ACSFM for IAC\u2014that leverage ACSA features to enhance consistency and interpretability. Experiments on RMSD and public benchmarks show that our approach alleviates data limitations and delivers competitive performance, highlighting the effectiveness of fine-grained sentiment modeling and multimodal learning in aesthetic evaluation.<\/jats:p>","DOI":"10.3390\/bdcc10010037","type":"journal-article","created":{"date-parts":[[2026,1,23]],"date-time":"2026-01-23T17:52:38Z","timestamp":1769190758000},"page":"37","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Unifying Aesthetic Evaluation via Multimodal Annotation and Fine-Grained Sentiment Analysis"],"prefix":"10.3390","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-3372-729X","authenticated-orcid":false,"given":"Kai","family":"Liu","sequence":"first","affiliation":[{"name":"School of Computer Science, Wuhan University, Wuhan 430072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0872-434X","authenticated-orcid":false,"given":"Hangyu","family":"Xiong","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Denmark (DTU), 2800 Copenhagen, Denmark"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinyi","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of California, Los Angeles, CA 90095, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Min","family":"Peng","sequence":"additional","affiliation":[{"name":"School of Computer Science, Wuhan University, Wuhan 430072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,1,22]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhang, J., Zhu, Y., Liu, Q., Wu, S., Wang, S., and Wang, L. (2021, January 20\u201324). Mining latent structures for multimedia recommendation. Proceedings of the 29th ACM International Conference on Multimedia, Chengdu, China.","DOI":"10.1145\/3474085.3475259"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"May, A., Chaintreau, A., Korula, N., and Lattanzi, S. (2014, January 16\u201320). Filter & follow: How social media foster content curation. Proceedings of the 2014 ACM International Conference on Measurement and Modeling of Computer Systems, Austin, TX, USA.","DOI":"10.1145\/2591971.2592010"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Yang, Y., Xu, L., Li, L., Qie, N., Li, Y., Zhang, P., and Guo, Y. (2022, January 18\u201324). Personalized Image Aesthetics Assessment with Rich Attributes. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01924"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Murray, N., Marchesotti, L., and Perronnin, F. (2012, January 16\u201321). AVA: A large-scale database for aesthetic visual analysis. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6247954"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Lu, X., Zhang, J., and Wang, J.Z. (2016, January 15\u201319). Joint Image and Text Representation for Aesthetics Analysis. Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands.","DOI":"10.1145\/2964284.2967223"},{"key":"ref_6","unstructured":"Chang, K.Y., Lu, K.H., and Chen, C.S. (2017, January 22\u201329). Aesthetic Critiques Generation for Photos. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"103368","DOI":"10.1016\/j.ipm.2023.103368","article-title":"Novel groundtruth transformations for the aesthetic assessment problem","volume":"60","author":"Flores","year":"2023","journal-title":"Inf. Process. Manag."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Obrador, P., Schmidt-Hackenberg, L., and Oliver, N. (2010, January 26\u201329). The role of image composition in image aesthetics. Proceedings of the 2010 IEEE International Conference on Image Processing, Hong Kong, China.","DOI":"10.1109\/ICIP.2010.5654231"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"103749","DOI":"10.1016\/j.ipm.2024.103749","article-title":"Enhancing image sentiment analysis: A user-centered approach through user emotions and visual features","volume":"61","author":"Liang","year":"2024","journal-title":"Inf. Process. Manag."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"112401","DOI":"10.1016\/j.patcog.2025.112401","article-title":"MIGF-Net: Multimodal interaction-guided fusion network for image aesthetics assessment","volume":"172","author":"Liu","year":"2026","journal-title":"Pattern Recognit."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"129954","DOI":"10.1016\/j.eswa.2025.129954","article-title":"Attribute-guided aesthetic assessment for artistic images based on multimodal hybrid network","volume":"299","author":"Xu","year":"2026","journal-title":"Expert Syst. Appl."},{"key":"ref_12","unstructured":"Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y.J. (2024). LLaVA-NeXT: A Strong Zero-shot Video Understanding Model. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., and Lu, L. (2024, January 16\u201322). Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02283"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Kao, Y., Wang, C., and Huang, K. (2015, January 27\u201330). Visual aesthetic quality assessment with a regression model. Proceedings of the 2015 IEEE International Conference on Image Processing (ICIP), Qu\u00e9bec City, QC, Canada.","DOI":"10.1109\/ICIP.2015.7351067"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"246","DOI":"10.1007\/s11263-014-0789-2","article-title":"Discovering beautiful attributes for aesthetic image analysis","volume":"113","author":"Marchesotti","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"94","DOI":"10.1109\/MSP.2011.941851","article-title":"Aesthetics and Emotions in Images","volume":"28","author":"Joshi","year":"2011","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Jin, X., Wu, L., Zhao, G., Li, X., Zhang, X., Ge, S., Zou, D., Zhou, B., and Zhou, X. (2019, January 21\u201325). Aesthetic Attributes Assessment of Images. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/3343031.3350970"},{"key":"ref_18","unstructured":"Nieto, D.V., Celona, L., and Labrador, C.F. (December, January 28). Understanding Aesthetics with Language: A Photo Critique Dataset for Aesthetic Assessment. Proceedings of the Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, New Orleans, LA, USA."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"2021","DOI":"10.1109\/TMM.2015.2477040","article-title":"Rating image aesthetics using deep learning","volume":"17","author":"Lu","year":"2015","journal-title":"IEEE Trans. Multimed."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Kong, S., Shen, X., Lin, Z., Mech, R., and Fowlkes, C. (2016, January 11\u201314). Photo aesthetics ranking network with attributes and content adaptation. Proceedings of the ECCV, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_40"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhang, B., Niu, L., and Zhang, L. (2021, January 23\u201325). Image composition assessment with saliency-augmented multi-pattern pooling. Proceedings of the BMVC, Virtual.","DOI":"10.5244\/C.35.106"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"110584","DOI":"10.1016\/j.patcog.2024.110584","article-title":"Emotion-aware hierarchical interaction network for multimodal image aesthetics assessment","volume":"154","author":"Zhu","year":"2024","journal-title":"Pattern Recognit."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"110227","DOI":"10.1016\/j.patcog.2023.110227","article-title":"Confidence-based dynamic cross-modal memory network for image aesthetic assessment","volume":"149","author":"Zhang","year":"2024","journal-title":"Pattern Recognit."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"5009","DOI":"10.1109\/TIP.2022.3191853","article-title":"Composition and Style Attributes Guided Image Aesthetic Assessment","volume":"31","author":"Celona","year":"2022","journal-title":"Trans. Img. Proc."},{"key":"ref_25","unstructured":"Li, J., Li, D., Xiong, C., and Hoi, S. (2022, January 17\u201323). BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. Proceedings of the International Conference on Machine Learning, PMLR, Baltimore, MD, USA."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Ke, J., Ye, K., Yu, J., Wu, Y., Milanfar, P., and Yang, F. (2023, January 18\u201322). VILA: Learning Image Aesthetics from User Comments with Vision-Language Pretraining. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00968"},{"key":"ref_27","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_28","unstructured":"Liu, H., Li, C., Wu, Q., and Lee, Y.J. (2023, January 10\u201316). Visual Instruction Tuning. Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_29","unstructured":"Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. (2023). Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv."},{"key":"ref_30","unstructured":"Zhao, Y., Yan, L., Sun, W., Xing, G., Wang, S., Meng, C., Cheng, Z., Ren, Z., and Yin, D. (2024, January 20\u201325). Improving the Robustness of Large Language Models via Consistency Alignment. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC\u2013COLING 2024), Torino, Italy."},{"key":"ref_31","unstructured":"Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., and Bhosale, S. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv."},{"key":"ref_32","unstructured":"Team, G., Anil, R., Borgeaud, S., Alayrac, J.B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A.M., Hauth, A., and Millican, K. (2023). Gemini: A family of highly capable multimodal models. arXiv."},{"key":"ref_33","unstructured":"Team, Q. (2024). Qwen2 technical report. arXiv."},{"key":"ref_34","unstructured":"OpenAI (2023). GPT-4 technical report. arXiv."},{"key":"ref_35","unstructured":"Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"7359","DOI":"10.1007\/s11192-023-04776-5","article-title":"Identifying interdisciplinary topics and their evolution based on BERTopic","volume":"129","author":"Wang","year":"2024","journal-title":"Scientometrics"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Gokcimen, T., and Das, B. (2024, January 29\u201330). Topic modelling using bertopic for robust spam detection. Proceedings of the 2024 12th International Symposium on Digital Forensics and Security (ISDFS), San Antonio, TX, USA.","DOI":"10.1109\/ISDFS60797.2024.10527342"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Ding, X., Zhou, J., Dou, L., Chen, Q., Wu, Y., Chen, A., and He, L. (2024, January 12\u201316). Boosting Large Language Models with Continual Learning for Aspect-based Sentiment Analysis. Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, FL, USA.","DOI":"10.18653\/v1\/2024.findings-emnlp.252"},{"key":"ref_39","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 2\u20137). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA."},{"key":"ref_40","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017, January 4\u20139). Attention is All You Need. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA."},{"key":"ref_41","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Liu, J., Teng, Z., Cui, L., Liu, H., and Zhang, Y. (2021, January 7\u201311). Solving Aspect Category Sentiment Analysis as a Text Generation Task. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominicana.","DOI":"10.18653\/v1\/2021.emnlp-main.361"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Cai, H., Tu, Y., Zhou, X., Yu, J., and Xia, R. (2020, January 13\u201318). Aspect-Category based Sentiment Analysis with Hierarchical Graph Convolutional Network. Proceedings of the 28th International Conference on Computational Linguistics, Barcelona, Spain.","DOI":"10.18653\/v1\/2020.coling-main.72"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Liang, B., Su, H., Yin, R., Gui, L., Yang, M., Zhao, Q., Yu, X., and Xu, R. (2021, January 1\u201311). Beta Distribution Guided Aspect-aware Graph for Aspect Category Sentiment Analysis with Affective Knowledge. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominicana.","DOI":"10.18653\/v1\/2021.emnlp-main.19"},{"key":"ref_45","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very Deep Convolutional Networks for Large-Scale Image Recognition. Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Tong, J., Zhang, G., Kong, P., Rao, Y., Wei, Z., Cui, H., and Guan, Q. (2022). An interpretable approach for automatic aesthetic assessment of remote sensing images. Front. Comput. Neurosci., 16.","DOI":"10.3389\/fncom.2022.1077439"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Ke, J., Wang, Q., Wang, Y., Milanfar, P., and Yang, F. (2021, January 11\u201317). Musiq: Multi-scale image quality transformer. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00510"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Bithel, S., and Bedathur, S. (2023, January 23\u201327). Evaluating Cross-Modal Generative Models Using Retrieval Task. Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, Taipei, Taiwan.","DOI":"10.1145\/3539618.3591979"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft coco: Common objects in context. Proceedings of the Computer Vision\u2013ECCV 2014: 13th European Conference, Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002, January 6\u201312). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_51","unstructured":"Lin, C.Y. (2004). ROUGE: A package for automatic evaluation of summaries. Text Summarization Branches Out: Proceedings of the ACL-04 Workshop, Barcelona, Spain, 25\u201326 July 2004, Association for Computational Linguistics."}],"container-title":["Big Data and Cognitive Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-2289\/10\/1\/37\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,24]],"date-time":"2026-01-24T05:26:40Z","timestamp":1769232400000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-2289\/10\/1\/37"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,22]]},"references-count":51,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["bdcc10010037"],"URL":"https:\/\/doi.org\/10.3390\/bdcc10010037","relation":{},"ISSN":["2504-2289"],"issn-type":[{"value":"2504-2289","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,22]]}}}