{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,25]],"date-time":"2025-06-25T01:36:50Z","timestamp":1750815410446,"version":"3.37.3"},"reference-count":62,"publisher":"Oxford University Press (OUP)","issue":"4","license":[{"start":{"date-parts":[[2020,10,24]],"date-time":"2020-10-24T00:00:00Z","timestamp":1603497600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Nature Science Foundation of China","doi-asserted-by":"crossref","award":["61672267"],"award-info":[{"award-number":["61672267"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Projects of the National Natural Science Foundation of China","award":["U1836220"],"award-info":[{"award-number":["U1836220"]}]},{"name":"Innovation Project of Undergraduate Students in Jiangsu University","award":["201810299103X"],"award-info":[{"award-number":["201810299103X"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,4,19]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Local information has significant contributions to visual sentiment analysis (VSA). Recent studies about local region discovery need manually annotate region location. Affective local information learning and automatic discovery of sentiment-specific region are still the challenges in VSA. In this paper, we propose an end-to-end VSA method for weakly supervised sentiment-specific region discovery. Our method contains two branches: an automatic sentiment-specific region discovery branch and a sentiment analysis branch. In the sentiment-specific region discovery branch, a region proposal network with multiple convolution kernels is proposed to generate candidate affective regions. Then, we design the multiple instance learning (MIL) loss to remove redundant and noisy candidate regions. Finally, the sentiment analysis branch integrates both holistic and localized information obtained in the first branch by feature map coupling for final sentiment classification. Our method automatically discovers sentiment-specific regions by the constraint of MIL loss function without object-level labels. Quantitative and qualitative evaluations on four benchmark affective datasets demonstrate that our proposed method outperforms the state-of-the-art methods.<\/jats:p>","DOI":"10.1093\/comjnl\/bxaa112","type":"journal-article","created":{"date-parts":[[2020,10,21]],"date-time":"2020-10-21T11:13:58Z","timestamp":1603278838000},"page":"818-830","source":"Crossref","is-referenced-by-count":2,"title":["Weakly Supervised Sentiment-Specific Region Discovery for VSA"],"prefix":"10.1093","volume":"65","author":[{"given":"Luoyang","family":"Xue","sequence":"first","affiliation":[{"name":"Computer Science and Communication Engineering, Zhenjaing, Jiangsu Province, 212013, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ang","family":"Xu","sequence":"additional","affiliation":[{"name":"Computer Science and Communication Engineering, Zhenjaing, Jiangsu Province, 212013, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qirong","family":"Mao","sequence":"additional","affiliation":[{"name":"Computer Science and Communication Engineering, Zhenjaing, Jiangsu Province, 212013, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lijian","family":"Gao","sequence":"additional","affiliation":[{"name":"Computer Science and Communication Engineering, Zhenjaing, Jiangsu Province, 212013, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jie","family":"Chen","sequence":"additional","affiliation":[{"name":"Computer Science and Communication Engineering, Zhenjaing, Jiangsu Province, 212013, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2020,10,24]]},"reference":[{"key":"2022041811432180700_ref1","first-page":"223","article-title":"Large-Scale Visual Sentiment Ontology and Detectors Using Adjective Noun Pairs","volume-title":"ACM Multimedia Conf., MM\u201913","author":"Borth","year":"2013"},{"key":"2022041811432180700_ref2","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1016\/j.imavis.2017.01.011","article-title":"From pixels to sentiment: fine-tuning cnns for visual sentiment prediction","volume":"65","author":"Campos","year":"2017","journal-title":"Image Vis. Comput."},{"key":"2022041811432180700_ref3","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1145\/2813524.2813530","article-title":"Diving Deep Into Sentiment: Understanding Fine-Tuned cnns for Visual Sentiment Prediction","volume-title":"The 1st Int. Workshop on Affect & Sentiment in Multimedia, ASM","author":"Campos","year":"2015"},{"key":"2022041811432180700_ref4","first-page":"27:1","article-title":"LIBSVM: a library for support vector machines","volume":"2","author":"Chang","year":"2011","journal-title":"ACM TIST"},{"article-title":"Deepsentibank: visual sentiment concept classification with deep convolutional neural networks","year":"2014","author":"Chen","key":"2022041811432180700_ref5"},{"key":"2022041811432180700_ref6","first-page":"367","article-title":"Object-Based Visual Sentiment Concept Analysis and Application","volume-title":"The ACM Int. Conf. Multimedia, MM\u201914","author":"Chen","year":"2014"},{"key":"2022041811432180700_ref7","doi-asserted-by":"crossref","first-page":"298","DOI":"10.1109\/TAFFC.2014.2388370","article-title":"Assistive image comment robot\u2014a novel mid-level concept-based representation","volume":"6","author":"Chen","year":"2015","journal-title":"IEEE Trans. Affective Comput."},{"key":"2022041811432180700_ref8","first-page":"16","article-title":"Weakly Supervised Learning of Part-Based Spatial Models for Visual Object Recognition","volume-title":"Computer Vision\u2014ECCV, 9th European Conf. Computer Vision","author":"Crandall","year":"2006"},{"key":"2022041811432180700_ref9","first-page":"248","article-title":"Imagenet: A Large-Scale Hierarchical Image Database","volume-title":"2009 IEEE Computer Society Conf. Computer Vision and Pattern Recognition (CVPR)","author":"Deng","year":"2009"},{"key":"2022041811432180700_ref10","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1146\/annurev.ne.18.030195.001205","article-title":"Neural mechanisms of selective visual attention","volume":"18","author":"Desimone","year":"1995","journal-title":"Annu. Rev. Neurosci."},{"key":"2022041811432180700_ref11","first-page":"5957","article-title":"WILDCAT: Weakly Supervised Learning of Deep Convnets for Image Classification, Pointwise Localization and Segmentation","volume-title":"IEEE Conf. Computer Vision and Pattern Recognition, CVPR","author":"Durand","year":"2017"},{"key":"2022041811432180700_ref12","first-page":"1134","article-title":"Object Detection via a Multi-Region and Semantic Segmentation-Aware CNN Model","volume-title":"IEEE Int. Conf. Computer Vision, ICCV","author":"Gidaris","year":"2015"},{"key":"2022041811432180700_ref13","first-page":"580","article-title":"Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation","volume-title":"IEEE Conf. Computer Vision and Pattern Recognition, CVPR","author":"Girshick","year":"2014"},{"key":"2022041811432180700_ref14","first-page":"770","article-title":"Deep Residual Learning for Image Recognition","volume-title":"IEEE Conf. Computer Vision and Pattern Recognition, CVPR","author":"He","year":"2016"},{"key":"2022041811432180700_ref15","doi-asserted-by":"crossref","first-page":"814","DOI":"10.1109\/TPAMI.2015.2465908","article-title":"What makes for effective detection proposals","volume":"38","author":"Hosang","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2022041811432180700_ref16","first-page":"4026","article-title":"Emotion-Aware Human Attention Prediction","volume-title":"IEEE Conf. Computer Vision and Pattern Recognition, CVPR","author":"Macario","year":"2019"},{"key":"2022041811432180700_ref17","first-page":"9841","article-title":"Facial Emotion Distribution Learning by Exploiting Low-Rank Label Correlations Locally","volume-title":"IEEE Conf. Computer Vision and Pattern Recognition, CVPR","author":"Jia","year":"2019"},{"key":"2022041811432180700_ref18","doi-asserted-by":"crossref","first-page":"94","DOI":"10.1109\/MSP.2011.941851","article-title":"Aesthetics and emotions in images","volume":"28","author":"Joshi","year":"2011","journal-title":"IEEE Signal Process. Mag."},{"key":"2022041811432180700_ref19","first-page":"1106","article-title":"Imagenet Classification With Deep Convolutional Neural Networks","volume-title":"The 26th Annual Conf. Neural Information Processing Systems","author":"Krizhevsky","year":"2012"},{"key":"2022041811432180700_ref20","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"Imagenet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2017","journal-title":"Commun. ACM"},{"key":"2022041811432180700_ref21","doi-asserted-by":"crossref","first-page":"541","DOI":"10.1162\/neco.1989.1.4.541","article-title":"Backpropagation applied to handwritten zip code recognition","volume":"1","author":"LeCun","year":"1989","journal-title":"Neural Comput."},{"key":"2022041811432180700_ref22","first-page":"10142","article-title":"Context-Aware Emotion Recognition Networks","volume-title":"IEEE\/CVF International Conf. Computer Vision, ICCV","author":"Lee","year":"2019"},{"key":"2022041811432180700_ref23","first-page":"721","article-title":"Context-Aware Affective Images Classification Based on Bilayer Sparse Representation","volume-title":"The 20th ACM Multimedia Conf., MM\u201912","author":"Li","year":"2012"},{"key":"2022041811432180700_ref24","doi-asserted-by":"crossref","first-page":"1115","DOI":"10.1007\/s11042-016-4310-5","article-title":"Image sentiment prediction based on textual descriptions with adjective noun pairs","volume":"77","author":"Li","year":"2018","journal-title":"Multimedia Tools Appl."},{"key":"2022041811432180700_ref25","first-page":"83","article-title":"Affective Image Classification Using Features Inspired by Psychology and Art Theory","volume-title":"The 18th Int. Conf. Multimedia","author":"Machajdik","year":"2010"},{"key":"2022041811432180700_ref26","doi-asserted-by":"crossref","first-page":"1605","DOI":"10.1093\/comjnl\/bxy016","article-title":"Cascaded multi-level transformed dirichlet process for multi-pose facial expression recognition","volume":"61","author":"Mao","year":"2018","journal-title":"Comput. J."},{"key":"2022041811432180700_ref27","first-page":"933","article-title":"A Multi-Layer Hybrid Framework for Dimensional Emotion Classification","volume-title":"The 19th Int. Conf. Multimedia","author":"Nicolaou","year":"2011"},{"key":"2022041811432180700_ref28","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1016\/j.patrec.2017.07.006","article-title":"Robust discriminative nonnegative dictionary learning for occluded face recognition","volume":"107","author":"Weihua","year":"2018","journal-title":"Pattern Recognit. Lett"},{"key":"2022041811432180700_ref29","first-page":"1","article-title":"Opinion mining and sentiment analysis","volume":"2","author":"Pang","year":"2007","journal-title":"Found. Trends Inf. Retriev."},{"key":"2022041811432180700_ref30","doi-asserted-by":"crossref","first-page":"2008","DOI":"10.1109\/TMM.2015.2482228","article-title":"Deep multimodal learning for affective analysis and retrieval","volume":"17","author":"Pang","year":"2015","journal-title":"IEEE Trans. Multimedia"},{"key":"2022041811432180700_ref31","first-page":"614","article-title":"Where Do Emotions Come From? Predicting the Emotion Stimuli Map","volume-title":"IEEE Int. Conf. Image Processing, ICIP","author":"Peng","year":"2016"},{"key":"2022041811432180700_ref32","doi-asserted-by":"crossref","first-page":"601","DOI":"10.1109\/TPAMI.2011.158","article-title":"Weakly supervised learning of interactions between humans and objects","volume":"34","author":"Prest","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2022041811432180700_ref33","first-page":"1137","article-title":"Faster R-CNN: towards real-time object detection with region proposal networks","volume-title":"IEEE Trans. Pattern Anal. Mach. Intell.","author":"Ren","year":"2017"},{"key":"2022041811432180700_ref34","first-page":"817","article-title":"Grounding of Textual Phrases in Images by Reconstruction","volume-title":"Computer Vision\u2014ECCV\u201414th European Conf.","author":"Rohrbach","year":"2016"},{"key":"2022041811432180700_ref35","first-page":"311","article-title":"Who\u2019s Afraid of Itten: Using the Art Theory of Color Combination to Analyze Emotions in Abstract Paintings","volume-title":"The 23rd Annual ACM Conf. Multimedia Conf., MM\u201915","author":"Sartori","year":"2015"},{"key":"2022041811432180700_ref36","article-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","volume-title":"3rd Int. Conf. Learning Representations, ICLR","author":"Simonyan","year":"2015"},{"key":"2022041811432180700_ref37","doi-asserted-by":"crossref","first-page":"573","DOI":"10.1007\/978-3-642-03767-2_70","article-title":"Color Based Bags-of-Emotions","volume-title":"Computer Analysis of Images and Patterns, 13th Int. Conf., CAIP","author":"Solli","year":"2009"},{"key":"2022041811432180700_ref38","first-page":"1","article-title":"Discovering Affective Regions in Deep Convolutional Neural Networks for Visual Sentiment Prediction","volume-title":"IEEE Int. Conf. Multimedia and Expo, ICME","author":"Sun","year":"2016"},{"key":"2022041811432180700_ref39","first-page":"1","article-title":"Going Deeper With Convolutions","volume-title":"IEEE Conf. Computer Vision and Pattern Recognition, CVPR","author":"Szegedy","year":"2015"},{"key":"2022041811432180700_ref40","first-page":"305","article-title":"Vistanet: Visual Aspect Attention Network for Multimodal Sentiment Analysis","volume-title":"The Thirty-Third AAAI Conf. Artificial Intelligence, AAAI","author":"Truong","year":"2019"},{"key":"2022041811432180700_ref41","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","article-title":"Selective search for object recognition","volume":"104","author":"Uijlings","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"2022041811432180700_ref42","doi-asserted-by":"crossref","first-page":"861","DOI":"10.1093\/comjnl\/bxv068","article-title":"Detecting review spammer groups via bipartite graph projection","volume":"59","author":"Wang","year":"2016","journal-title":"Comput. J."},{"key":"2022041811432180700_ref43","first-page":"7584","article-title":"Weakly Supervised Coupled Networks for Visual Sentiment Analysis","volume-title":"IEEE Conf. Computer Vision and Pattern Recognition, CVPR","author":"Yang","year":"2018"},{"key":"2022041811432180700_ref44","first-page":"3266","article-title":"Joint Image Emotion Classification and Distribution Learning Via Deep Convolutional Neural Network","volume-title":"The Twenty-Sixth Int. Joint Conf. Artificial Intelligence, IJCAI","author":"Yang","year":"2017"},{"key":"2022041811432180700_ref45","first-page":"3266","article-title":"Joint Image Emotion Classification and Distribution Learning via Deep Convolutional Neural Network","volume-title":"The Twenty-Sixth Int. Joint Conf. Artificial Intelligence, IJCAI","author":"Yang","year":"2017"},{"key":"2022041811432180700_ref46","doi-asserted-by":"crossref","first-page":"2513","DOI":"10.1109\/TMM.2018.2803520","article-title":"Visual sentiment prediction based on automatic discovery of affective regions","volume":"20","author":"Yang","year":"2018","journal-title":"IEEE Trans. Multimedia"},{"key":"2022041811432180700_ref47","first-page":"224","article-title":"Learning Visual Sentiment Distributions via Augmented Conditional Probability Neural Network","volume-title":"The Thirty-First AAAI Conf. Artificial Intelligence, AAAI","author":"Yang","year":"2017"},{"key":"2022041811432180700_ref48","first-page":"101","article-title":"Emotional Valence Categorization Using Holistic Image Features","volume-title":"The Int. Conf. Image Processing, ICIP","author":"Yanulevskaya","year":"2008"},{"key":"2022041811432180700_ref49","first-page":"211","article-title":"A 3d Facial Expression Database for Facial Behavior Research","volume-title":"Seventh IEEE Int. Conf. Automatic Face and Gesture Recognition (FGR)","author":"Yin","year":"2006"},{"key":"2022041811432180700_ref50","first-page":"231","article-title":"Visual Sentiment Analysis by Attending on Local Image Regions","volume-title":"The Thirty-First AAAI Conf. Artificial Intelligence, AAAI","author":"You","year":"2017"},{"key":"2022041811432180700_ref51","first-page":"381","article-title":"Robust Image Sentiment Analysis Using Progressively Trained and Domain Transferred Deep Networks","volume-title":"The Twenty-Ninth AAAI Conf. Artificial Intelligence, AAAI","author":"You","year":"2015"},{"key":"2022041811432180700_ref52","first-page":"308","article-title":"Building a Large Scale Dataset for Image Emotion Recognition: The Fine Print and the Benchmark","volume-title":"The Thirtieth AAAI Conf. Artificial Intelligence","author":"You","year":"2016"},{"key":"2022041811432180700_ref53","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1145\/2835776.2835779","article-title":"Cross-Modality Consistent Regression for Joint Visual-Textual Sentiment Analysis of Social Multimedia","volume-title":"The Ninth ACM Int. Conf. Web Search and Data Mining","author":"You","year":"2016"},{"key":"2022041811432180700_ref54","doi-asserted-by":"crossref","first-page":"10:1","DOI":"10.1145\/2502069.2502079","article-title":"Sentribute: Image Sentiment Analysis From a Mid-Level Perspective","volume-title":"The Second Int. Workshop on Issues of Sentiment Discovery and Opinion Mining, WISDOM","author":"Yuan","year":"2013"},{"key":"2022041811432180700_ref55","doi-asserted-by":"crossref","first-page":"2965","DOI":"10.1007\/s00521-017-2900-4","article-title":"Improving sparsity of coefficients for robust sparse and collaborative representation-based image classification","volume":"30","author":"Zeng","year":"2018","journal-title":"Neural Comput. Appl."},{"key":"2022041811432180700_ref56","first-page":"834","article-title":"Part-Based r-cnns for Fine-Grained Category Detection","volume-title":"Computer Vision\u2014ECCV\u201413th European Conf.","author":"Zhang","year":"2014"},{"key":"2022041811432180700_ref57","first-page":"4669","article-title":"Approximating Discrete Probability Distribution of Image Emotions by Multi-Modal Features Fusion","volume-title":"The Twenty-Sixth Int. Joint Conf. Artificial Intelligence, IJCAI","author":"Zhao","year":"2017"},{"key":"2022041811432180700_ref58","doi-asserted-by":"crossref","first-page":"3218","DOI":"10.1109\/TCYB.2017.2762344","article-title":"Real-time multimedia social event detection in microblog","volume":"48","author":"Zhao","year":"2018","journal-title":"IEEE Trans. Cybernetics"},{"key":"2022041811432180700_ref59","first-page":"47","article-title":"Exploring Principles-of-Art Features for Image Emotion Recognition","volume-title":"The ACM Int. Conf. Multimedia, MM\u201914","author":"Zhao","year":"2014"},{"key":"2022041811432180700_ref60","doi-asserted-by":"crossref","first-page":"1224","DOI":"10.1109\/TPAMI.2017.2709749","article-title":"SIFT meets CNN: a decade survey of instance retrieval","volume":"40","author":"Zheng","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"2022041811432180700_ref61","first-page":"2921","article-title":"Learning Deep Features for Discriminative Localization","volume-title":"IEEE Conf. Computer Vision and Pattern Recognition, CVPR","author":"Zhou","year":"2016"},{"key":"2022041811432180700_ref62","first-page":"1859","article-title":"Soft Proposal Networks for Weakly Supervised Object Localization","volume-title":"IEEE Int. Conf. Computer Vision, ICCV","author":"Zhu","year":"2017"}],"container-title":["The Computer Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/65\/4\/818\/43377608\/bxaa112.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/65\/4\/818\/43377608\/bxaa112.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,4,18]],"date-time":"2022-04-18T11:47:15Z","timestamp":1650282435000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/comjnl\/article\/65\/4\/818\/5900572"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,24]]},"references-count":62,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2020,10,24]]},"published-print":{"date-parts":[[2022,4,19]]}},"URL":"https:\/\/doi.org\/10.1093\/comjnl\/bxaa112","relation":{},"ISSN":["0010-4620","1460-2067"],"issn-type":[{"type":"print","value":"0010-4620"},{"type":"electronic","value":"1460-2067"}],"subject":[],"published-other":{"date-parts":[[2022,4]]},"published":{"date-parts":[[2020,10,24]]}}}