{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:12:20Z","timestamp":1782400340216,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":36,"publisher":"ACM","license":[{"start":{"date-parts":[[2018,10,15]],"date-time":"2018-10-15T00:00:00Z","timestamp":1539561600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Natural Science Foundation of China","award":["61672520,61573348,61702488,61720106006"],"award-info":[{"award-number":["61672520,61573348,61702488,61720106006"]}]},{"name":"Beijing Natural Science Foundation","award":["4162056"],"award-info":[{"award-number":["4162056"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2018,10,15]]},"DOI":"10.1145\/3240508.3240554","type":"proceedings-article","created":{"date-parts":[[2018,10,18]],"date-time":"2018-10-18T17:52:08Z","timestamp":1539885128000},"page":"879-886","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":104,"title":["Attention-based Multi-Patch Aggregation for Image Aesthetic Assessment"],"prefix":"10.1145","author":[{"given":"Kekai","family":"Sheng","sequence":"first","affiliation":[{"name":"NLPR, Institute of Automation, Chinese Academy of Sciences &amp; University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weiming","family":"Dong","sequence":"additional","affiliation":[{"name":"NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chongyang","family":"Ma","sequence":"additional","affiliation":[{"name":"Snap Inc., Los Angeles, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xing","family":"Mei","sequence":"additional","affiliation":[{"name":"Snap Inc., Los Angeles, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Feiyue","family":"Huang","sequence":"additional","affiliation":[{"name":"Tencent, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bao-Gang","family":"Hu","sequence":"additional","affiliation":[{"name":"NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2018,10,15]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/2037676.2037678"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.338"},{"key":"e_1_3_2_1_3_1","volume-title":"Aesthetic Critiques Generation for Photos. In IEEE International Conference on Computer Vision. IEEE, 3534--3543","author":"Chang Kuang-Yu","year":"2017","unstructured":"Kuang-Yu Chang , Kung-Hung Lu , and Chu-Song Chen . 2017 . Aesthetic Critiques Generation for Photos. In IEEE International Conference on Computer Vision. IEEE, 3534--3543 . Kuang-Yu Chang, Kung-Hung Lu, and Chu-Song Chen. 2017. Aesthetic Critiques Generation for Photos. In IEEE International Conference on Computer Vision. IEEE, 3534--3543."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123274"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/11744078_23"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2017.2696576"},{"key":"e_1_3_2_1_7_1","volume-title":"The Complete Guide to Light and Lighting in Digital Photography (A Lark Photography Book)","author":"Freeman Michael","unstructured":"Michael Freeman . 2006. The Complete Guide to Light and Lighting in Digital Photography (A Lark Photography Book) . Lark Books . Michael Freeman. 2006. The Complete Guide to Light and Lighting in Digital Photography (A Lark Photography Book) .Lark Books."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10578-9_23"},{"key":"e_1_3_2_1_9_1","volume-title":"Deep Residual Learning for Image Recognition","author":"He Kaiming","unstructured":"Kaiming He , Xiangyu Zhang , Shaoqing Ren , and Jian Sun . 2016. Deep Residual Learning for Image Recognition . In IEEE Computer Vision and Pattern Recognition. IEEE , 770--778. Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In IEEE Computer Vision and Pattern Recognition. IEEE, 770--778."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1038\/35058500"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2651399"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.image.2016.05.004"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.303"},{"key":"e_1_3_2_1_14_1","volume-title":"On large-batch training for deep learning: Generalization gap and sharp minima. arXiv preprint:1609.04836","author":"Keskar Nitish Shirish","year":"2016","unstructured":"Nitish Shirish Keskar , Dheevatsa Mudigere , Jorge Nocedal , Mikhail Smelyanskiy , and Ping Tak Peter Tang . 2016. On large-batch training for deep learning: Generalization gap and sharp minima. arXiv preprint:1609.04836 ( 2016 ). Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2016. On large-batch training for deep learning: Generalization gap and sharp minima. arXiv preprint:1609.04836 (2016)."},{"key":"e_1_3_2_1_15_1","volume-title":"Photo Aesthetics Ranking Network with Attributes and Content Adaptation. European Conference on Computer Vision","author":"Kong Shu","year":"2016","unstructured":"Shu Kong , Xiaohui Shen , Zhe Lin , Radomir Mech , and Charless C Fowlkes . 2016 b. Photo Aesthetics Ranking Network with Attributes and Content Adaptation. European Conference on Computer Vision (2016), 662--679. Shu Kong, Xiaohui Shen, Zhe Lin, Radomir Mech, and Charless C Fowlkes. 2016b. Photo Aesthetics Ranking Network with Attributes and Content Adaptation. European Conference on Computer Vision (2016), 662--679."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2016.2515614"},{"key":"e_1_3_2_1_17_1","volume-title":"Focal Loss for Dense Object Detection. In IEEE International Conference on Computer Vision. IEEE, 2999--3007","author":"Lin Tsung Yi","year":"2017","unstructured":"Tsung Yi Lin , Priya Goyal , Ross Girshick , Kaiming He , and Piotr Dollar . 2017 . Focal Loss for Dense Object Detection. In IEEE International Conference on Computer Vision. IEEE, 2999--3007 . Tsung Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. 2017. Focal Loss for Dense Object Detection. In IEEE International Conference on Computer Vision. IEEE, 2999--3007."},{"key":"e_1_3_2_1_18_1","volume-title":"Fully convolutional networks for semantic segmentation","author":"Long Jonathan","unstructured":"Jonathan Long , Evan Shelhamer , and Trevor Darrell . 2015. Fully convolutional networks for semantic segmentation . In IEEE Computer Vision and Pattern Recognition. IEEE , 3431--3440. Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015. Fully convolutional networks for semantic segmentation. In IEEE Computer Vision and Pattern Recognition. IEEE, 3431--3440."},{"key":"e_1_3_2_1_19_1","volume-title":"Composition-preserving deep photo aesthetics assessment","author":"Long Mai","unstructured":"Mai Long , Jin Hailin , and Liu Feng . 2016. Composition-preserving deep photo aesthetics assessment . In IEEE Computer Vision and Pattern Recognition. IEEE , 497--506. Mai Long, Jin Hailin, and Liu Feng. 2016. Composition-preserving deep photo aesthetics assessment. In IEEE Computer Vision and Pattern Recognition. IEEE, 497--506."},{"key":"e_1_3_2_1_20_1","volume-title":"Online Batch Selection for Faster Training of Neural Networks. Mathematics","author":"Loshchilov Ilya","year":"2015","unstructured":"Ilya Loshchilov and Frank Hutter . 2015. Online Batch Selection for Faster Training of Neural Networks. Mathematics ( 2015 ). Ilya Loshchilov and Frank Hutter. 2015. Online Batch Selection for Faster Training of Neural Networks. Mathematics (2015)."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2015.2477040"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.119"},{"key":"e_1_3_2_1_23_1","volume-title":"A-Lamp: Adaptive layout-aware multi-patch deep convolutional neural network for photo aesthetic assessment","author":"Ma Shuang","unstructured":"Shuang Ma , Jing Liu , and Wen Chen Chang . 2017. A-Lamp: Adaptive layout-aware multi-patch deep convolutional neural network for photo aesthetic assessment . In IEEE Computer Vision and Pattern Recognition. IEEE , 722--731. Shuang Ma, Jing Liu, and Wen Chen Chang. 2017. A-Lamp: Adaptive layout-aware multi-patch deep convolutional neural network for photo aesthetic assessment. In IEEE Computer Vision and Pattern Recognition. IEEE, 722--731."},{"key":"e_1_3_2_1_24_1","volume-title":"SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability. In Advances in Neural Information Processing Systems. 6076--6085.","author":"Maithra Raghu","year":"2017","unstructured":"Raghu Maithra , Gilmer Justin , Yosinski Jason , and Jascha Sohl-Dickstein . 2017 . SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability. In Advances in Neural Information Processing Systems. 6076--6085. Raghu Maithra, Gilmer Justin, Yosinski Jason, and Jascha Sohl-Dickstein. 2017. SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability. In Advances in Neural Information Processing Systems. 6076--6085."},{"key":"e_1_3_2_1_25_1","volume-title":"AVA: A large-scale database for aesthetic visual analysis","author":"Murray Naila","year":"2012","unstructured":"Naila Murray , Luca Marchesotti , and Florent Perronnin . 2012 . AVA: A large-scale database for aesthetic visual analysis . In IEEE Computer Vision and Pattern Recognition. IEEE , 2408--2415. Naila Murray, Luca Marchesotti, and Florent Perronnin. 2012. AVA: A large-scale database for aesthetic visual analysis. In IEEE Computer Vision and Pattern Recognition. IEEE, 2408--2415."},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"crossref","unstructured":"Clark V. Poling. 1975. Johannes Itten Design and Form: The Basic Course at the Bauhaus and Later .Thames and Hudson. 368--370 pages.  Clark V. Poling. 1975. Johannes Itten Design and Form: The Basic Course at the Bauhaus and Later .Thames and Hudson. 368--370 pages.","DOI":"10.1080\/00043249.1977.10793387"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_1_28_1","volume-title":"Training Region-Based Object Detectors with Online Hard Example Mining","author":"Shrivastava Abhinav","unstructured":"Abhinav Shrivastava , Abhinav Gupta , and Ross Girshick . 2016. Training Region-Based Object Detectors with Online Hard Example Mining . In IEEE Computer Vision and Pattern Recognition. IEEE , 761--769. Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick. 2016. Training Region-Based Object Detectors with Online Hard Example Mining. In IEEE Computer Vision and Pattern Recognition. IEEE, 761--769."},{"key":"e_1_3_2_1_29_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint:1409.1556 (2014)."},{"key":"e_1_3_2_1_30_1","unstructured":"Marijn F Stollenga Jonathan Masci Faustino Gomez and J\u00fcrgen Schmidhuber. 2014. Deep networks with internal selective attention through feedback connections. In Advances in neural information processing systems. 3545--3553.   Marijn F Stollenga Jonathan Masci Faustino Gomez and J\u00fcrgen Schmidhuber. 2014. Deep networks with internal selective attention through feedback connections. In Advances in neural information processing systems. 3545--3553."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2831899"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-30541-5_25"},{"key":"e_1_3_2_1_33_1","volume-title":"Residual attention network for image classification","author":"Wang Fei","unstructured":"Fei Wang , Mengqing Jiang , Chen Qian , Shuo Yang , Cheng Li , Honggang Zhang , Xiaogang Wang , and Xiaoou Tang . 2017. Residual attention network for image classification . IEEE Computer Vision and Pattern Recognition , 3156--3164. Fei Wang, Mengqing Jiang, Chen Qian, Shuo Yang, Cheng Li, Honggang Zhang, Xiaogang Wang, and Xiaoou Tang. 2017. Residual attention network for image classification. IEEE Computer Vision and Pattern Recognition, 3156--3164."},{"key":"e_1_3_2_1_34_1","volume-title":"Brain-inspired deep networks for image aesthetics assessment. arXiv preprint:1601.04155","author":"Wang Zhangyang","year":"2016","unstructured":"Zhangyang Wang , Shiyu Chang , Florin Dolcos , Diane Beck , Ding Liu , and Thomas S Huang . 2016. Brain-inspired deep networks for image aesthetics assessment. arXiv preprint:1601.04155 ( 2016 ). Zhangyang Wang, Shiyu Chang, Florin Dolcos, Diane Beck, Ding Liu, and Thomas S Huang. 2016. Brain-inspired deep networks for image aesthetics assessment. arXiv preprint:1601.04155 (2016)."},{"key":"e_1_3_2_1_35_1","volume-title":"IEEE International Conference on Computer Vision. IEEE, 2186--2194","author":"Wenguan Wang","year":"2017","unstructured":"Wang Wenguan and Shen Jianbing . 2017 . Deep cropping via attention box prediction and aesthetic assessment . In IEEE International Conference on Computer Vision. IEEE, 2186--2194 . Wang Wenguan and Shen Jianbing. 2017. Deep cropping via attention box prediction and aesthetic assessment. In IEEE International Conference on Computer Vision. IEEE, 2186--2194."},{"key":"e_1_3_2_1_36_1","volume-title":"Describing Human Aesthetic Perception by Deeply-learned Attributes from Flickr. arXiv preprint:1605.07699","author":"Zhang Luming","year":"2016","unstructured":"Luming Zhang . 2016. Describing Human Aesthetic Perception by Deeply-learned Attributes from Flickr. arXiv preprint:1605.07699 ( 2016 ). Luming Zhang. 2016. Describing Human Aesthetic Perception by Deeply-learned Attributes from Flickr. arXiv preprint:1605.07699 (2016)."}],"event":{"name":"MM '18: ACM Multimedia Conference","location":"Seoul Republic of Korea","acronym":"MM '18","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 26th ACM international conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3240508.3240554","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3240508.3240554","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:44:01Z","timestamp":1750207441000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3240508.3240554"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,10,15]]},"references-count":36,"alternative-id":["10.1145\/3240508.3240554","10.1145\/3240508"],"URL":"https:\/\/doi.org\/10.1145\/3240508.3240554","relation":{},"subject":[],"published":{"date-parts":[[2018,10,15]]},"assertion":[{"value":"2018-10-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}