{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T16:52:49Z","timestamp":1780764769881,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":35,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key R&D Program","award":["No.2017YFC0113000; No.2016YFB1001503"],"award-info":[{"award-number":["No.2017YFC0113000; No.2016YFB1001503"]}]},{"name":"Key R&D Program of Jiangxi Province","award":["No. 20171ACH80022"],"award-info":[{"award-number":["No. 20171ACH80022"]}]},{"name":"National Natural Science Foundation of China","award":["No.U1705262; No.61772443; No.61572410; No.61802324; No.61702136"],"award-info":[{"award-number":["No.U1705262; No.61772443; No.61572410; No.61802324; No.61702136"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3414008","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T12:26:53Z","timestamp":1602505613000},"page":"4199-4207","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Exploring Language Prior for Mode-Sensitive Visual Attention Modeling"],"prefix":"10.1145","author":[{"given":"Xiaoshuai","family":"Sun","sequence":"first","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xuying","family":"Zhang","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liujuan","family":"Cao","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yongjian","family":"Wu","sequence":"additional","affiliation":[{"name":"Youtu Lab, Tencent, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Feiyue","family":"Huang","sequence":"additional","affiliation":[{"name":"Youtu Lab, Tencent, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rongrong","family":"Ji","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"crossref","unstructured":"Peter Anderson Xiaodong He Chris Buehler Damien Teney Mark Johnson Stephen Gould and Lei Zhang. 2018. Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In CVPR. 6077--6086.  Peter Anderson Xiaodong He Chris Buehler Damien Teney Mark Johnson Stephen Gould and Lei Zhang. 2018. Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In CVPR. 6077--6086.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_2_2_2_1","volume-title":"Saliency Prediction in the Deep Learning Era: Successes and Limitations","author":"Borji Ali","year":"2019","unstructured":"Ali Borji . 2019. Saliency Prediction in the Deep Learning Era: Successes and Limitations . IEEE Transactions on Pattern Analysis and Machine Intelligence ( 2019 ), 1--1. https:\/\/doi.org\/10.1109\/tpami.2019.2935715 10.1109\/tpami.2019.2935715 Ali Borji. 2019. Saliency Prediction in the Deep Learning Era: Successes and Limitations. IEEE Transactions on Pattern Analysis and Machine Intelligence (2019), 1--1. https:\/\/doi.org\/10.1109\/tpami.2019.2935715"},{"key":"e_1_3_2_2_3_1","first-page":"185","article-title":"State-of-the-art in visual attention modeling. Pattern Analysis and Machine Intelligence","volume":"35","author":"Borji Ali","year":"2013","unstructured":"Ali Borji and Laurent Itti . 2013 . State-of-the-art in visual attention modeling. Pattern Analysis and Machine Intelligence , IEEE Transactions on , Vol. 35 , 1 (2013), 185 -- 207 . Ali Borji and Laurent Itti. 2013. State-of-the-art in visual attention modeling. Pattern Analysis and Machine Intelligence, IEEE Transactions on, Vol. 35, 1 (2013), 185--207.","journal-title":"IEEE Transactions on"},{"key":"e_1_3_2_2_4_1","unstructured":"N. Bruce and J. Tsotsos. 2006. Saliency Based on Information Maximization. In Advances in Neural Information Processing Systems (NIPS). 155--162.  N. Bruce and J. Tsotsos. 2006. Saliency Based on Information Maximization. In Advances in Neural Information Processing Systems (NIPS). 155--162."},{"key":"e_1_3_2_2_5_1","unstructured":"Zoya Bylinskii Tilke Judd Ali Borji Laurent Itti Fr\u00e9do Durand Aude Oliva and Antonio Torralba. [n. d.]. MIT Saliency Benchmark.  Zoya Bylinskii Tilke Judd Ali Borji Laurent Itti Fr\u00e9do Durand Aude Oliva and Antonio Torralba. [n. d.]. MIT Saliency Benchmark."},{"key":"e_1_3_2_2_6_1","volume-title":"What do different evaluation metrics tell us about saliency models? arXiv preprint arXiv:1604.03605","author":"Bylinskii Zoya","year":"2016","unstructured":"Zoya Bylinskii , Tilke Judd , Aude Oliva , Antonio Torralba , and Fr\u00e9do Durand . 2016. What do different evaluation metrics tell us about saliency models? arXiv preprint arXiv:1604.03605 ( 2016 ). Zoya Bylinskii, Tilke Judd, Aude Oliva, Antonio Torralba, and Fr\u00e9do Durand. 2016. What do different evaluation metrics tell us about saliency models? arXiv preprint arXiv:1604.03605 (2016)."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01252-6_5"},{"key":"e_1_3_2_2_8_1","volume-title":"Human Attention in Image Captioning: Dataset and Analysis. In The IEEE International Conference on Computer Vision (ICCV).","author":"He Sen","year":"2019","unstructured":"Sen He , Hamed R. Tavakoli , Ali Borji , and Nicolas Pugeault . 2019 . Human Attention in Image Captioning: Dataset and Analysis. In The IEEE International Conference on Computer Vision (ICCV). Sen He, Hamed R. Tavakoli, Ali Borji, and Nicolas Pugeault. 2019. Human Attention in Image Captioning: Dataset and Analysis. In The IEEE International Conference on Computer Vision (ICCV)."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/2919332.2919806"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1038\/35058500"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298710"},{"key":"e_1_3_2_2_13_1","volume-title":"Computer Vision, 2009 IEEE 12th International Conference on. IEEE, 2106--2113","author":"Judd T.","unstructured":"T. Judd , K. Ehinger , F. Durand , and A. Torralba . 2009. Learning to predict where humans look . In Computer Vision, 2009 IEEE 12th International Conference on. IEEE, 2106--2113 . T. Judd, K. Ehinger, F. Durand, and A. Torralba. 2009. Learning to predict where humans look. In Computer Vision, 2009 IEEE 12th International Conference on. IEEE, 2106--2113."},{"key":"e_1_3_2_2_14_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba . 2014 . Adam : A Method for Stochastic Optimization. CoRR , Vol. abs\/ 1412 .6980 (2014). Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Optimization. CoRR, Vol. abs\/1412.6980 (2014)."},{"key":"e_1_3_2_2_15_1","volume-title":"Deepfix: A fully convolutional neural network for predicting human eye fixations. arXiv preprint arXiv:1510.02927","author":"Kruthiventi Srinivas SS","year":"2015","unstructured":"Srinivas SS Kruthiventi , Kumar Ayush , and R Venkatesh Babu . 2015 . Deepfix: A fully convolutional neural network for predicting human eye fixations. arXiv preprint arXiv:1510.02927 (2015). Srinivas SS Kruthiventi, Kumar Ayush, and R Venkatesh Babu. 2015. Deepfix: A fully convolutional neural network for predicting human eye fixations. arXiv preprint arXiv:1510.02927 (2015)."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"crossref","unstructured":"Matthias K\u00fcmmerer Thomas S. A. Wallis and Matthias Bethge. 2018. Saliency Benchmarking Made Easy: Separating Models Maps and Metrics. In ECCV.  Matthias K\u00fcmmerer Thomas S. A. Wallis and Matthias Bethge. 2018. Saliency Benchmarking Made Easy: Separating Models Maps and Metrics. In ECCV.","DOI":"10.1007\/978-3-030-01270-0_47"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.513"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1510393112"},{"key":"e_1_3_2_2_19_1","volume-title":"Dynamic whitening saliency","author":"Lebor\u00e1n V'ictor","year":"2017","unstructured":"V'ictor Lebor\u00e1n , Anton Garcia-Diaz , Xos\u00e9 R Fdez-Vidal , and Xos\u00e9 M Pardo . 2017. Dynamic whitening saliency . IEEE transactions on pattern analysis and machine intelligence, Vol. 39 , 5 ( 2017 ), 893--907. V'ictor Lebor\u00e1n, Anton Garcia-Diaz, Xos\u00e9 R Fdez-Vidal, and Xos\u00e9 M Pardo. 2017. Dynamic whitening saliency. IEEE transactions on pattern analysis and machine intelligence, Vol. 39, 5 (2017), 893--907."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.43"},{"key":"e_1_3_2_2_21_1","volume-title":"Microsoft COCO: Common Objects in Context. In European Conference on Computer Vision. 740--755","author":"Lin Tsungyi","year":"2014","unstructured":"Tsungyi Lin , Michael Maire , Serge J Belongie , James Hays , Pietro Perona , Deva Ramanan , Piotr Dollar , and C Lawrence Zitnick . 2014 . Microsoft COCO: Common Objects in Context. In European Conference on Computer Vision. 740--755 . Tsungyi Lin, Michael Maire, Serge J Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. In European Conference on Computer Vision. 740--755."},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.71"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.visres.2005.03.019"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_2_2_25_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.272"},{"key":"e_1_3_2_2_28_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In NeurIPS. 5998--6008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In NeurIPS. 5998--6008."},{"key":"e_1_3_2_2_29_1","volume-title":"CIDEr: Consensus-Based Image Description Evaluation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Vedantam Ramakrishna","year":"2015","unstructured":"Ramakrishna Vedantam , C. Lawrence Zitnick , and Devi Parikh . 2015 . CIDEr: Consensus-Based Image Description Evaluation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. CIDEr: Consensus-Based Image Description Evaluation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_2_30_1","volume-title":"Show and Tell: A Neural Image Caption Generator. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Vinyals Oriol","year":"2015","unstructured":"Oriol Vinyals , Alexander Toshev , Samy Bengio , and Dumitru Erhan . 2015 . Show and Tell: A Neural Image Caption Generator. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015. Show and Tell: A Neural Image Caption Generator. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"crossref","unstructured":"Wenguan Wang Jianbing Shen Ming-Ming Cheng and L. M. Shao. 2019. An Iterative and Cooperative Top-Down and Bottom-Up Inference Network for Salient Object Detection. In CVPR.  Wenguan Wang Jianbing Shen Ming-Ming Cheng and L. M. Shao. 2019. An Iterative and Cooperative Top-Down and Bottom-Up Inference Network for Salient Object Detection. In CVPR.","DOI":"10.1109\/CVPR.2019.00612"},{"key":"e_1_3_2_2_32_1","unstructured":"Kelvin Xu Jimmy Ba Ryan Kiros Kyunghyun Cho Aaron C Courville Ruslan Salakhudinov Rich Zemel and Yoshua Bengio. 2015. Show Attend and Tell: Neural Image Caption Generation with Visual Attention. In ICML. 2048--2057.  Kelvin Xu Jimmy Ba Ryan Kiros Kyunghyun Cho Aaron C Courville Ruslan Salakhudinov Rich Zemel and Yoshua Bengio. 2015. Show Attend and Tell: Neural Image Caption Generation with Visual Attention. In ICML. 2048--2057."},{"key":"e_1_3_2_2_33_1","unstructured":"Quanzeng You Hailin Jin Zhaowen Wang Chen Fang and Jiebo Luo. 2016. Image Captioning with Semantic Attention. In CVPR. 4651--4659.  Quanzeng You Hailin Jin Zhaowen Wang Chen Fang and Jiebo Luo. 2016. Image Captioning with Semantic Attention. In CVPR. 4651--4659."},{"key":"e_1_3_2_2_34_1","volume-title":"Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Computer Society Conference on. IEEE.","author":"Yun Kiwon","unstructured":"Kiwon Yun , Yifan Peng , Dimitris Samaras , Gregory J. Zelinsky , and Tamara L. Berg . 2013. Studying Relationships Between Human Gaze, Description, and Computer Vision . In Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Computer Society Conference on. IEEE. Kiwon Yun, Yifan Peng, Dimitris Samaras, Gregory J. Zelinsky, and Tamara L. Berg. 2013. Studying Relationships Between Human Gaze, Description, and Computer Vision. In Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Computer Society Conference on. IEEE."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"crossref","unstructured":"Jianming Zhang and Stan Sclaroff. 2013. Saliency detection: a boolean map approach. In ICCV.  Jianming Zhang and Stan Sclaroff. 2013. Saliency detection: a boolean map approach. In ICCV.","DOI":"10.1109\/ICCV.2013.26"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1167\/8.7.32"}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","location":"Seattle WA USA","acronym":"MM '20","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3414008","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3414008","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:32:07Z","timestamp":1750195927000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3414008"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":35,"alternative-id":["10.1145\/3394171.3414008","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3414008","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}