{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,25]],"date-time":"2026-01-25T12:07:11Z","timestamp":1769342831852,"version":"3.49.0"},"reference-count":32,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2020,5,13]],"date-time":"2020-05-13T00:00:00Z","timestamp":1589328000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,5,13]],"date-time":"2020-05-13T00:00:00Z","timestamp":1589328000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100004358","name":"Samsung","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100004358","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Data Sci. Eng."],"published-print":{"date-parts":[[2020,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Neural attention mechanism has been used as a form of explanation for model behavior. Users can either passively consume explanation or actively disagree with explanation and then supervise attention into more proper values (attention supervision). Though attention supervision was shown to be effective in some tasks, we find the existing attention supervision is biased, for which we propose to augment counterfactual observations to debias and contribute to accuracy gains. To this end, we propose a counterfactual method to estimate such missing observations and debias the existing supervisions. We validate the effectiveness of our counterfactual supervision on widely adopted image benchmark datasets: CUFED and PEC.<\/jats:p>","DOI":"10.1007\/s41019-020-00119-z","type":"journal-article","created":{"date-parts":[[2020,5,13]],"date-time":"2020-05-13T08:02:43Z","timestamp":1589356963000},"page":"193-204","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Meta-supervision for Attention Using Counterfactual Estimation"],"prefix":"10.1007","volume":"5","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3570-0907","authenticated-orcid":false,"given":"Seungtaek","family":"Choi","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haeju","family":"Park","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Seung-won","family":"Hwang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,5,13]]},"reference":[{"key":"119_CR1","unstructured":"Xu K, Ba J, Kiros R, Cho K, Courville A, Salakhudinov R, Zemel R, Bengio Y (2015) Show, attend and tell: neural image caption generation with visual attention. In: ICML. pp 2048\u20132057"},{"key":"119_CR2","doi-asserted-by":"crossref","unstructured":"Yang Z, Yang D, Dyer C, He X, Smola AJ, Hovy EH (2016) Hierarchical attention networks for document classification. In: HLT-NAACL. pp 1480\u20131489","DOI":"10.18653\/v1\/N16-1174"},{"key":"119_CR3","doi-asserted-by":"crossref","unstructured":"Yu L, Bansal M, Berg T (2017) Hierarchically-attentive RNN for album summarization and storytelling. In: EMNLP. pp 977\u2013982","DOI":"10.18653\/v1\/D17-1101"},{"issue":"7","key":"119_CR4","doi-asserted-by":"publisher","first-page":"1837","DOI":"10.1109\/TMM.2017.2777664","volume":"20","author":"C Guo","year":"2018","unstructured":"Guo C, Tian X, Mei T (2018) Multigranular event recognition of personal photo albums. IEEE Trans Multimed 20(7):1837\u20131847","journal-title":"IEEE Trans Multimed"},{"key":"119_CR5","unstructured":"Liu L, Utiyama M, Finch A, Sumita E (2016) Neural machine translation with supervised attention. In: COLING. pp 3093\u20133102"},{"key":"119_CR6","doi-asserted-by":"crossref","unstructured":"Mi H, Wang Z, Ittycheriah A (2016) Supervised attentions for neural machine translation. In: EMNLP. pp 2283\u20132288","DOI":"10.18653\/v1\/D16-1249"},{"key":"119_CR7","doi-asserted-by":"crossref","unstructured":"Liu C, Mao J, Sha F, Yuille AL (2017) Attention correctness in neural image captioning. In: AAAI. pp 4176\u20134182","DOI":"10.1609\/aaai.v31i1.11197"},{"key":"119_CR8","doi-asserted-by":"crossref","unstructured":"Gan C, Li Y, Li H, Sun C, Gong B (2017) Vqs: Linking segmentations to questions and answers for supervised attention in VQA and question-focused semantic segmentation. In: Proceedings of the IEEE international conference on computer vision. vol 3","DOI":"10.1109\/ICCV.2017.201"},{"key":"119_CR9","unstructured":"Qiao T, Dong J, Xu D (2017) Exploring human-like attention supervision in visual question answering. arXiv preprint arXiv:1709.06308"},{"key":"119_CR10","doi-asserted-by":"crossref","unstructured":"Yu Y, Choi J, Kim Y, Yoo K, Lee S-H, Kim G (2017) Supervising neural attention models for video captioning by human gaze data. In: CVPR. pp 2680\u201329","DOI":"10.1109\/CVPR.2017.648"},{"key":"119_CR11","doi-asserted-by":"crossref","unstructured":"Wang Y, Lin Z, Shen X, Mech R, Miller G, Cottrell GW (2017) Recognizing and curating photo albums via event-specific image importance. In: BMVC","DOI":"10.5244\/C.31.94"},{"key":"119_CR12","doi-asserted-by":"crossref","unstructured":"Wang Y, Lin Z, Shen X, Mech R, Miller G, Cottrell GW (2017) Event-specific image importance. In: CVPR. pp 4810\u20134819","DOI":"10.1109\/CVPR.2016.520"},{"key":"119_CR13","doi-asserted-by":"crossref","unstructured":"Agarwal A, Takatsu K, Zaitsev I, Joachims T (2019) A general framework for counterfactual learning-to-rank. In: ACM conference on research and development in information retrieval (SIGIR)","DOI":"10.1145\/3331184.3331202"},{"issue":"1","key":"119_CR14","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1093\/biomet\/70.1.41","volume":"70","author":"PR Rosenbaum","year":"1983","unstructured":"Rosenbaum PR, Rubin DB (1983) The central role of the propensity score in observational studies for causal effects. Biometrika 70(1):41\u201355","journal-title":"Biometrika"},{"key":"119_CR15","doi-asserted-by":"crossref","unstructured":"Choi S, Park H, Hwang S (2019) Counterfactual attention supervision. In: 2019 IEEE international conference on data mining (ICDM). IEEE, pp 1006\u20131011","DOI":"10.1109\/ICDM.2019.00115"},{"issue":"8","key":"119_CR16","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735\u20131780","journal-title":"Neural Comput"},{"key":"119_CR17","unstructured":"Bahdanau D, Cho K, Bengio Y (2014) Neural machine translation by jointly learning to align and translate. arXiv"},{"key":"119_CR18","doi-asserted-by":"crossref","unstructured":"Jagerman R, Oosterhuis H, de\u00a0Rijke M (2019) To model or to intervene: a comparison of counterfactual and online learning to rank from user interactions","DOI":"10.1145\/3331184.3331269"},{"key":"119_CR19","doi-asserted-by":"crossref","unstructured":"Bossard L, Guillaumin M, Van\u00a0Gool L (2013) Event recognition in photo collections with a stopwatch HMM. In: ICCV","DOI":"10.1109\/ICCV.2013.151"},{"key":"119_CR20","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: CVPR. pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"issue":"11","key":"119_CR21","doi-asserted-by":"publisher","first-page":"1960","DOI":"10.1109\/TMM.2015.2477681","volume":"17","author":"Z Wu","year":"2015","unstructured":"Wu Z, Huang Y, Wang L (2015) Learning representative deep features for image set analysis. IEEE Trans Multimed 17(11):1960\u20131968","journal-title":"IEEE Trans Multimed"},{"key":"119_CR22","unstructured":"Kingma DP, Ba JL (2014) Adam: A method for stochastic optimization. In: Proceedings of the 3rd international conference on learning representations"},{"key":"119_CR23","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1016\/j.cviu.2017.10.001","volume":"163","author":"A Das","year":"2017","unstructured":"Das A, Agrawal H, Zitnick L, Parikh D, Batra D (2017) Human attention in visual question answering: Do humans and deep networks look at the same regions? Comput Vis Image Underst 163:90\u2013100","journal-title":"Comput Vis Image Underst"},{"key":"119_CR24","doi-asserted-by":"crossref","unstructured":"Zhang Y, Niebles JC, Soto A (2018) Interpretable visual question answering by visual grounding from attention supervision mining. arXiv preprint arXiv:1808.00265","DOI":"10.1109\/WACV.2019.00043"},{"key":"119_CR25","doi-asserted-by":"crossref","unstructured":"Wang Z, Liu X, Chen L, Wang L, Qiao Y, Xie X, Fowlkes C (2018) Structured triplet learning with pos-tag guided attention for visual question answering. arXiv preprint arXiv:1801.07853","DOI":"10.1109\/WACV.2018.00209"},{"key":"119_CR26","unstructured":"Jain S, Wallace BC (2019) Attention is not explanation. arXiv preprint"},{"key":"119_CR27","doi-asserted-by":"crossref","unstructured":"Serrano S, Smith NA (2019) Is attention interpretable? arXiv preprint","DOI":"10.18653\/v1\/P19-1282"},{"key":"119_CR28","unstructured":"Zhong R, Shao S, McKeown K (2019) Fine-grained sentiment analysis with faithful attention. arXiv preprint"},{"key":"119_CR29","unstructured":"Zou Y, Gui T, Zhang Q, Huang X (2018) A lexicon-based supervised attention model for neural sentiment analysis. In: Proceedings of the 27th international conference on computational linguistics. pp 868\u2013877"},{"key":"119_CR30","doi-asserted-by":"crossref","unstructured":"Kim J-K, Kim Y-B (2018) Supervised domain enablement attention for personalized domain classification. In: EMNLP","DOI":"10.18653\/v1\/D18-1106"},{"key":"119_CR31","doi-asserted-by":"crossref","unstructured":"Strubell E, Verga P, Andor D, Weiss D, McCallum A (2018) Linguistically-informed self-attention for semantic role labeling. In: Proceedings of the 2018 conference on empirical methods in natural language processing. pp 5027\u20135038","DOI":"10.18653\/v1\/D18-1548"},{"key":"119_CR32","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556"}],"container-title":["Data Science and Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41019-020-00119-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s41019-020-00119-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41019-020-00119-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,23]],"date-time":"2022-10-23T16:32:00Z","timestamp":1666542720000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s41019-020-00119-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,5,13]]},"references-count":32,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2020,6]]}},"alternative-id":["119"],"URL":"https:\/\/doi.org\/10.1007\/s41019-020-00119-z","relation":{},"ISSN":["2364-1185","2364-1541"],"issn-type":[{"value":"2364-1185","type":"print"},{"value":"2364-1541","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,5,13]]},"assertion":[{"value":"16 March 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 April 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}