{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T23:07:15Z","timestamp":1777676835332,"version":"3.51.4"},"reference-count":64,"publisher":"Association for Computing Machinery (ACM)","issue":"CSCW2","license":[{"start":{"date-parts":[[2023,9,28]],"date-time":"2023-09-28T00:00:00Z","timestamp":1695859200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000050","name":"National Heart, Lung, and Blood Institute","doi-asserted-by":"publisher","award":["R01HL164906"],"award-info":[{"award-number":["R01HL164906"]}],"id":[{"id":"10.13039\/100000050","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Hum.-Comput. Interact."],"published-print":{"date-parts":[[2023,9,28]]},"abstract":"<jats:p>Recent developments in explainable AI (XAI) aim to improve the transparency of black-box models. However, empirically evaluating the interpretability of these XAI techniques is still an open challenge. The most common evaluation method is algorithmic performance, but such an approach may not accurately represent how interpretable these techniques are to people. A less common but growing evaluation strategy is to leverage crowd-workers to provide feedback on multiple XAI techniques to compare them. However, these tasks often feel like work and may limit participation. We propose a novel, playful, human-centered method for evaluating XAI techniques: a Game With a Purpose (GWAP), Eye into AI, that allows researchers to collect human evaluations of XAI at scale. We provide an empirical study demonstrating how our GWAP supports evaluating and comparing the agreement between three popular XAI techniques (LIME, Grad-CAM, and Feature Visualization) and humans, as well as evaluating and comparing the interpretability of those three XAI techniques applied to a deep learning model for image classification. The data collected from Eye into AI offers convincing evidence that GWAPs can be used to evaluate and compare XAI techniques. Eye into AI is available to the public: https:\/\/dig.cmu.edu\/eyeintoai\/.<\/jats:p>","DOI":"10.1145\/3610064","type":"journal-article","created":{"date-parts":[[2023,10,4]],"date-time":"2023-10-04T15:54:10Z","timestamp":1696434850000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Eye into AI: Evaluating the Interpretability of Explainable AI Techniques through a Game with a Purpose"],"prefix":"10.1145","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2644-4422","authenticated-orcid":false,"given":"Katelyn","family":"Morrison","sequence":"first","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-0250-2022","authenticated-orcid":false,"given":"Mayank","family":"Jain","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4102-2575","authenticated-orcid":false,"given":"Jessica","family":"Hammer","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8369-3847","authenticated-orcid":false,"given":"Adam","family":"Perer","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,10,4]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173574.3174156"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/access.2018.2870052"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327546.3327621"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377325.3377519"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3485447.3512241"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445717"},{"key":"e_1_2_1_7_1","volume-title":"https:\/\/distill.pub\/2019\/activation-atlas","author":"Carter Shan","year":"2019","unstructured":"Shan Carter, Zan Armstrong, Ludwig Schubert, Ian Johnson, and Chris Olah. 2019. Activation Atlas. Distill (2019). https:\/\/distill.pub\/2019\/activation-atlas"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cub.2004.09.041"},{"key":"e_1_2_1_9_1","unstructured":"Sabrina Culyba. 2018. The Transformational Framework: A process tool for the development of Transformational games. figshare. https:\/\/kilthub.cmu.edu\/articles\/journal_contribution\/The_Transformational_Framework_A_Process_Tool_for_the_Development_of_Transformational_Games\/7130594\/files\/13117568.pdf"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1177\/1075547015609322"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_12_1","unstructured":"Finale Doshi-Velez and Been Kim. 2017. Towards A Rigorous Science of Interpretable Machine Learning. arxiv: 1702.08608 [stat.ML] https:\/\/arxiv.org\/pdf\/1702.08608"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445188"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-60117-1_33"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411763.3441342"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3334480.3382831"},{"key":"e_1_2_1_17_1","unstructured":"Jacob Gildenblat and contributors. 2021. PyTorch library for CAM methods. https:\/\/github.com\/jacobgil\/pytorch-grad-cam."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3235765.3235803"},{"key":"e_1_2_1_19_1","first-page":"1","article-title":"KissKissBan","volume":"12","author":"Ho Chien-Ju","year":"2010","unstructured":"Chien-Ju Ho, Tao-Hsuan Chang, Jong-Chuan Lee, Jane Yung-jen Hsu, and Kuan-Ta Chen. 2010. KissKissBan: A Competitive Human Computation Game for Image Annotation. SIGKDD Explor. Newsl., Vol. 12, 1 (Nov. 2010), 21--24. http:\/\/doi.acm.org\/10.1145\/1882471.1882475","journal-title":"A Competitive Human Computation Game for Image Annotation. SIGKDD Explor. Newsl."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/tvcg.2018.2843369"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2858036.2858558"},{"key":"e_1_2_1_22_1","volume-title":"Examples are not enough, learn to criticize! criticism for interpretability. Advances in neural information processing systems","author":"Kim Been","year":"2016","unstructured":"Been Kim, Rajiv Khanna, and Oluwasanmi O Koyejo. 2016. Examples are not enough, learn to criticize! criticism for interpretability. Advances in neural information processing systems, Vol. 29 (2016). https:\/\/proceedings.neurips.cc\/paper\/2016\/file\/5680522b8e2bb01943234bce7bf84534-Paper.pdf"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--3-031--19775--8_17"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2858036.2858529"},{"key":"e_1_2_1_25_1","volume-title":"ISMIR","volume":"3","author":"Law Edith LM","year":"2007","unstructured":"Edith LM Law, Luis Von Ahn, Roger B Dannenberg, and Mike Crawford. 2007. TagATune: A Game for Music and Sound Annotation.. In ISMIR, Vol. 3. 2. https:\/\/www.cs.cmu.edu\/ elaw\/papers\/ISMIR2007.pdf"},{"key":"e_1_2_1_26_1","volume-title":"Human-centered explainable AI (XAI): From algorithms to user experiences. arXiv preprint arXiv:2110.10790","author":"Vera Liao Q","year":"2021","unstructured":"Q Vera Liao and Kush R Varshney. 2021. Human-centered explainable AI (XAI): From algorithms to user experiences. arXiv preprint arXiv:2110.10790 (2021). https:\/\/arxiv.org\/pdf\/2110.10790"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1177\/1555412014559978"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467213"},{"key":"e_1_2_1_29_1","volume-title":"Stanislav Bochkarev, Michael St Jules, Xiao Yu Wang, and Alexander Wong.","author":"Lin Zhong Qiu","year":"2019","unstructured":"Zhong Qiu Lin, Mohammad Javad Shafiee, Stanislav Bochkarev, Michael St Jules, Xiao Yu Wang, and Alexander Wong. 2019. Do explanations reflect decisions? A machine-centric strategy to quantify the performance of explainability algorithms. arXiv preprint arXiv:1910.07387 (2019). https:\/\/arxiv.org\/pdf\/1910.07387.pdf"},{"key":"e_1_2_1_30_1","volume-title":"The Mythos of Model Interpretability. CoRR","author":"Lipton Zachary Chase","year":"2016","unstructured":"Zachary Chase Lipton. 2016. The Mythos of Model Interpretability. CoRR, Vol. abs\/1606.03490 (2016). arxiv: 1606.03490 http:\/\/arxiv.org\/abs\/1606.03490"},{"key":"e_1_2_1_31_1","volume-title":"Crowdsourcing Evaluation of Saliency-based XAI Methods. arXiv:2107.00456 [cs] (Aug","author":"Lu Xiaotian","year":"2021","unstructured":"Xiaotian Lu, Arseny Tolmachev, Tatsuya Yamamoto, Koh Takeuchi, Seiji Okajima, Tomoyoshi Takebayashi, Koji Maruhashi, and Hisashi Kashima. 2021. Crowdsourcing Evaluation of Saliency-based XAI Methods. arXiv:2107.00456 [cs] (Aug 2021). http:\/\/arxiv.org\/abs\/2107.00456 arXiv: 2107.00456."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3130859.3131332"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2018.07.007"},{"key":"e_1_2_1_34_1","volume-title":"Ragan","author":"Mohseni Sina","year":"2018","unstructured":"Sina Mohseni and Eric D. Ragan. 2018. A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning. CoRR, Vol. abs\/1801.05075 (2018). showeprint[arXiv]1801.05075 http:\/\/arxiv.org\/abs\/1801.05075"},{"key":"e_1_2_1_35_1","unstructured":"Christoph Molnar. 2019. Model-Agnostic Methods. Lulu. https:\/\/christophm.github.io\/interpretable-ml-book\/agnostic.html"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1609\/hcomp.v7i1.5284"},{"key":"e_1_2_1_37_1","volume-title":"Feature Visualization. Distill","author":"Olah Chris","year":"2017","unstructured":"Chris Olah, Alexander Mordvintsev, and Ludwig Schubert. 2017. Feature Visualization. Distill (2017). https:\/\/distill.pub\/2017\/feature-visualization"},{"key":"e_1_2_1_38_1","volume-title":"The Building Blocks of Interpretability. Distill","author":"Olah Chris","year":"2018","unstructured":"Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev. 2018. The Building Blocks of Interpretability. Distill (2018). https:\/\/distill.pub\/2018\/building-blocks"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/1866029.1866038"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1080\/0144929x.2013.862304"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1002\/asi.23863"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1978942.1979148"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939778"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939778"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/iccvw54120.2021.00462"},{"key":"e_1_2_1_47_1","volume-title":"Digra conference. [PDF] digra.org","author":"Henrik","unstructured":"Henrik Schoenau-Fog et al. 2011. The Player Engagement Process-An Exploration of Continuation Desire in Digital Games.. In Digra conference. [PDF] digra.org"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/1837885.1837905"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-019-01228--7"},{"key":"e_1_2_1_50_1","volume-title":"Proceedings of the 12th International Conference on the Foundations of Digital Games","author":"Siu Kristin","year":"2071","unstructured":"Kristin Siu, Matthew Guzdial, and Mark O. Riedl. 2017a. Evaluating Singleplayer and Multiplayer in Human Computation Games. In Proceedings of the 12th International Conference on the Foundations of Digital Games (Hyannis, Massachusetts) (FDG '17). ACM, New York, NY, USA, Article 34, 10 pages. http:\/\/doi.acm.org\/10.1145\/3102071.3102077"},{"key":"e_1_2_1_51_1","volume-title":"Proceedings of the 2016 Annual Symposium on Computer-Human Interaction in Play (Austin, Texas, USA) (CHI PLAY '16). ACM","author":"Siu Kristin","unstructured":"Kristin Siu and Mark O. Riedl. 2016. Reward Systems in Human Computation Games. In Proceedings of the 2016 Annual Symposium on Computer-Human Interaction in Play (Austin, Texas, USA) (CHI PLAY '16). ACM, New York, NY, USA, 266--275. http:\/\/doi.acm.org\/10.1145\/2967934.2968083"},{"key":"e_1_2_1_52_1","volume-title":"Proceedings of the 12th International Conference on the Foundations of Digital Games","author":"Siu Kristin","year":"2071","unstructured":"Kristin Siu, Alexander Zook, and Mark O. Riedl. 2017b. A Framework for Exploring and Evaluating Mechanics in Human Computation Games. In Proceedings of the 12th International Conference on the Foundations of Digital Games (Hyannis, Massachusetts) (FDG '17). ACM, New York, NY, USA, Article 38, 4 pages. http:\/\/doi.acm.org\/10.1145\/3102071.3106344"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijhcs.2009.03.004"},{"key":"e_1_2_1_54_1","volume-title":"Going Deeper with Convolutions. arXiv:1409.4842 [cs] (Sep","author":"Szegedy Christian","year":"2014","unstructured":"Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2014. Going Deeper with Convolutions. arXiv:1409.4842 [cs] (Sep 2014). http:\/\/arxiv.org\/abs\/1409.4842 arXiv: 1409.4842."},{"key":"e_1_2_1_55_1","unstructured":"Tensorflow. 2020. Lucid. https:\/\/github.com\/tensorflow\/lucid."},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/mic.2012.67"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.3389\/frai.2022.826499"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/1378704.1378719"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/1124772.1124782"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397481.3450650"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/tvcg.2019.2934619"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3351095.3372852"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3359158"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/CIG.2018.8490433"}],"container-title":["Proceedings of the ACM on Human-Computer Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3610064","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3610064","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T04:26:26Z","timestamp":1755750386000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3610064"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,28]]},"references-count":64,"journal-issue":{"issue":"CSCW2","published-print":{"date-parts":[[2023,9,28]]}},"alternative-id":["10.1145\/3610064"],"URL":"https:\/\/doi.org\/10.1145\/3610064","relation":{},"ISSN":["2573-0142"],"issn-type":[{"value":"2573-0142","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,28]]},"assertion":[{"value":"2023-10-04","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}