{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T05:55:08Z","timestamp":1782280508913,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":116,"publisher":"ACM","license":[{"start":{"date-parts":[[2025,4,25]],"date-time":"2025-04-25T00:00:00Z","timestamp":1745539200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,4,26]]},"DOI":"10.1145\/3706598.3713491","type":"proceedings-article","created":{"date-parts":[[2025,4,24]],"date-time":"2025-04-24T04:45:58Z","timestamp":1745469958000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Exploring Empty Spaces: Human-in-the-Loop Data Augmentation"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-0429-4770","authenticated-orcid":false,"given":"Catherine","family":"Yeh","sequence":"first","affiliation":[{"name":"Harvard University, Boston, Massachusetts, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8666-7241","authenticated-orcid":false,"given":"Donghao","family":"Ren","sequence":"additional","affiliation":[{"name":"Apple, Seattle, Washington, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6646-0961","authenticated-orcid":false,"given":"Yannick","family":"Assogba","sequence":"additional","affiliation":[{"name":"Apple, Cambridge, Massachusetts, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3110-1053","authenticated-orcid":false,"given":"Dominik","family":"Moritz","sequence":"additional","affiliation":[{"name":"Apple, Pittsburgh, Pennsylvania, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4164-844X","authenticated-orcid":false,"given":"Fred","family":"Hohman","sequence":"additional","affiliation":[{"name":"Apple, Seattle, Washington, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,4,25]]},"reference":[{"key":"e_1_3_3_3_2_2","doi-asserted-by":"crossref","unstructured":"Yongsu Ahn and Yu-Ru Lin. 2019. Fairsight: Visual analytics for fairness in decision making. IEEE transactions on visualization and computer graphics 26 1 (2019) 1086\u20131095.","DOI":"10.1109\/TVCG.2019.2934262"},{"key":"e_1_3_3_3_3_2","doi-asserted-by":"crossref","unstructured":"Mehmet\u00a0Eren Ahsen Mehmet Ulvi\u00a0Saygi Ayvaci and Srinivasan Raghunathan. 2019. When algorithmic predictions use human-generated data: A bias-aware classification algorithm for breast cancer diagnosis. Information Systems Research 30 1 (2019) 97\u2013116.","DOI":"10.1287\/isre.2018.0789"},{"key":"e_1_3_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642858"},{"key":"e_1_3_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642016"},{"key":"e_1_3_3_3_6_2","unstructured":"Yannick Assogba Adam Pearce and Madison Elliott. 2023. Large scale qualitative evaluation of generative image model outputs."},{"key":"e_1_3_3_3_7_2","unstructured":"Yuntao Bai Andy Jones Kamal Ndousse Amanda Askell Anna Chen Nova DasSarma Dawn Drain Stanislav Fort Deep Ganguli Tom Henighan et\u00a0al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. 74\u00a0pages."},{"key":"e_1_3_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1609\/hcomp.v7i1.5285"},{"key":"e_1_3_3_3_9_2","doi-asserted-by":"publisher","unstructured":"Markus Bayer Marc-Andr\u00e9 Kaufhold and Christian Reuter. 2022. A Survey on Data Augmentation for Text Classification. ACM Comput. Surv. 55 7 Article 146 (dec 2022) 39\u00a0pages. 10.1145\/3544558","DOI":"10.1145\/3544558"},{"key":"e_1_3_3_3_10_2","doi-asserted-by":"crossref","unstructured":"Emma Beauxis-Aussalet Michael Behrisch Rita Borgo Duen\u00a0Horng Chau Christopher Collins David Ebert Mennatallah El-Assady Alex Endert Daniel\u00a0A Keim J\u00f6rn Kohlhammer et\u00a0al. 2021. The role of interactive visualization in fostering trust in AI. IEEE Computer Graphics and Applications 41 6 (2021) 7\u201312.","DOI":"10.1109\/MCG.2021.3107875"},{"key":"e_1_3_3_3_11_2","unstructured":"Sarah Bird Miro Dud\u00edk Richard Edgar Brandon Horn Roman Lutz Vanessa Milan Mehrnoosh Sameki Hanna Wallach and Kathleen Walker. 2020. Fairlearn: A toolkit for assessing and improving fairness in AI. Microsoft Tech. Rep. MSR-TR-2020-32 1 (2020) 7\u00a0pages."},{"key":"e_1_3_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3586183.3606725"},{"key":"e_1_3_3_3_13_2","unstructured":"Richard Brath Daniel Keim Johannes Knittel Shimei Pan Pia Sommerauer and Hendrik Strobelt. 2023. The role of interactive visualization in explaining (large) NLP models: from data to inference."},{"key":"e_1_3_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3311790.3396666"},{"key":"e_1_3_3_3_15_2","doi-asserted-by":"crossref","unstructured":"Benedetta Cevoli Chris Watkins and Kathleen Rastle. 2021. What is semantic diversity and why does it facilitate visual word recognition? Behavior research methods 53 (2021) 247\u2013263.","DOI":"10.3758\/s13428-020-01440-1"},{"key":"e_1_3_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Guoqing Chao Jingyao Liu Mingyu Wang and Dianhui Chu. 2023. Data augmentation for sentiment classification with semantic preservation and diversity. Knowledge-Based Systems 280 (2023) 111038.","DOI":"10.1016\/j.knosys.2023.111038"},{"key":"e_1_3_3_3_17_2","doi-asserted-by":"crossref","unstructured":"Nitesh\u00a0V Chawla Kevin\u00a0W Bowyer Lawrence\u00a0O Hall and W\u00a0Philip Kegelmeyer. 2002. SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research 16 (2002) 321\u2013357.","DOI":"10.1613\/jair.953"},{"key":"e_1_3_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.988"},{"key":"e_1_3_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.194"},{"key":"e_1_3_3_3_20_2","unstructured":"Heejae Chon Seonghyeon Lee Jinyoung Yeo and Dongha Lee. 2024. Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes."},{"key":"e_1_3_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3586183.3606777"},{"key":"e_1_3_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3491102.3501819"},{"key":"e_1_3_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW50498.2020.00359"},{"key":"e_1_3_3_3_24_2","unstructured":"Dan Dan\u00a0Friedman and Adji\u00a0Bousso Dieng. 2023. The vendi score: A diversity evaluation metric for machine learning. Transactions on machine learning research 1 (2023) 26\u00a0pages."},{"key":"e_1_3_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.97"},{"key":"e_1_3_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-2003"},{"key":"e_1_3_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-3405"},{"key":"e_1_3_3_3_28_2","unstructured":"Hugging Face. 2020. Wikipedia English Sentences Dataset. https:\/\/huggingface.co\/datasets\/sentence-transformers\/wikipedia-en-sentences. Accessed: 2024-08-29."},{"key":"e_1_3_3_3_29_2","unstructured":"Hugging Face. 2021. Introducing the Data Measurements Tool: an Interactive Tool for Looking at Datasets. https:\/\/huggingface.co\/blog\/data-measurements-tool."},{"key":"e_1_3_3_3_30_2","doi-asserted-by":"crossref","unstructured":"Michael Feffer Anusha Sinha Zachary\u00a0C Lipton and Hoda Heidari. 2024. Red-Teaming for Generative AI: Silver Bullet or Security Theater? 37\u00a0pages.","DOI":"10.1609\/aies.v7i1.31647"},{"key":"e_1_3_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.deelio-1.4"},{"key":"e_1_3_3_3_32_2","unstructured":"Yingchaojie Feng Zhizhang Chen Zhining Kang Sijia Wang Minfeng Zhu Wei Zhang and Wei Chen. 2024. Jailbreaklens: Visual analysis of jailbreak attacks against large language models."},{"key":"e_1_3_3_3_33_2","unstructured":"Yingchaojie Feng Xingbo Wang Kam\u00a0Kwai Wong Sijia Wang Yuhong Lu Minfeng Zhu Baicheng Wang and Wei Chen. 2023. Promptmagician: Interactive prompt engineering for text-to-image creation. IEEE Transactions on Visualization and Computer Graphics 30 (2023) 295\u2013305."},{"key":"e_1_3_3_3_34_2","unstructured":"Deep Ganguli Liane Lovitt Jackson Kernion Amanda Askell Yuntao Bai Saurav Kadavath Ben Mann Ethan Perez Nicholas Schiefer Kamal Ndousse et\u00a0al. 2022. Red teaming language models to reduce harms: Methods scaling behaviors and lessons learned. 30\u00a0pages."},{"key":"e_1_3_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642139"},{"key":"e_1_3_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.13175"},{"key":"e_1_3_3_3_37_2","unstructured":"Madeleine Grunde-McLaughlin Michelle\u00a0S Lam Ranjay Krishna Daniel\u00a0S Weld and Jeffrey Heer. 2023. Designing LLM chains by adapting techniques from crowdsourcing workflows."},{"key":"e_1_3_3_3_38_2","unstructured":"John Guerra. 2024. CHI 2024 Papers. https:\/\/observablehq.com\/@john-guerra\/chi2024-papers."},{"key":"e_1_3_3_3_39_2","unstructured":"Xu Guo and Yiqiang Chen. 2024. Generative AI for Synthetic Data Generation: Methods Challenges and the Future."},{"key":"e_1_3_3_3_40_2","doi-asserted-by":"crossref","unstructured":"Yuhan Guo Hanning Shao Can Liu Kai Xu and Xiaoru Yuan. 2024. PrompTHis: Visualizing the Process and Influence of Prompt Editing during Text-to-Image Creation. IEEE Transactions on Visualization and Computer Graphics Preprints (2024) 1\u201312.","DOI":"10.1109\/TVCG.2024.3408255"},{"key":"e_1_3_3_3_41_2","unstructured":"Vernon Toh\u00a0Yan Han Rishabh Bhardwaj and Soujanya Poria. 2024. Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming."},{"key":"e_1_3_3_3_42_2","unstructured":"Nicolas Heulot Jean-Daniel Fekete and Michael Aupetit. 2017. Visualizing dimensionality reduction artifacts: An evaluation."},{"key":"e_1_3_3_3_43_2","doi-asserted-by":"crossref","unstructured":"Fred Hohman Haekyu Park Caleb Robinson and Duen Horng\u00a0Polo Chau. 2019. S ummit: Scaling deep learning interpretability by visualizing activation and attribution summarizations. IEEE transactions on visualization and computer graphics 26 1 (2019) 1096\u20131106.","DOI":"10.1109\/TVCG.2019.2934659"},{"key":"e_1_3_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3313831.3376177"},{"key":"e_1_3_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305545"},{"key":"e_1_3_3_3_46_2","unstructured":"IBM. 2021. Data Quality for AI. https:\/\/www.ibm.com\/products\/dqaiapi."},{"key":"e_1_3_3_3_47_2","unstructured":"Albert\u00a0Q Jiang Alexandre Sablayrolles Arthur Mensch Chris Bamford Devendra\u00a0Singh Chaplot Diego de\u00a0las Casas Florian Bressand Gianna Lengyel Guillaume Lample Lucile Saulnier et\u00a0al. 2023. Mistral 7B. 9\u00a0pages."},{"key":"e_1_3_3_3_48_2","volume-title":"Next Generation of AI Safety Workshop","author":"Jiang Liwei","year":"2024","unstructured":"Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Nouha Dziri, and Yejin Choi. 2024. WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models. In Next Generation of AI Safety Workshop. ICML, Online, 51\u00a0pages. https:\/\/openreview.net\/forum?id=IRwWOprAPo"},{"key":"e_1_3_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613905.3650755"},{"key":"e_1_3_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIEM48762.2020.9160048"},{"key":"e_1_3_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ITNT60778.2024.10582292"},{"key":"e_1_3_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-6101"},{"key":"e_1_3_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2016-727"},{"key":"e_1_3_3_3_54_2","unstructured":"Yi-An Lai Xuan Zhu Yi Zhang and Mona Diab. 2020. Diversity density and homogeneity: Quantitative characteristic metrics for text collections."},{"key":"e_1_3_3_3_55_2","unstructured":"Harsh Lara and Manoj Tiwari. 2022. Evaluation of synthetic datasets for conversational recommender systems."},{"key":"e_1_3_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1014"},{"key":"e_1_3_3_3_57_2","unstructured":"Zhuoyan Li Hangxiao Zhu Zhuoran Lu and Ming Yin. 2023. Synthetic data generation with large language models for text classification: Potential and limitations."},{"key":"e_1_3_3_3_58_2","unstructured":"Lilac. 2023. Lilac: Better data better AI. https:\/\/www.lilacml.com."},{"key":"e_1_3_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.467"},{"key":"e_1_3_3_3_60_2","doi-asserted-by":"crossref","unstructured":"Leland McInnes John Healy and James Melville. 2018. Umap: Uniform manifold approximation and projection for dimension reduction. 63\u00a0pages.","DOI":"10.21105\/joss.00861"},{"key":"e_1_3_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3589335.3651929"},{"key":"e_1_3_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.765"},{"key":"e_1_3_3_3_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP48485.2024.10446564"},{"key":"e_1_3_3_3_64_2","unstructured":"OpenAI. 2024. GPT-4o mini. https:\/\/platform.openai.com\/docs\/models\/gpt-4o-mini."},{"key":"e_1_3_3_3_65_2","doi-asserted-by":"crossref","unstructured":"Will Orr and Kate Crawford. 2023. The social construction of datasets: On the practices processes and challenges of dataset creation for machine learning.","DOI":"10.31235\/osf.io\/8c9uh"},{"key":"e_1_3_3_3_66_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.208"},{"key":"e_1_3_3_3_67_2","doi-asserted-by":"crossref","unstructured":"Letian Peng Yuwei Zhang and Jingbo Shang. 2023. Controllable Data Augmentation for Few-Shot Text Mining with Chain-of-Thought Attribute Manipulation.","DOI":"10.18653\/v1\/2024.findings-acl.1"},{"key":"e_1_3_3_3_68_2","doi-asserted-by":"crossref","unstructured":"Tuan Pham Rob Hess Crystal Ju Eugene Zhang and Ronald Metoyer. 2010. Visualization of diversity in large multivariate data sets. IEEE Transactions on Visualization and Computer Graphics 16 6 (2010) 1053\u20131062.","DOI":"10.1109\/TVCG.2010.216"},{"key":"e_1_3_3_3_69_2","volume-title":"International Conference on Learning Representations","author":"Qu Yanru","year":"2021","unstructured":"Yanru Qu, Dinghan Shen, Yelong Shen, Sandra Sajeev, Weizhu Chen, and Jiawei Han. 2021. CoDA: Contrast-enhanced and Diversity-promoting Data Augmentation for Natural Language Understanding. In International Conference on Learning Representations. ICLR, Online, 14\u00a0pages."},{"key":"e_1_3_3_3_70_2","unstructured":"Senthooran Rajamanoharan Arthur Conmy Lewis Smith Tom Lieberum Vikrant Varma J\u00e1nos Kram\u00e1r Rohin Shah and Neel Nanda. 2024. Improving dictionary learning with gated sparse autoencoders. 37\u00a0pages."},{"key":"e_1_3_3_3_71_2","doi-asserted-by":"crossref","unstructured":"Gonzalo Ramos Christopher Meek Patrice Simard Jina Suh and Soroush Ghorashi. 2020. Interactive machine teaching: a human-centered approach to building machine-learned models. Human\u2013Computer Interaction 35 5-6 (2020) 413\u2013451.","DOI":"10.1080\/07370024.2020.1734931"},{"key":"e_1_3_3_3_72_2","first-page":"133","volume-title":"Conference on empirical methods in natural language processing","author":"Ratnaparkhi Adwait","year":"1996","unstructured":"Adwait Ratnaparkhi. 1996. A maximum entropy model for part-of-speech tagging. In Conference on empirical methods in natural language processing. EMNLP, Online, 133\u2013142."},{"key":"e_1_3_3_3_73_2","unstructured":"Sylvestre-Alvise Rebuffi Sven Gowal Dan\u00a0Andrei Calian Florian Stimberg Olivia Wiles and Timothy\u00a0A Mann. 2021. Data augmentation can improve robustness. Advances in Neural Information Processing Systems 34 (2021) 29935\u201329948."},{"key":"e_1_3_3_3_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/VIS54172.2023.00056"},{"key":"e_1_3_3_3_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613905.3650798"},{"key":"e_1_3_3_3_76_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_3_3_77_2","unstructured":"Google\u00a0People+AI Research. 2017. Facets - Visualizations for ML datasets. https:\/\/pair-code.github.io\/facets\/."},{"key":"e_1_3_3_3_78_2","unstructured":"Google\u00a0People+AI Research. 2019. Understanding UMAP. https:\/\/pair-code.github.io\/understanding-umap."},{"key":"e_1_3_3_3_79_2","unstructured":"Google\u00a0People+AI Research. 2021. Know Your Data. https:\/\/knowyourdata.withgoogle.com."},{"key":"e_1_3_3_3_80_2","unstructured":"Google\u00a0People+AI Research. 2021. Measuring Diversity. https:\/\/pair.withgoogle.com\/explorables\/measuring-diversity\/."},{"key":"e_1_3_3_3_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445518"},{"key":"e_1_3_3_3_82_2","unstructured":"Mikayel Samvelyan Sharath\u00a0Chandra Raparthy Andrei Lupu Eric Hambro Aram\u00a0H Markosyan Manish Bhatt Yuning Mao Minqi Jiang Jack Parker-Holder Jakob Foerster et\u00a0al. 2024. Rainbow teaming: Open-ended generation of diverse adversarial prompts."},{"key":"e_1_3_3_3_83_2","first-page":"123","volume-title":"Graphics Interface","author":"Sarvghad Ali","year":"2015","unstructured":"Ali Sarvghad and Melanie Tory. 2015. Exploiting analysis history to support collaborative data analysis.. In Graphics Interface. Academia.edu, Online, 123\u2013130."},{"key":"e_1_3_3_3_84_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401287"},{"key":"e_1_3_3_3_85_2","first-page":"68","volume-title":"Fourth International Workshop on Learning with Imbalanced Domains: Theory and Applications","author":"Shi Yiwen","year":"2022","unstructured":"Yiwen Shi, Taha ValizadehAslani, Jing Wang, Ping Ren, Yi Zhang, Meng Hu, Liang Zhao, and Hualou Liang. 2022. Improving imbalanced learning by pre-finetuning with data augmentation. In Fourth International Workshop on Learning with Imbalanced Domains: Theory and Applications. PMLR, Online, 68\u201382."},{"key":"e_1_3_3_3_86_2","doi-asserted-by":"crossref","unstructured":"Connor Shorten and Taghi\u00a0M Khoshgoftaar. 2019. A survey on image data augmentation for deep learning. Journal of big data 6 1 (2019) 1\u201348.","DOI":"10.1186\/s40537-019-0197-0"},{"key":"e_1_3_3_3_87_2","doi-asserted-by":"publisher","DOI":"10.1109\/IV.2016.57"},{"key":"e_1_3_3_3_88_2","unstructured":"Daniel Smilkov Nikhil Thorat Charles Nicholson Emily Reif Fernanda\u00a0B Vi\u00e9gas and Martin Wattenberg. 2016. Embedding projector: Interactive visualization and interpretation of embeddings."},{"key":"e_1_3_3_3_89_2","doi-asserted-by":"crossref","unstructured":"Hendrik Strobelt Sebastian Gehrmann Michael Behrisch Adam Perer Hanspeter Pfister and Alexander\u00a0M Rush. 2018. S eq 2s eq-v is: A visual debugging tool for sequence-to-sequence models. IEEE transactions on visualization and computer graphics 25 1 (2018) 353\u2013363.","DOI":"10.1109\/TVCG.2018.2865044"},{"key":"e_1_3_3_3_90_2","unstructured":"Yongduo Sui Qitian Wu Jiancan Wu Qing Cui Longfei Li Jun Zhou Xiang Wang and Xiangnan He. 2024. Unleashing the power of graph data augmentation on covariate distribution shift. Advances in Neural Information Processing Systems 36 (2024) 18109\u201318131."},{"key":"e_1_3_3_3_91_2","doi-asserted-by":"publisher","unstructured":"Mengdi Sun Ligan Cai Weiwei Cui Yanqiu Wu Yang Shi and Nan Cao. 2023. Erato: Cooperative Data Story Editing via Fact Interpolation. IEEE Transactions on Visualization and Computer Graphics 29 1 (2023) 983\u2013993. 10.1109\/TVCG.2022.3209428","DOI":"10.1109\/TVCG.2022.3209428"},{"key":"e_1_3_3_3_92_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.272"},{"key":"e_1_3_3_3_93_2","unstructured":"Simone Tedeschi Felix Friedrich Patrick Schramowski Kristian Kersting Roberto Navigli Huu Nguyen and Bo Li. 2024. ALERT: A Comprehensive Benchmark for Assessing Large Language Models\u2019 Safety through Red Teaming."},{"key":"e_1_3_3_3_94_2","unstructured":"Teknium. 2023. OpenHermes-2.5 Dataset. https:\/\/huggingface.co\/datasets\/teknium\/OpenHermes-2.5. Accessed: 2024-08-29."},{"key":"e_1_3_3_3_95_2","volume-title":"Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet","author":"Templeton Adly","year":"2024","unstructured":"Adly Templeton. 2024. Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. Anthropic, \u2019Online\u2019."},{"key":"e_1_3_3_3_96_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-demos.15"},{"key":"e_1_3_3_3_97_2","unstructured":"Miguel Tissera. 2023. Synthia v1.3 Dataset. https:\/\/huggingface.co\/datasets\/migtissera\/Synthia-v1.3. Accessed: 2024-08-29."},{"key":"e_1_3_3_3_98_2","first-page":"1730","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics","author":"Van\u00a0Miltenburg Emiel","year":"2018","unstructured":"Emiel Van\u00a0Miltenburg, Desmond Elliott, and Piek Vossen. 2018. Measuring the diversity of automatic image descriptions. In Proceedings of the 27th International Conference on Computational Linguistics. COLING, Online, 1730\u20131741."},{"key":"e_1_3_3_3_99_2","unstructured":"Stefan\u00a0Sylvius Wagner Maike Behrendt Marc Ziegele and Stefan Harmeling. 2024. SQBC: Active Learning using LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions."},{"key":"e_1_3_3_3_100_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.200"},{"key":"e_1_3_3_3_101_2","unstructured":"Su Wang Greg Durrett and Katrin Erk. 2020. Narrative interpolation for generating and understanding stories. 5\u00a0pages."},{"key":"e_1_3_3_3_102_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642335"},{"key":"e_1_3_3_3_103_2","doi-asserted-by":"crossref","unstructured":"Zijie\u00a0J Wang Robert Turko Omar Shaikh Haekyu Park Nilaksh Das Fred Hohman Minsuk Kahng and Duen Horng\u00a0Polo Chau. 2020. CNN explainer: learning convolutional neural networks with interactive visualization. IEEE Transactions on Visualization and Computer Graphics 27 2 (2020) 1396\u20131406.","DOI":"10.1109\/TVCG.2020.3030418"},{"key":"e_1_3_3_3_104_2","doi-asserted-by":"crossref","unstructured":"James Wexler Mahima Pushkarna Tolga Bolukbasi Martin Wattenberg Fernanda Vi\u00e9gas and Jimbo Wilson. 2019. The what-if tool: Interactive probing of machine learning models. IEEE transactions on visualization and computer graphics 26 1 (2019) 56\u201365.","DOI":"10.1109\/TVCG.2019.2934619"},{"key":"e_1_3_3_3_105_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.523"},{"key":"e_1_3_3_3_106_2","doi-asserted-by":"publisher","DOI":"10.1145\/3313831.3376451"},{"key":"e_1_3_3_3_107_2","doi-asserted-by":"crossref","unstructured":"Peipei Xia Li Zhang and Fanzhang Li. 2015. Learning similarity with cosine similarity ensemble. Information sciences 307 (2015) 39\u201352.","DOI":"10.1016\/j.ins.2015.02.024"},{"key":"e_1_3_3_3_108_2","unstructured":"Qizhe Xie Zihang Dai Eduard Hovy Thang Luong and Quoc Le. 2020. Unsupervised data augmentation for consistency training. Advances in neural information processing systems 33 (2020) 6256\u20136268."},{"key":"e_1_3_3_3_109_2","unstructured":"Catherine Yeh Gonzalo Ramos Rachel Ng Andy Huntington and Richard Banks. 2024. GhostWriter: Augmenting Collaborative Human-AI Writing Experiences Through Personalization and Agency. 29\u00a0pages."},{"key":"e_1_3_3_3_110_2","doi-asserted-by":"publisher","DOI":"10.1145\/3544548.3581388"},{"key":"e_1_3_3_3_111_2","doi-asserted-by":"crossref","unstructured":"Alice\u00a0Qian Zhang Ryland Shaw Jacy\u00a0Reese Anthis Ashlee Milton Emily Tseng Jina Suh Lama Ahmad Ram Shankar\u00a0Siva Kumar Julian Posada Benjamin Shestakofsky et\u00a0al. 2024. The Human Factor in AI Red Teaming: Perspectives from Social and Collaborative Computing.","DOI":"10.1145\/3678884.3687147"},{"key":"e_1_3_3_3_112_2","volume-title":"International Conference on Learning Representations","author":"Zhang Hongyi","year":"2018","unstructured":"Hongyi Zhang, Moustapha Cisse, Yann\u00a0N. Dauphin, and David Lopez-Paz. 2018. mixup: Beyond Empirical Risk Minimization. In International Conference on Learning Representations. ICLR, Online, 13\u00a0pages. https:\/\/openreview.net\/forum?id=r1Ddp1-Rb"},{"key":"e_1_3_3_3_113_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413772"},{"key":"e_1_3_3_3_114_2","unstructured":"Dorothy Zhao Jerone\u00a0TA Andrews AI Sony Tokyo\u00a0Orestis Papakyriakopoulos and Alice Xiang. 2024. Measuring Diversity in Datasets. International Conference on Learning Representations 1 (2024) 36\u00a0pages."},{"key":"e_1_3_3_3_115_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Zhao Wenting","year":"2024","unstructured":"Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024. WildChat: 1M ChatGPT Interaction Logs in the Wild. In The Twelfth International Conference on Learning Representations. ICLR, Online, 16\u00a0pages. https:\/\/openreview.net\/forum?id=Bl8u7ZRlbM"},{"key":"e_1_3_3_3_116_2","unstructured":"Lianmin Zheng Wei-Lin Chiang Ying Sheng Tianle Li Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zhuohan Li Zi Lin Eric.\u00a0P Xing Joseph\u00a0E. Gonzalez Ion Stoica and Hao Zhang. 2023. LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset. arxiv:https:\/\/arXiv.org\/abs\/2309.11998\u00a0[cs.CL]"},{"key":"e_1_3_3_3_117_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210080"}],"event":{"name":"CHI 2025: CHI Conference on Human Factors in Computing Systems","location":"Yokohama Japan","acronym":"CHI '25","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3706598.3713491","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3706598.3713491","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,4]],"date-time":"2025-07-04T06:07:38Z","timestamp":1751609258000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3706598.3713491"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,25]]},"references-count":116,"alternative-id":["10.1145\/3706598.3713491","10.1145\/3706598"],"URL":"https:\/\/doi.org\/10.1145\/3706598.3713491","relation":{},"subject":[],"published":{"date-parts":[[2025,4,25]]},"assertion":[{"value":"2025-04-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}