{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,21]],"date-time":"2026-04-21T14:50:54Z","timestamp":1776783054684,"version":"3.51.2"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,4,18]],"date-time":"2024-04-18T00:00:00Z","timestamp":1713398400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Engineering Research Center Program through the National Research Foundation of Korea","award":["NRF-2018R1A5A1059921"],"award-info":[{"award-number":["NRF-2018R1A5A1059921"]}]},{"name":"Institute of Information & Communications Technology Planning & Evaluation","award":["2018-0-00769"],"award-info":[{"award-number":["2018-0-00769"]}]},{"name":"Swedish Scientific Council","award":["2020-05272"],"award-info":[{"award-number":["2020-05272"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2024,5,31]]},"abstract":"<jats:p>Due to the rapid adoption of Deep Neural Networks (DNNs) into larger software systems, testing of DNN-based systems has received much attention recently. While many different test adequacy criteria have been suggested, we lack effective test input generation techniques. Inputs such as images of real-world objects and scenes are not only expensive to collect but also difficult to randomly sample. Consequently, current testing techniques for DNNs tend to apply small local perturbations to existing inputs to generate new inputs. We propose SINVAD (Search-based Input space Navigation using Variational AutoencoDers), a way to sample from, and navigate over, a space of realistic inputs that resembles the true distribution in the training data. Our input space is constructed using Variational Autoencoders (VAEs), and navigated through their latent vector space. Our analysis shows that the VAE-based input space is well-aligned with human perception of what constitutes realistic inputs. Further, we show that this space can be effectively searched to achieve various testing scenarios, such as boundary testing of two different DNNs or analyzing class labels that are difficult for the given DNN to distinguish. Guidelines on how to design VAE architectures are presented as well. Our results have the potential to open the field to meaningful exploration through the space of highly structured images.<\/jats:p>","DOI":"10.1145\/3635706","type":"journal-article","created":{"date-parts":[[2023,12,21]],"date-time":"2023-12-21T11:53:05Z","timestamp":1703159585000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Deceiving Humans and Machines Alike: Search-based Test Input Generation for DNNs Using Variational Autoencoders"],"prefix":"10.1145","volume":"33","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0298-5320","authenticated-orcid":false,"given":"Sungmin","family":"Kang","sequence":"first","affiliation":[{"name":"Korea Advanced Institute of Science and Technology, Daejeon, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5179-4205","authenticated-orcid":false,"given":"Robert","family":"Feldt","sequence":"additional","affiliation":[{"name":"Chalmers University of Technology, Gothenburg, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0836-6993","authenticated-orcid":false,"given":"Shin","family":"Yoo","sequence":"additional","affiliation":[{"name":"Korea Advanced Institute of Science and Technology, Daejeon, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,4,18]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Andrea Agostinelli Timo I. Denk Zal\u00e1n Borsos Jesse Engel Mauro Verzetti Antoine Caillon Qingqing Huang Aren Jansen Adam Roberts Marco Tagliasacchi Matt Sharifi Neil Zeghidour and Christian Frank. 2023. MusicLM: Generating Music From Text. https:\/\/arxiv.org\/pdf\/2301.11325.pdf"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.153.3731.34"},{"key":"e_1_3_2_4_2","unstructured":"Tom Brown Dandelion Mane Aurko Roy Martin Abadi and Justin Gilmer. 2017. Adversarial patch. Retrieved from https:\/\/arxiv.org\/pdf\/1712.09665.pdf"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377816.3381734"},{"key":"e_1_3_2_6_2","first-page":"15","volume-title":"Proceedings of the IEEE International Conference on Artificial Intelligence Testing (AITest\u201920)","author":"Byun Taejoon","year":"2020","unstructured":"Taejoon Byun, Abhishek Vijayakumar, Sanjai Rayadurgam, and Darren Cofer. 2020. Manifold-based test generation for image classifiers. In Proceedings of the IEEE International Conference on Artificial Intelligence Testing (AITest\u201920). IEEE, 15\u201322."},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","unstructured":"Nicholas Carlini and David A. Wagner. 2016. Towards evaluating the robustness of neural networks. Retrieved from http:\/\/arxiv.org\/abs\/1608.04644","DOI":"10.1109\/SP.2017.49"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.312"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.691"},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","unstructured":"Gregory Cohen Saeed Afshar Jonathan Tapson and Andr\u00e9 van Schaik. 2017. EMNIST: An Extension of MNIST to Handwritten letters. https:\/\/arxiv.org\/pdf\/1702.05373.pdf","DOI":"10.1109\/IJCNN.2017.7966217"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Ekin Dogus Cubuk Barret Zoph Dandelion Mane Vijay Vasudevan and Quoc V. Le. 2019. AutoAugment: Learning augmentation policies from data. Retrieved from https:\/\/arxiv.org\/pdf\/1805.09501.pdf","DOI":"10.1109\/CVPR.2019.00020"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_13_2","volume-title":"Generative Deep Learning","author":"Foster David","year":"2019","unstructured":"David Foster. 2019. Generative Deep Learning. O\u2019Reilly."},{"key":"e_1_3_2_14_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Goodfellow Ian","year":"2015","unstructured":"Ian Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and harnessing adversarial examples. In Proceedings of the International Conference on Learning Representations. Retrieved from http:\/\/arxiv.org\/abs\/1412.6572"},{"key":"e_1_3_2_15_2","first-page":"2672","volume-title":"Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS\u201914)","author":"Goodfellow Ian J.","year":"2014","unstructured":"Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS\u201914). MIT Press, 2672\u20132680."},{"key":"e_1_3_2_16_2","unstructured":"Louay Hazami Rayhane Mama and Ragavan Thurairatnam. 2022. Efficient-VDVAE: Less is more. Retrieved from https:\/\/arXiv:2203.13751"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-63387-9_1"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"Robin Jia and Percy Liang. 2017. Adversarial examples for evaluating reading comprehension systems. Retrieved from https:\/\/arxiv.org\/abs\/1707.07328","DOI":"10.18653\/v1\/D17-1215"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3387940.3391456"},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","unstructured":"Guy Katz Clark W. Barrett David L. Dill Kyle Julian and Mykel J. Kochenderfer. 2017. Reluplex: An efficient SMT solver for verifying deep neural networks. Retrieved from http:\/\/arxiv.org\/abs\/1702.01135","DOI":"10.1007\/978-3-319-63387-9_5"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00108"},{"key":"e_1_3_2_22_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201914)","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Max Welling. 2014. Auto-encoding variational bayes. In Proceedings of the International Conference on Learning Representations (ICLR\u201914)."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3178876.3186133"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_2_25_2","unstructured":"Alexey Kurakin Ian J. Goodfellow and Samy Bengio. 2016. Adversarial examples in the physical world. Retrieved from https:\/\/arxiv.org\/abs\/1607.02533"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_2_27_2","unstructured":"Yann LeCun Corinna Cortes and C. J. Burges. 2010. MNIST handwritten digit database. Retrieved from http:\/\/yann.lecun.com\/exdb\/mnist"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3238147.3238202"},{"key":"e_1_3_2_29_2","first-page":"427","article-title":"Deep neural networks are easily fooled: High confidence predictions for unrecognizable images","author":"Nguyen Anh Mai","year":"2015","unstructured":"Anh Mai Nguyen, Jason Yosinski, and Jeff Clune. 2015. Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915). 427\u2013436.","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915)"},{"key":"e_1_3_2_30_2","doi-asserted-by":"crossref","unstructured":"Nicolas Papernot Patrick D. McDaniel Somesh Jha Matt Fredrikson Z. Berkay Celik and Ananthram Swami. 2015. The limitations of deep learning in adversarial settings. Retrieved from https:\/\/arxiv.org\/abs\/1511.07528","DOI":"10.1109\/EuroSP.2016.36"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132747.3132785"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","unstructured":"Andr\u00e9 Reichstaller and Alexander Knapp. 2017. Compressing Uniform Test Suites Using Variational Autoencoders. In Proceedings of the IEEE International Conference on Software Quality Reliability and Security Companion (QRS-C\u201917)10.1109\/QRS-C.2017.128","DOI":"10.1109\/QRS-C.2017.128"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409730"},{"key":"e_1_3_2_34_2","unstructured":"Fanny Roche Thomas Hueber Samuel Limier and Laurent Girin. 2018. Autoencoders for Music Sound Synthesis: A Comparison of Linear Shallow Deep and Variational Models. https:\/\/arxiv.org\/pdf\/1806.04096.pdf"},{"key":"e_1_3_2_35_2","first-page":"1","volume-title":"Proceedings of the Annual Meeting of the Southern Association for Institutional Research","author":"Romano Jeanine","year":"2006","unstructured":"Jeanine Romano, Jeffrey D. Kromrey, Jesse Coraggio, Jeff Skowronek, and Linda Devine. 2006. Exploring methods for evaluating group differences on the NSSE and other surveys: Are the t-test and Cohen\u2019s d indices the most appropriate choices. In Proceedings of the Annual Meeting of the Southern Association for Institutional Research. Citeseer, 1\u201351."},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","unstructured":"Mostafa Sadeghi Simon Leglaive Xavier Alameda-Pineda Laurent Girin and Radu Horaud. 2019. Audio-Visual Speech Enhancement Using Conditional Variational Auto-Encoder. https:\/\/arxiv.org\/pdf\/1908.02590.pdf","DOI":"10.1109\/TASLP.2020.3000593"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","unstructured":"Lea Sch\u00f6nherr Katharina Kohls Steffen Zeiler Thorsten Holz and Dorothea Kolossa. 2018. Adversarial attacks against automatic speech recognition systems via psychoacoustic hiding. Retrieved from https:\/\/arxiv.org\/abs\/1808.05665","DOI":"10.14722\/ndss.2019.23288"},{"key":"e_1_3_2_38_2","volume-title":"Advances in Neural Information Processing Systems","author":"Sohn Kihyuk","year":"2015","unstructured":"Kihyuk Sohn, Honglak Lee, and Xinchen Yan. 2015. Learning structured output representation using deep conditional generative models. In Advances in Neural Information Processing Systems, C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett (Eds.), Vol. 28. Curran Associates. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2015\/file\/8d55a249e6baa5c06772297520da2051-Paper.pdf"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.5555\/3327757.3327924"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3180155.3180220"},{"key":"e_1_3_2_41_2","unstructured":"Han Xiao Kashif Rasul and Roland Vollgraf. 2017. Fashion-MNIST: A Novel Image Dataset for Benchmarking Machine Learning Algorithms. https:\/\/arxiv.org\/pdf\/1708.07747.pdf"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/SBST.2019.000-2"},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","unstructured":"Jie M. Zhang Mark Harman Lei Ma and Yang Liu. 2020. Machine learning testing: Survey landscapes and horizons. IEEE Transactions on Software Engineering 48 1 (2020) 1\u201336.","DOI":"10.1109\/TSE.2019.2962027"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3238147.3238187"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3635706","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3635706","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:56:59Z","timestamp":1750291019000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3635706"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,18]]},"references-count":43,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,5,31]]}},"alternative-id":["10.1145\/3635706"],"URL":"https:\/\/doi.org\/10.1145\/3635706","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,18]]},"assertion":[{"value":"2022-04-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-07","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}