{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:46:11Z","timestamp":1787017571699,"version":"3.56.0"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"ISSTA","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,22]]},"abstract":"<jats:p>Despite advances in automated testing, manual testing remains prevalent due to the high maintenance demands associated with test script fragility\u2014scripts often break with minor changes in application structure. Recent developments in Large Language Models (LLMs) offer a potential alternative by powering Autonomous Web Agents (AWAs) that can autonomously interact with applications. These agents may serve as Autonomous Test Agents (ATAs), potentially reducing the need for maintenance-heavy automated scripts by utilising natural language instructions similar to those used by human testers.  \nThis paper investigates the feasibility of adapting AWAs for natural language test case execution and how to evaluate them.<\/jats:p>\n                  <jats:p>We contribute with (1) a benchmark of three offline web applications, and a suite of 113 manual test cases, split between passing and failing cases, to evaluate and compare ATAs performance, (2) SeeAct-ATA and pinATA, two open-source ATA implementations capable of executing test steps, verifying assertions and giving verdicts, and (3) comparative experiments using our benchmark that quantifies our ATAs effectiveness. Finally we also proceed to a qualitative evaluation to identify the limitations of PinATA, our best performing implementation.<\/jats:p>\n                  <jats:p>Our findings reveal that our simple implementation, SeeAct-ATA, does not perform well compared to our more advanced PinATA implementation when executing test cases (50% performance improvement). However, while PinATA obtains around 60% of correct verdict and up to a promising 94% specificity, we identify several limitations that need to be addressed to develop more resilient and reliable ATAs, paving the way for robust, low maintenance test automation.<\/jats:p>","DOI":"10.1145\/3728879","type":"journal-article","created":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T10:52:56Z","timestamp":1750589576000},"page":"206-228","source":"Crossref","is-referenced-by-count":9,"title":["Are Autonomous Web Agents Good Testers?"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3677-5150","authenticated-orcid":false,"given":"Antoine","family":"Chevrot","sequence":"first","affiliation":[{"name":"Smartesting, Besan\u00e7on, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2113-4900","authenticated-orcid":false,"given":"Alexandre","family":"Vernotte","sequence":"additional","affiliation":[{"name":"Smartesting, Besan\u00e7on, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8284-7218","authenticated-orcid":false,"given":"Jean-R\u00e9my","family":"Falleri","sequence":"additional","affiliation":[{"name":"University of Bordeaux - LaBRI - UMR 5800, Bordeaux, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1783-0708","authenticated-orcid":false,"given":"Xavier","family":"Blanc","sequence":"additional","affiliation":[{"name":"University of Bordeaux, Bordeaux, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4986-7097","authenticated-orcid":false,"given":"Bruno","family":"Legeard","sequence":"additional","affiliation":[{"name":"Smartesting, Besan\u00e7on, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-7686-136X","authenticated-orcid":false,"given":"Aymeric","family":"Cretin","sequence":"additional","affiliation":[{"name":"Smartesting, Besan\u00e7on, France"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,22]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2011.5958253"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/S1389-1286(00)00073-6"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2021.106754"},{"key":"e_1_2_1_4_1","volume-title":"Black","author":"Chandu Khyathi Raghavi","year":"2021","unstructured":"Khyathi Raghavi Chandu, Yonatan Bisk, and Alan W. Black. 2021. Grounding \u2019Grounding\u2019 in NLP. arxiv:2106.02192 arXiv:2106.02192 [cs]"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","unstructured":"Qi Chen Dileepa Pitawela Chongyang Zhao Gengze Zhou Hsiang-Ting Chen and Qi Wu. 2023. WebVLN: Vision-and-Language Navigation on Websites. https:\/\/doi.org\/10.48550\/arXiv.2312.15820 arXiv:2312.15820 [cs] 10.48550\/arXiv.2312.15820","DOI":"10.48550\/arXiv.2312.15820"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-74296-6_29"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","unstructured":"Antoine Chevrot Alexandre Vernotte Jean-R\u00e9my Falleri-Vialard Xavier Blanc and Bruno Legeard. 2025. Autonomous Tester Agent Benchmark. https:\/\/doi.org\/10.5281\/zenodo.15198569 10.5281\/zenodo.15198569","DOI":"10.5281\/zenodo.15198569"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","unstructured":"Xiang Deng Yu Gu Boyuan Zheng Shijie Chen Samuel Stevens Boshi Wang Huan Sun and Yu Su. 2023. Mind2Web: Towards a Generalist Agent for the Web. https:\/\/doi.org\/10.48550\/arXiv.2306.06070 arXiv:2306.06070 [cs] 10.48550\/arXiv.2306.06070","DOI":"10.48550\/arXiv.2306.06070"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3661167.3661179"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.23919\/CISTI.2019.8760848"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","unstructured":"Robert Feldt Sungmin Kang Juyeon Yoon and Shin Yoo. 2023. Towards Autonomous Testing Agents via Conversational Large Language Models. https:\/\/doi.org\/10.48550\/arXiv.2306.05152 arXiv:2306.05152 [cs] 10.48550\/arXiv.2306.05152","DOI":"10.48550\/arXiv.2306.05152"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-70245-7_10"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","unstructured":"Luca Gioacchini Giuseppe Siracusano Davide Sanvito Kiril Gashteovski David Friede Roberto Bifulco and Carolin Lawrence. 2024. AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents. https:\/\/doi.org\/10.48550\/arXiv.2404.06411 arXiv:2404.06411 [cs] 10.48550\/arXiv.2404.06411","DOI":"10.48550\/arXiv.2404.06411"},{"key":"e_1_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Yu Gu Xiang Deng and Yu Su. 2023. Don\u2019t Generate Discriminate: A Proposal for Grounding Language Models to Real-World Environments. arxiv:2212.09736 arXiv:2212.09736 [cs]","DOI":"10.18653\/v1\/2023.acl-long.270"},{"key":"e_1_2_1_15_1","volume-title":"Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust.","author":"Gur Izzeddin","year":"2023","unstructured":"Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. 2023. A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis. arxiv:2307.12856 arXiv:2307.12856 [cs]"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2950290.2950294"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICST.2016.16"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","unstructured":"Hongliang He Wenlin Yao Kaixin Ma Wenhao Yu Yong Dai Hongming Zhang Zhenzhong Lan and Dong Yu. 2024. WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models. https:\/\/doi.org\/10.48550\/arXiv.2401.13919 arXiv:2401.13919 [cs] 10.48550\/arXiv.2401.13919","DOI":"10.48550\/arXiv.2401.13919"},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations)","author":"Iong Iat Long","year":"2024","unstructured":"Iat Long Iong, Xiao Liu, Yuxuan Chen, Hanyu Lai, Shuntian Yao, Pengbo Shen, Hao Yu, Yuxiao Dong, and Jie Tang. 2024. OpenWebAgent: An Open Toolkit to Enable Web Agents on Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), Yixin Cao, Yang Feng, and Deyi Xiong (Eds.). Association for Computational Linguistics, Bangkok, Thailand. 72\u201381. https:\/\/aclanthology.org\/2024.acl-demos.8"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3637528.3671620"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-08245-5_19"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICST57152.2023.00039"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSREW.2014.17"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICST.2015.7102611"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639180"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3417069"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1002\/stvr.1760"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","unstructured":"Yujia Qin Yining Ye Junjie Fang Haoming Wang Shihao Liang Shizuo Tian Junda Zhang Jiahao Li Yunxin Li Shijue Huang Wanjun Zhong Kuanye Li Jiale Yang Yu Miao Woyu Lin Longxiang Liu Xu Jiang Qianli Ma Jingyu Li Xiaojun Xiao Kai Cai Chuang Li Yaowei Zheng Chaolin Jin Chen Li Xiao Zhou Minchao Wang Haoli Chen Zhaojian Li Haihua Yang Haifeng Liu Feng Lin Tao Peng Xin Liu and Guang Shi. 2025. UI-TARS: Pioneering Automated GUI Interaction with Native Agents. https:\/\/doi.org\/10.48550\/arXiv.2501.12326 arXiv:2501.12326 10.48550\/arXiv.2501.12326","DOI":"10.48550\/arXiv.2501.12326"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","unstructured":"Filippo Ricca Maurizio Leotta and Andrea Stocco. 2019. Chapter Three - Three Open Problems in the Context of E2E Web Testing and a Vision: NEONATE. In Advances in Computers Atif M. Memon (Ed.). 113 Elsevier Philadelphia PA USA. 89\u2013133. https:\/\/doi.org\/10.1016\/bs.adcom.2018.10.005 ISSN: 0065-2458 10.1016\/bs.adcom.2018.10.005","DOI":"10.1016\/bs.adcom.2018.10.005"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.2105\/AJPH.89.8.1175"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3236024.3236063"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3640794.3665887"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2668930.2688819"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3368208"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-024-40231-1"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2201.11903"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4615-4625-2"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2412.14161"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","unstructured":"Jianwei Yang Hao Zhang Feng Li Xueyan Zou Chunyuan Li and Jianfeng Gao. 2023. Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V. https:\/\/doi.org\/10.48550\/arXiv.2310.11441 arXiv:2310.11441 [cs] 10.48550\/arXiv.2310.11441","DOI":"10.48550\/arXiv.2310.11441"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3691620.3695529"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/QRS60937.2023.00029"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","unstructured":"Yao Zhang Zijian Ma Yunpu Ma Zhen Han Yu Wu and Volker Tresp. 2024. WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration. https:\/\/doi.org\/10.48550\/arXiv.2408.15978 arXiv:2408.15978 [cs] 10.48550\/arXiv.2408.15978","DOI":"10.48550\/arXiv.2408.15978"},{"key":"e_1_2_1_43_1","doi-asserted-by":"crossref","unstructured":"Zeyu Zhang Xiaohe Bo Chen Ma Rui Li Xu Chen Quanyu Dai Jieming Zhu Zhenhua Dong and Ji-Rong Wen. 2024. A Survey on the Memory Mechanism of Large Language Model based Agents. arxiv:2404.13501 arXiv:2404.13501 [cs]","DOI":"10.1145\/3748302"},{"key":"e_1_2_1_44_1","unstructured":"Zhehao Zhang Ryan Rossi Tong Yu Franck Dernoncourt Ruiyi Zhang Jiuxiang Gu Sungchul Kim Xiang Chen Zichao Wang and Nedim Lipka. 2024. VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use. arxiv:2410.16400 arXiv:2410.16400"},{"key":"e_1_2_1_45_1","unstructured":"Boyuan Zheng Boyu Gou Jihyung Kil Huan Sun and Yu Su. 2024. GPT-4V(ision) is a Generalist Web Agent if Grounded. In Proceedings of the 41st International Conference on Machine Learning Ruslan Salakhutdinov Zico Kolter Katherine Heller Adrian Weller Nuria Oliver Jonathan Scarlett and Felix Berkenkamp (Eds.) (Proceedings of Machine Learning Research Vol. 235). PMLR Vienna Austria. 61349\u201361385. https:\/\/proceedings.mlr.press\/v235\/zheng24e.html"},{"key":"e_1_2_1_46_1","unstructured":"Shuyan Zhou Frank F. Xu Hao Zhu Xuhui Zhou Robert Lo Abishek Sridhar Xianyi Cheng Tianyue Ou Yonatan Bisk Daniel Fried Uri Alon and Graham Neubig. 2024. WebArena: A Realistic Web Environment for Building Autonomous Agents. arxiv:2307.13854 arXiv:2307.13854 [cs]"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","unstructured":"Mingchen Zhuge Changsheng Zhao Dylan Ashley Wenyi Wang Dmitrii Khizbullin Yunyang Xiong Zechun Liu Ernie Chang Raghuraman Krishnamoorthi Yuandong Tian Yangyang Shi Vikas Chandra and J\u00fcrgen Schmidhuber. 2024. Agent-as-a-Judge: Evaluate Agents with Agents. https:\/\/doi.org\/10.48550\/arXiv.2410.10934 arXiv:2410.10934 [cs] version: 1 10.48550\/arXiv.2410.10934","DOI":"10.48550\/arXiv.2410.10934"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-804206-9.00027-1"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728879","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,16]],"date-time":"2025-07-16T16:55:34Z","timestamp":1752684934000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728879"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,22]]},"references-count":48,"journal-issue":{"issue":"ISSTA","published-print":{"date-parts":[[2025,6,22]]}},"alternative-id":["10.1145\/3728879"],"URL":"https:\/\/doi.org\/10.1145\/3728879","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,22]]}}}