{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T03:56:03Z","timestamp":1782878163625,"version":"3.54.5"},"reference-count":94,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>Automated web GUI testing approaches aim to maximize the code coverage of a web app within a specific time budget. However, due to the highly dynamic characteristics of web apps, testing approaches often get stuck in loops or repeatedly explore the same app areas. To address this issue, existing approaches conduct state abstraction, grouping similar pages into the same state in an effort to approximate the ideal state (i.e., a state that encompasses all-and-only those pages exhibiting the same behavior from a testing perspective) to reduce repetitive explorations. Typically, these approaches rely on the Document Object Model (DOM) or visual similarity, using predefined thresholds or learning-based classifiers to determine which pages should belong to the same state. However, pages within the same ideal state still exhibit discrepancies, caused by factors such as dynamically loaded data and dynamically expanded UI elements. The varying page complexities and design styles among apps bring even more challenges. These phenomena present substantial obstacles to existing approaches in determining desirable classification thresholds or training desirable classifiers, preventing them from conducting satisfactory state abstraction to guide the testing process.<\/jats:p>\n                  <jats:p>To address the preceding challenges, in this article, we propose Judge, a novel approach based on structure merging and contrastive learning for state abstraction. Judge includes a \u201cmerge-and-classify\u201d strategy. In the \u201cmerge\u201d phase, Judge iterates through the DOM tree of each given page and merges web element siblings that share the same subtree structure into a single one to abstract and simplify the page, while discarding text contents and HTML attributes of web elements in the process. In this way, Judge mitigates the negative effects introduced by dynamically loaded data and dynamically expanded UI elements, substantially reducing discrepancies between pages in the same ideal state. In the \u201cclassify\u201d phase, Judge uses a dedicated contrastive learning model to embed simplified page DOMs into vectors and further conducts classification with a Support Vector Machine (SVM), enabling classification in high-dimensional vector space and improving generalizability across diverse web apps.<\/jats:p>\n                  <jats:p>We evaluate Judge against 13 widely used baseline approaches. The results highlight that Judge outperforms these baseline approaches in classifying page pairs, with an average margin ranging from 8.95% to 28.90% in the F1 score across three manually labeled datasets. Additionally, when compared to the five most effective baseline approaches, Judge demonstrates superiority in guiding the exploration of automated web GUI testing in six widely studied apps, with code coverage improved by an average of 2.62\u201314.12%. The code and data of Judge are publicly accessible.<\/jats:p>","DOI":"10.1145\/3736162","type":"journal-article","created":{"date-parts":[[2025,5,16]],"date-time":"2025-05-16T12:49:28Z","timestamp":1747399768000},"page":"1-40","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Judge: Effective State Abstraction for Guiding Automated Web GUI Testing"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4604-491X","authenticated-orcid":false,"given":"Chenxu","family":"Liu","sequence":"first","affiliation":[{"name":"Key Lab of HCST (PKU), MOE, SCS, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-3215-9700","authenticated-orcid":false,"given":"Junheng","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Software and Microelectronics, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5338-7347","authenticated-orcid":false,"given":"Wei","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Richardson, Texas, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-6924-2319","authenticated-orcid":false,"given":"Ying","family":"Zhang","sequence":"additional","affiliation":[{"name":"Key Lab of HCST (PKU) &amp; National Engineering Research Center of Software Engineering, MOE, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6731-216X","authenticated-orcid":false,"given":"Tao","family":"Xie","sequence":"additional","affiliation":[{"name":"Key Lab of HCST (PKU), MOE, SCS, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,2,13]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10270-004-0077-7"},{"key":"e_1_3_2_3_2","unstructured":"Spring Petclinic Angular. 2025. Retrieved from https:\/\/github.com\/spring-petclinic\/spring-petclinic-angular"},{"key":"e_1_3_2_4_2","unstructured":"Iz Beltagy Matthew E. Peters and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv:2004.05150. Retrieved from https:\/\/arxiv.org\/abs\/2004.05150"},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1007\/978-3-319-66299-2_2","volume-title":"Proceedings of the 9th International Symposium on Search-Based Software Engineering","author":"Biagiola Matteo","year":"2017","unstructured":"Matteo Biagiola, Filippo Ricca, and Paolo Tonella. 2017. Search based path and input data generation for web application testing. In Proceedings of the 9th International Symposium on Search-Based Software Engineering, 18\u201332."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3338906.3338970"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0169-7552(97)00031-7"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3495883"},{"key":"e_1_3_2_9_2","first-page":"406","volume-title":"Proceedings of the 5th Asian-Pacific Web Conference","author":"Cai Deng","year":"2003","unstructured":"Deng Cai, Shipeng Yu, Ji-Rong Wen, and Wei-Ying Ma. 2003. Extracting content structure for web pages based on visual representation. In Proceedings of the 5th Asian-Pacific Web Conference, 406\u2013417."},{"key":"e_1_3_2_10_2","first-page":"13","volume-title":"Proceedings of the 2023 IEEE\/ACM International Conference on Automation of Software Test","author":"Chang Xiaoning","year":"2023","unstructured":"Xiaoning Chang, Zheheng Liang, Yifei Zhang, Lei Cui, Zhenyue Long, Guoquan Wu, Yu Gao, Wei Chen, Jun Wei, and Tao Huang. 2023. A reinforcement learning approach to generating test cases for web applications. In Proceedings of the 2023 IEEE\/ACM International Conference on Automation of Software Test, 13\u201323."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/509907.509965"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.5555\/3524938.3525087"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3497589"},{"key":"e_1_3_2_14_2","unstructured":"Xinlei Chen Haoqi Fan Ross B. Girshick and Kaiming He. 2020. Improved baselines with momentum contrastive learning. arXiv:2003.04297. Retrieved from https:\/\/arxiv.org\/abs\/2003.04297"},{"key":"e_1_3_2_15_2","first-page":"1","volume-title":"Proceedings of the 15th ACM\/IEEE International Symposium on Empirical Software Engineering and Measurement","volume":"37","author":"Corazza Anna","year":"2021","unstructured":"Anna Corazza, Sergio Di Martino, Adriano Peron, and Luigi Libero Lucio Starace. 2021. Web application testing: Using tree kernels to detect near-duplicate states in automated model inference. In Proceedings of the 15th ACM\/IEEE International Symposium on Empirical Software Engineering and Measurement, Article 37, 1\u20136."},{"key":"e_1_3_2_16_2","unstructured":"Near-Duplicate Study Dataset. 2019. Retrieved from https:\/\/zenodo.org\/records\/3376730"},{"key":"e_1_3_2_17_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171\u20134186."},{"key":"e_1_3_2_18_2","unstructured":"Difflib. 2025. Retrieved from https:\/\/docs.python.org\/3\/library\/difflib.html"},{"key":"e_1_3_2_19_2","unstructured":"Dimeshift. 2025. Retrieved from https:\/\/github.com\/jeka-kiselyov\/dimeshift"},{"key":"e_1_3_2_20_2","unstructured":"EmberJS. 2025. Retrieved from https:\/\/emberjs.com\/"},{"key":"e_1_3_2_21_2","unstructured":"Hugging Face. 2025. Retrieved from https:\/\/huggingface.co\/"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Hongchao Fang and Pengtao Xie. 2020. CERT: Contrastive self-supervised learning for language understanding. arXiv:2005.12766. Retrieved from https:\/\/arxiv.org\/abs\/2005.12766","DOI":"10.36227\/techrxiv.12308378"},{"key":"e_1_3_2_23_2","first-page":"278","volume-title":"Proceedings of the 24th International Symposium on Software Reliability Engineering","author":"Fard Amin Milani","year":"2013","unstructured":"Amin Milani Fard and Ali Mesbah. 2013. Feedback-directed exploration of web applications to derive test models. In Proceedings of the 24th International Symposium on Software Reliability Engineering, 278\u2013287."},{"issue":"4","key":"e_1_3_2_24_2","first-page":"228","article-title":"On the evolution of clusters of near-duplicate web pages","volume":"2","author":"Fetterly Dennis","year":"2004","unstructured":"Dennis Fetterly, Mark S. Manasse, and Marc Najork. 2004. On the evolution of clusters of near-duplicate web pages. Journal of Web Engineering 2, 4 (2004), 228\u2013246.","journal-title":"Journal of Web Engineering"},{"key":"e_1_3_2_25_2","first-page":"536","volume-title":"Proceedings of the 2nd IEEE International Conference on Multimedia Computing and Systems","volume":"2","author":"Jiri Fridrich","year":"1999","unstructured":"Jiri Fridrich. 1999. Robust bit extraction from images. In Proceedings of the 2nd IEEE International Conference on Multimedia Computing and Systems, Vol. 2, 536\u2013540."},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.552"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.72"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.5555\/3086952"},{"key":"e_1_3_2_29_2","unstructured":"Ian J. Goodfellow Jonathon Shlens and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv:1412.6572. Retrieved from https:\/\/arxiv.org\/abs\/1412.6572"},{"key":"e_1_3_2_30_2","volume-title":"Flask Web Development: Developing Web Applications with Python","author":"Grinberg Miguel","year":"2014","unstructured":"Miguel Grinberg. 2014. Flask Web Development: Developing Web Applications with Python. O\u2019Reilly Media, Inc."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00042"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.100"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","first-page":"284","DOI":"10.1145\/1148170.1148222","volume-title":"Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Henzinger Monika Rauch","year":"2006","unstructured":"Monika Rauch Henzinger. 2006. Finding near-duplicate web pages: A large-scale evaluation of algorithms. In Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 284\u2013291."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_37_2","first-page":"594","volume-title":"Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Hou Zhenyu","year":"2022","unstructured":"Zhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong, Hongxia Yang, Chunjie Wang, and Jie Tang. 2022. GraphMAE: Self-supervised masked graph autoencoders. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 594\u2013604."},{"key":"e_1_3_2_38_2","unstructured":"dockerhub. Docker Images of App Subjects on Docker Hub. 2025. Retrieved from https:\/\/hub.docker.com\/u\/dockercontainervm"},{"key":"e_1_3_2_39_2","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_2_40_2","first-page":"1188","volume-title":"Proceedings of the 31th International Conference on Machine Learning","author":"Le Quoc V.","year":"2014","unstructured":"Quoc V. Le and Tom\u00e1s Mikolov. 2014. Distributed representations of sentences and documents. In Proceedings of the 31th International Conference on Machine Learning, 1188\u20131196."},{"issue":"8","key":"e_1_3_2_41_2","first-page":"707","article-title":"Binary codes capable of correcting deletions, insertions, and reversals","volume":"10","author":"Levenshtein Vladimir I.","year":"1966","unstructured":"Vladimir I. Levenshtein. 1966. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady 10, 8 (1966), 707\u2013710.","journal-title":"Soviet Physics Doklady"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.aiopen.2022.03.001"},{"key":"e_1_3_2_43_2","first-page":"167","volume-title":"Proceedings of the 34th Pacific Asia Conference on Language, Information and Computation","author":"Louvan Samuel","year":"2020","unstructured":"Samuel Louvan and Bernardo Magnini. 2020. Simple is better! Lightweight data augmentation for low resource slot filling and intent classification. In Proceedings of the 34th Pacific Asia Conference on Language, Information and Computation, 167\u2013177."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.5555\/850924.851523"},{"key":"e_1_3_2_45_2","first-page":"141","volume-title":"Proceedings of the 16th International Conference on World Wide Web","author":"Singh Manku Gurmeet","year":"2007","unstructured":"Gurmeet Singh Manku, Arvind Jain, and Anish Das Sarma. 2007. Detecting near-duplicates for web crawling. In Proceedings of the 16th International Conference on World Wide Web, 141\u2013150."},{"key":"e_1_3_2_46_2","first-page":"121","volume-title":"Proceedings of the 1st International Conference on Software Testing, Verification, and Validation","author":"Marchetto Alessandro","year":"2008","unstructured":"Alessandro Marchetto, Paolo Tonella, and Filippo Ricca. 2008. State-based testing of ajax web applications. In Proceedings of the 1st International Conference on Software Testing, Verification, and Validation, 121\u2013130."},{"key":"e_1_3_2_47_2","first-page":"369","volume-title":"Proceedings of the 1987 Conference on the Theory and Applications of Cryptographic Techniques","author":"Merkle Ralph C.","year":"1987","unstructured":"Ralph C. Merkle. 1987. A digital signature based on a conventional encryption function. In Proceedings of the 1987 Conference on the Theory and Applications of Cryptographic Techniques, 369\u2013378."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/2109205.2109208"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2011.28"},{"key":"e_1_3_2_50_2","first-page":"3111","volume-title":"Proceedings of the 27th International Conference on Neural Information Processing Systems","author":"Mikolov Tom\u00e1s","year":"2013","unstructured":"Tom\u00e1s Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Proceedings of the 27th International Conference on Neural Information Processing Systems, 3111\u20133119."},{"key":"e_1_3_2_51_2","unstructured":"Android Monkey. 2025. Retrieved from https:\/\/developer.android.com\/studio\/test\/monkey"},{"key":"e_1_3_2_52_2","first-page":"7","volume-title":"Proceedings of the 4th Cybercrime and Trustworthy Computing Workshop","author":"Oliver Jonathan","year":"2013","unstructured":"Jonathan Oliver, Chun Cheng, and Yanggui Chen. 2013. TLSH\u2014A locality sensitive hash. In Proceedings of the 4th Cybercrime and Trustworthy Computing Workshop, 7\u201313."},{"key":"e_1_3_2_53_2","unstructured":"FragGen Replication Package. 2022. Retrieved from https:\/\/zenodo.org\/records\/5981993"},{"key":"e_1_3_2_54_2","unstructured":"Replication Package and Evaluation Results of Judge. 2025. Retrieved from https:\/\/zenodo.org\/records\/15223709"},{"key":"e_1_3_2_55_2","unstructured":"Pagekit. 2023. Retrieved from https:\/\/github.com\/pagekit\/pagekit"},{"key":"e_1_3_2_56_2","first-page":"8024","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, 8024\u20138035."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/2699485"},{"key":"e_1_3_2_58_2","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.OpenAI. Retrieved from https:\/\/cdn.openai.com\/research-covers\/language-unsupervised\/language_understanding_paper.pdf"},{"issue":"8","key":"e_1_3_2_59_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510457.3513027"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00083"},{"key":"e_1_3_2_62_2","unstructured":"React. 2025. Retrieved from https:\/\/react.dev\/"},{"key":"e_1_3_2_63_2","unstructured":"HTML Element Reference. 2025. Retrieved from https:\/\/www.w3schools.com\/TAGS\/default.asp"},{"key":"e_1_3_2_64_2","unstructured":"NDStudy GitHub Repository. 2025. Retrieved from https:\/\/github.com\/NDStudyICSE2019\/NDStudy"},{"key":"e_1_3_2_65_2","unstructured":"Retroboard. 2025. Retrieved from https:\/\/github.com\/antoinejaussoin\/retro-board"},{"key":"e_1_3_2_66_2","unstructured":"Dinghan Shen Mingzhi Zheng Yelong Shen Yanru Qu and Weizhu Chen. 2020. A simple but tough-to-beat data augmentation approach for natural language understanding and generation. arXiv:2009.13818. Retrieved from https:\/\/arxiv.org\/abs\/2009.13818"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2022.111512"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2009.03.002"},{"key":"e_1_3_2_69_2","unstructured":"Splittypie. 2025. Retrieved from https:\/\/github.com\/tsubik\/splittypie"},{"key":"e_1_3_2_70_2","unstructured":"W3C HTML Living Standard. 2025. Retrieved from https:\/\/www.w3.org\/TR\/html5\/"},{"key":"e_1_3_2_71_2","first-page":"132","volume-title":"Proceedings of the 16th International Conference on Web Engineering","author":"Stocco Andrea","year":"2016","unstructured":"Andrea Stocco, Maurizio Leotta, Filippo Ricca, and Paolo Tonella. 2016. Clustering-aided page object generation for web testing. In Proceedings of the 16th International Conference on Web Engineering, 132\u2013151."},{"key":"e_1_3_2_72_2","unstructured":"Andrea Stocco Alexandra Willi Luigi Libero Lucio Starace Matteo Biagiola and Paolo Tonella. 2023. Neural embeddings for web testing. arXiv:2306.07400. Retrieved from https:\/\/arxiv.org\/abs\/2306.07400"},{"key":"e_1_3_2_73_2","first-page":"245","volume-title":"Proceedings of the 11th European Software Engineering Conference and Symposium on the Foundations of Software Engineering","author":"Su Ting","year":"2017","unstructured":"Ting Su, Guozhu Meng, Yuting Chen, Ke Wu, Weiming Yang, Yao Yao, Geguang Pu, Yang Liu, and Zhendong Su. 2017. Guided, stochastic model-based GUI testing of android apps. In Proceedings of the 11th European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 245\u2013256."},{"key":"e_1_3_2_74_2","first-page":"390","volume-title":"Proceedings of the 3rd International Conference on Computer Vision","author":"Swain Michael J.","year":"1990","unstructured":"Michael J. Swain and Dana H. Ballard. 1990. Indexing via color histograms. In Proceedings of the 3rd International Conference on Computer Vision, 390\u2013393."},{"key":"e_1_3_2_75_2","unstructured":"Phoenix Trello. 2025. Retrieved from https:\/\/github.com\/bigardone\/phoenix-trello"},{"key":"e_1_3_2_76_2","unstructured":"A\u00e4ron van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv:1807.03748. Retrieved from https:\/\/arxiv.org\/abs\/1807.03748"},{"key":"e_1_3_2_77_2","first-page":"5998","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems, 5998\u20136008."},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3468554"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"e_1_3_2_81_2","unstructured":"Web Content Accessibility Guidelines (WCAG). 2024. Retrieved from https:\/\/www.w3.org\/TR\/WCAG21\/"},{"key":"e_1_3_2_82_2","unstructured":"W3C Official Website. 2025. Retrieved from https:\/\/www.w3.org\/"},{"key":"e_1_3_2_83_2","unstructured":"Alexa Top Websites. 2019. Retrieved from https:\/\/www.expireddomains.net\/alexa-top-websites"},{"key":"e_1_3_2_84_2","first-page":"6381","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing","author":"Wei Jason W.","year":"2019","unstructured":"Jason W. Wei and Kai Zou. 2019. EDA: Easy data augmentation techniques for boosting performance on text classification tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 6381\u20136387."},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00393"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.393"},{"key":"e_1_3_2_87_2","first-page":"186","volume-title":"Proceedings of the 42nd International Conference on Software Engineering","author":"Yandrapally Rahulkrishna","year":"2020","unstructured":"Rahulkrishna Yandrapally, Andrea Stocco, and Ali Mesbah. 2020. Near-duplicate detection in web app model inference. In Proceedings of the 42nd International Conference on Software Engineering, 186\u2013197."},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2022.3171295"},{"key":"e_1_3_2_89_2","first-page":"167","volume-title":"Proceedings of the 2nd International Conference on Intelligent Information Hiding and Multimedia Signal Processing","author":"Yang Bian","year":"2006","unstructured":"Bian Yang, Fan Gu, and Xiamu Niu. 2006. Block mean value based image perceptual hashing. In Proceedings of the 2nd International Conference on Intelligent Information Hiding and Multimedia Signal Processing, 167\u2013172."},{"key":"e_1_3_2_90_2","doi-asserted-by":"publisher","DOI":"10.1145\/383745.383748"},{"key":"e_1_3_2_91_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3414672"},{"key":"e_1_3_2_92_2","doi-asserted-by":"publisher","DOI":"10.1145\/3674728"},{"key":"e_1_3_2_93_2","first-page":"1601","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing","author":"Zhang Yan","year":"2020","unstructured":"Yan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim, and Lidong Bing. 2020. An unsupervised sentence embedding method by mutual information maximization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 1601\u20131610."},{"key":"e_1_3_2_94_2","doi-asserted-by":"publisher","DOI":"10.1145\/3533767.3543292"},{"key":"e_1_3_2_95_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00048"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3736162","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,13]],"date-time":"2026-02-13T14:37:55Z","timestamp":1770993475000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3736162"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,13]]},"references-count":94,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3736162"],"URL":"https:\/\/doi.org\/10.1145\/3736162","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,13]]},"assertion":[{"value":"2024-08-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}