{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:04:44Z","timestamp":1782846284063,"version":"3.54.5"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>In recent years, Extended Reality (XR) has emerged as a transformative technology, offering users immersive and interactive experiences across diversified virtual or virtual-real environments. Users can interact with XR applications (apps) through interactable GUI elements (IGEs) on the stereoscopic three-dimensional (3D) graphical user interface (GUI). IGE constitutes the fundamental element of XR GUI, embodying rich semantic information. The accurate recognition and precise understanding of these IGEs is instrumental, serving as the foundation of GUI grounding, which can facilitates downstream tasks, including automated XR testing. A straightforward XR test generator can interact randomly within the app\u2019s 3D environment, making it trapped in uninteractable space and resulting in an ineffective and inefficient testing process. In contrast, a more intelligent test generator, informed by the accurate locations and semantics of IGEs, can make wiser decisions on interaction targets and orders, forming test sequences that cover more functionalities faster. The most recent IGE detection approaches in SE are designed for 2D mobile apps and typically train a supervised object detection model based on a large-scale manually-labeled GUI dataset, usually with a pre-defined set of clickable GUI element categories like buttons and spinners. Such approaches can hardly be applied to IGE detection in XR apps, due to a multitude of challenges including complexities posed by open-vocabulary and heterogeneous IGE categories, intricacies of context-sensitive interactability, and the necessities of precise spatial perception and visual-semantic alignment for accurate IGE detection results. Thus, it is necessary to embark on the IGE research tailored to XR apps.<\/jats:p>\n                  <jats:p>\n                    In this paper, we propose the first zero-shot context-sensitive interactable GUI element detection framework for Extended Reality apps, named Orienter. Rather than relying on generic visual grounding which fails in 3D environments, Orienter introduces a structured workflow tailored to XR constraints. It first synthesizes XR-specific semantic contexts (e.g., global interaction paradigms and 3D spatial properties) before performing detection. To overcome severe spatial hallucinations inherent in LMMs, the detection process is iterated within an XR-constrained reflection loop. Specifically, Orienter contains three components, including (1)\n                    <jats:italic toggle=\"yes\">Semantic context comprehension<\/jats:italic>\n                    for capturing the apps\u2019 GUI context, (2)\n                    <jats:italic toggle=\"yes\">Reflection-directed IGE candidate detection<\/jats:italic>\n                    for identifying and localizing valid GUI elements based on multi-perspective description guided IGE detection, as well as feedback-directed reflection, and (3)\n                    <jats:italic toggle=\"yes\">Context-sensitive interactability classification<\/jats:italic>\n                    which integrates semantic contexts for interactability prediction. To evaluate our approach and facilitate follow-up research, we construct the first benchmark dataset which contains 1,552 images from 100 industrial-setting apps on Steam, with 4,470 interactable annotations across 766 semantics categories. Extensive experiments on the dataset demonstrate that Orienter is more effective than the state-of-the-art GUI element detection approaches, including general or GUI-automation-targeted vision language models, and deep learning based models, surpassing their F1 Score by at least 34.9% and 20.1\u00d7 in distinguishing the interactibility and semantics of the IGEs, respectively. Orienter is beneficial for boosting the performance of automatic testing by isolating the interactable action space from the whole space, regardless of the testing strategies employed. Experiments demonstrate that Orienter-guided testing covers 103.1% more IGEs with 125.7% more effective interactions than testing without action space isolation.\n                  <\/jats:p>","DOI":"10.1145\/3808134","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"2858-2881","source":"Crossref","is-referenced-by-count":0,"title":["Look Before You Leap: Context-Sensitive GUI Grounding for Boosting Automated Extended Reality (XR) Testing"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6323-1402","authenticated-orcid":false,"given":"Shuqing","family":"Li","sequence":"first","affiliation":[{"name":"Chinese University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-5995-4040","authenticated-orcid":false,"given":"Binchang","family":"Li","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8147-8126","authenticated-orcid":false,"given":"Yepang","family":"Liu","sequence":"additional","affiliation":[{"name":"Southern University of Science and Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4774-2434","authenticated-orcid":false,"given":"Cuiyun","family":"Gao","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-5978-8888","authenticated-orcid":false,"given":"Jianping","family":"Zhang","sequence":"additional","affiliation":[{"name":"Chinese University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3508-7172","authenticated-orcid":false,"given":"Shing-Chi","family":"Cheung","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3666-5798","authenticated-orcid":false,"given":"Michael R.","family":"Lyu","sequence":"additional","affiliation":[{"name":"Chinese University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2021. [BUG] Crash on game start and left eye issue in home. https:\/\/github.com\/ValveSoftware\/SteamVR-for- Linux\/issues\/453."},{"key":"e_1_2_1_2_1","unstructured":"2021. FILM XR. https:\/\/vrfilmreview.ru\/."},{"key":"e_1_2_1_3_1","unstructured":"2021. VirtualSkill -Virtual Reality Training. https:\/\/virtualskill.com\/."},{"key":"e_1_2_1_4_1","unstructured":"2021. XR Games. https:\/\/www.xrgames.io\/."},{"key":"e_1_2_1_5_1","unstructured":"2023. COCO API. https:\/\/github.com\/cocodataset\/cocoapi."},{"key":"e_1_2_1_6_1","unstructured":"2023. Munch VR on Steam. https:\/\/store.steampowered.com\/app\/549000\/Munch_VR\/."},{"key":"e_1_2_1_7_1","unstructured":"2023. VR Content on Steam App Store. https:\/\/store.steampowered.com\/search\/?vrsupport=401."},{"key":"e_1_2_1_8_1","unstructured":"2023. VR The Diner Duo on Steam. https:\/\/store.steampowered.com\/app\/530120\/VR_The_Diner_Duo\/."},{"key":"e_1_2_1_9_1","unstructured":"2024. Baseball Kings VR. https:\/\/store.steampowered.com\/app\/802300\/Baseball_Kings_VR\/."},{"key":"e_1_2_1_10_1","unstructured":"2024. GPT-4o Release Page. https:\/\/openai.com\/index\/hello-gpt-4o\/."},{"key":"e_1_2_1_11_1","unstructured":"2024. Job Simulator on Steam. https:\/\/store.steampowered.com\/app\/448280\/Job_Simulator\/."},{"key":"e_1_2_1_12_1","unstructured":"2024. LabTrainingVR: Biosafety Cabinet Edition. https:\/\/store.steampowered.com\/app\/1337060\/LabTrainingVR_ Biosafety_Cabinet_Edition\/."},{"key":"e_1_2_1_13_1","unstructured":"2024. Potioneer: The VR Gardening Simulator. https:\/\/store.steampowered.com\/app\/544410\/Potioneer_The_VR_ Gardening_Simulator\/."},{"key":"e_1_2_1_14_1","unstructured":"2024. Storm VR on Steam. https:\/\/store.steampowered.com\/app\/457380\/Storm_VR\/."},{"key":"e_1_2_1_15_1","unstructured":"2024. Ultimate Fishing Simulator VR. https:\/\/store.steampowered.com\/app\/1024010\/Ultimate_Fishing_Simulator_VR\/."},{"key":"e_1_2_1_16_1","unstructured":"2025. Claude-4.5 Sonnet Release Page. https:\/\/www.anthropic.com\/news\/claude-sonnet-4-5."},{"key":"e_1_2_1_17_1","unstructured":"2025. Crazy graphic glitch. What is going on. https:\/\/www.reddit.com\/r\/hoggit\/comments\/1i1dt6p\/crazy_graphic_ glitch_what_is_going_on\/."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/SVR.2019.00044"},{"key":"e_1_2_1_19_1","unstructured":"Shuai Bai Yuxuan Cai Ruizhe Chen Keqin Chen Xionghui Chen Zesen Cheng Lianghao Deng Wei Ding Chang Gao Chunjiang Ge Wenbin Ge Zhifang Guo Qidong Huang Jie Huang Fei Huang Binyuan Hui Shutong Jiang Zhaohai Li Mingsheng Li Mei Li Kaixin Li Zicheng Lin Junyang Lin Xuejing Liu Jiawei Liu Chenglong Liu Yang Liu Dayiheng Liu Shixuan Liu Dunjie Lu Ruilin Luo Chenxu Lv Rui Men Lingchen Meng Xuancheng Ren Xingzhang Ren Sibo Song Yuchong Sun Jun Tang Jianhong Tu Jianqiang Wan Peng Wang Pengfei Wang Qiuyue Wang Yuxuan Wang Tianbao Xie Yiheng Xu Haiyang Xu Jin Xu Zhibo Yang Mingkun Yang Jianxin Yang An Yang Bowen Yu Fei Zhang Hang Zhang Xi Zhang Bo Zheng Humen Zhong Jingren Zhou Fan Zhou Jing Zhou Yuanzhi Zhu and Ke Zhu. 2025. Qwen3-VL Technical Report. CoRR abs\/2511.21631 (2025). https:\/\/doi.org\/10.48550\/ARXIV.2511.21631 arXiv:2511.21631 10.48550\/ARXIV.2511.21631"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2015.220"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10055-016-0284-x"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3180155.3180240"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409691"},{"key":"e_1_2_1_24_1","volume-title":"Qualitative inquiry and research design: Choosing among five approaches","author":"Creswell John W","unstructured":"John W Creswell and Cheryl N Poth. 2016. Qualitative inquiry and research design: Choosing among five approaches. Sage publications."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1002\/stvr.1863"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/1753326.1753554"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3563213"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2505.07062"},{"key":"e_1_2_1_29_1","unstructured":"Glenn Jocher Ayush Chaurasia and Jing Qiu. 2023. YOLO by Ultralytics. https:\/\/github.com\/ultralytics\/ultralytics"},{"key":"e_1_2_1_30_1","article-title":"Augmented and virtual reality in surgery-the digital surgical environment: applications, limitations and legal pitfalls","volume":"4","author":"Khor Wee Sim","year":"2016","unstructured":"Wee Sim Khor, Benjamin Baker, Kavit Amin, Adrian Chan, Ketan Patel, and Jason Wong. 2016. Augmented and virtual reality in surgery-the digital surgical environment: applications, limitations and legal pitfalls. Annals of translational medicine 4, 23 (2016).","journal-title":"Annals of translational medicine"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3660803"},{"key":"e_1_2_1_32_1","volume-title":"Lyu","author":"Li Shuqing","year":"2024","unstructured":"Shuqing Li, Binchang Li, Cuiyun Gao, and Michael R. Lyu. 2024. An Interaction Simulation and Automated Testing Framework for Spatial Computing Extended Reality Applications. (2024)."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.06783"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2018.2844788"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2015.32"},{"key":"e_1_2_1_37_1","unstructured":"OpenAI. 2023. GPT-4 Technical Report. CoRR abs\/2303.08774 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2303.08774 arXiv:2303.08774 10.48550\/ARXIV.2303.08774"},{"key":"e_1_2_1_38_1","unstructured":"OpenAI. 2025. GPT-5 Release Page. https:\/\/openai.com\/index\/introducing-gpt-5."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3395363.3397354"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3556966"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3561160"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-SEIP55303.2022.9793948"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598134"},{"key":"e_1_2_1_45_1","doi-asserted-by":"crossref","unstructured":"Yunhang Shen Chaoyou Fu Peixian Chen Mengdan Zhang Ke Li Xing Sun Yunsheng Wu Shaohui Lin and Rongrong Ji. 2024. Aligning and Prompting Everything All at Once for Universal Visual Perception. CVPR.","DOI":"10.1109\/CVPR52733.2024.01253"},{"key":"e_1_2_1_46_1","unstructured":"Statista. 2022. Report of Active Virtual Reality Users Worldwide. https:\/\/www.statista.com\/statistics\/426469\/activevirtual-reality-users-worldwide\/."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2507.06261"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01481"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3613868"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2508.18265"},{"key":"e_1_2_1_51_1","volume-title":"CogVLM: Visual Expert for Pretrained Language Models. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024","author":"Wang Weihan","year":"2024","unstructured":"Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Keqin Chen, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, and Jie Tang. 2024. CogVLM: Visual Expert for Pretrained Language Models. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 -15, 2024, Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.). http:\/\/papers.nips.cc\/paper_files\/paper\/2024\/hash\/dc06d4d2792265fb5454a6092bfd5c6a-Abstract-Conference.html"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510454.3516870"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE56229.2023.00197"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3293882.3330551"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3417940"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.findings-acl.1152"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3473935"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/1622176.1622213"},{"key":"e_1_2_1_60_1","unstructured":"Zhipu AI. 2024. Zhipu AI Open Platform. https:\/\/open.bigmodel.cn\/."},{"key":"e_1_2_1_61_1","volume-title":"Probabilistic two-stage detection. CoRR abs\/2103.07461","author":"Zhou Xingyi","year":"2021","unstructured":"Xingyi Zhou, Vladlen Koltun, and Philipp Kr\u00e4henb\u00fchl. 2021. Probabilistic two-stage detection. CoRR abs\/2103.07461 (2021). arXiv:2103.07461"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3808134","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:10:52Z","timestamp":1782843052000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3808134"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":61,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3808134"],"URL":"https:\/\/doi.org\/10.1145\/3808134","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}