{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T15:45:03Z","timestamp":1781797503778,"version":"3.54.5"},"reference-count":81,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,9,18]],"date-time":"2025-09-18T00:00:00Z","timestamp":1758153600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,9,18]],"date-time":"2025-09-18T00:00:00Z","timestamp":1758153600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Autom Softw Eng"],"published-print":{"date-parts":[[2026,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    As the Virtual Reality (VR) industry expands, the need for automated GUI testing is growing rapidly. Large Language Models (LLMs), capable of retaining information long-term and analyzing both visual and textual data, are emerging as a potential key to deciphering the complexities of VR\u2019s evolving user interfaces. In this paper, we conduct a case study to investigate the capability of using LLMs, particularly GPT-4o, for field of view (FOV) analysis in VR exploration testing. Specifically, we validate that LLMs can identify test entities in FOVs and that prompt engineering can effectively enhance the accuracy of test entity identification from\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\varvec{41.67\\%}$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    to\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\varvec{71.30\\%}$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    . Our study also shows that LLMs can accurately describe identified entities\u2019 features with at least a\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\varvec{90\\%}$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    accuracy rate. We further find out that the core features that effectively represent an entity are color, placement, and shape. Furthermore, the combination of the three features can especially be used to improve the accuracy of determining identical entities in multiple FOVs with the highest F1-score of\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\varvec{0.70}$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    . Additionally, our study demonstrates that LLMs are capable of scene recognition and spatial understanding in VR with precisely designed structured prompts. Finally, we find that LLMs fail to label the identified test entities, and we discuss potential solutions as future research directions.\n                  <\/jats:p>","DOI":"10.1007\/s10515-025-00535-3","type":"journal-article","created":{"date-parts":[[2025,9,18]],"date-time":"2025-09-18T11:11:53Z","timestamp":1758193913000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Harnessing large language models for virtual reality exploration testing: a case study"],"prefix":"10.1007","volume":"33","author":[{"given":"Zhenyu","family":"Qi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haotang","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hao","family":"Qin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kebin","family":"Peng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sen","family":"He","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xue","family":"Qin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,9,18]]},"reference":[{"key":"535_CR1","unstructured":"3DGEp: Understanding the View Matrix. Accessed 05 Jan 2025 (2023) https:\/\/www.3dgep.com\/understanding-the-view-matrix\/"},{"key":"535_CR2","doi-asserted-by":"crossref","unstructured":"Adelson, E.H.: On seeing stuff: the perception of materials by humans and machines. In: Human Vision and Electronic Imaging VI, vol. 4299, pp. 1\u201312. SPIE (2001)","DOI":"10.1117\/12.429489"},{"key":"535_CR3","unstructured":"Armi, L., Fekri-Ershad, S.: Texture image analysis and texture classification methods-a review. arXiv:1904.06554 (2019)"},{"key":"535_CR4","unstructured":"Becker, E., Soatto, S.: Cycles of Thought: Measuring LLM confidence through stable explanations (2024). https:\/\/arxiv.org\/abs\/2406.03441"},{"key":"535_CR5","doi-asserted-by":"publisher","unstructured":"Bierbaum, A., Hartling, P., Cruz-Neira, C.: Automated testing of virtual reality application interfaces. In: Proceedings of the Workshop on Virtual Environments 2003. EGVE \u201903, pp. 107\u2013114. Association for Computing Machinery, New York, NY, USA (2003). https:\/\/doi.org\/10.1145\/769953.769966","DOI":"10.1145\/769953.769966"},{"key":"535_CR6","doi-asserted-by":"publisher","unstructured":"Boy, G., Mazzone, R., Conroy, M.: The virtual camera concept: A third person view, pp. 551\u2013559 (2010). https:\/\/doi.org\/10.1201\/EBK1439834916-c55","DOI":"10.1201\/EBK1439834916-c55"},{"key":"535_CR7","doi-asserted-by":"publisher","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: Computer Vision \u2013 ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part I, pp. 213\u2013229. Springer, Berlin, Heidelberg (2020). https:\/\/doi.org\/10.1007\/978-3-030-58452-8_13","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"535_CR8","doi-asserted-by":"publisher","unstructured":"Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., Ye, W., Zhang, Y., Chang, Y., Yu, P.S., Yang, Q., Xie, X.: A survey on evaluation of large language models. ACM Trans. Intell. Syst. Technol. 15(3) (2024). https:\/\/doi.org\/10.1145\/3641289","DOI":"10.1145\/3641289"},{"key":"535_CR9","volume-title":"Constructing Grounded Theory: A Practical Guide Through Qualitative Analysis","author":"K Charmaz","year":"2006","unstructured":"Charmaz, K.: Constructing Grounded Theory: A Practical Guide Through Qualitative Analysis. Sage Publications, Thousand Oaks, CA (2006)"},{"key":"535_CR10","doi-asserted-by":"publisher","unstructured":"Chen, J., Xie, M., Xing, Z., Chen, C., Xu, X., Zhu, L., Li, G.: Object detection for graphical user interface: Old fashioned or deep learning or a combination? In: Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. ESEC\/FSE 2020, pp. 1202\u20131214. Association for Computing Machinery, New York, NY, USA (2020). https:\/\/doi.org\/10.1145\/3368089.3409691","DOI":"10.1145\/3368089.3409691"},{"key":"535_CR11","unstructured":"coinse: Autonomous Large Language Model Agents Enabling Intent-Driven Mobile GUI Testing. https:\/\/github.com\/coinse\/droidagent. Accessed 08 Aug 2024 (2024)"},{"key":"535_CR12","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1007\/s11263-006-8711-1","volume":"72","author":"D Cremers","year":"2007","unstructured":"Cremers, D., Rousson, M., Deriche, R.: A review of statistical approaches to level set segmentation: integrating color, texture, motion and shape. Int. J. Comput. Vision 72, 195\u2013215 (2007)","journal-title":"Int. J. Comput. Vision"},{"key":"535_CR13","doi-asserted-by":"publisher","unstructured":"Devlin, J., Chang, M.-W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171\u20134186. Association for Computational Linguistics, Minneapolis, Minnesota (2019). https:\/\/doi.org\/10.18653\/v1\/N19-1423","DOI":"10.18653\/v1\/N19-1423"},{"key":"535_CR14","unstructured":"Diplaros, A., Gevers, T., Patras, I., et al.: Color-shape context for object recognition. In: IEEE Workshop on Color and Photometric Methods in Computer Vision, pp. 1\u20138. Citeseer (2003)"},{"key":"535_CR15","doi-asserted-by":"crossref","unstructured":"Duong, T.A., Duong, V.A., Stubberud, A.R.: Shape and color features for object recognition search. In: Handbook of Pattern Recognition and Computer Vision, pp. 101\u2013125. World Scientific, Singapore (2010)","DOI":"10.1142\/9789814273398_0005"},{"key":"535_CR16","doi-asserted-by":"publisher","unstructured":"Ferdous, R., Kifetew, F., Prandi, D., Susi, A.: Towards agent-based testing of 3d games using reinforcement learning. In: Proceedings of the 37th IEEE\/ACM International Conference on Automated Software Engineering. ASE \u201922. Association for Computing Machinery, New York, NY, USA (2023). https:\/\/doi.org\/10.1145\/3551349.3560507","DOI":"10.1145\/3551349.3560507"},{"issue":"3","key":"535_CR17","doi-asserted-by":"publisher","first-page":"284","DOI":"10.1007\/s11263-009-0270-9","volume":"87","author":"V Ferrari","year":"2010","unstructured":"Ferrari, V., Jurie, F., Schmid, C.: From images to shape models for object detection. Int. J. Comput. Vision 87(3), 284\u2013303 (2010)","journal-title":"Int. J. Comput. Vision"},{"key":"535_CR18","doi-asserted-by":"publisher","unstructured":"Fischer, F., Ikkala, A., Klar, M., Fleig, A., Bachinski, M., Murray-Smith, R., H\u00e4m\u00e4l\u00e4inen, P., Oulasvirta, A., M\u00fcller, J.: Sim2vr: Towards automated biomechanical testing in vr. In: Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. UIST \u201924. Association for Computing Machinery, New York, NY, USA (2024). https:\/\/doi.org\/10.1145\/3654777.3676452","DOI":"10.1145\/3654777.3676452"},{"key":"535_CR19","doi-asserted-by":"publisher","unstructured":"Gadille, M., Corvasce, C., Impedovo, M.: Material and socio-cognitive effects of immersive virtual reality in a french secondary school: Conditions for innovation. Education Sci. 13(3) (2023). https:\/\/doi.org\/10.3390\/educsci13030251","DOI":"10.3390\/educsci13030251"},{"key":"535_CR20","doi-asserted-by":"publisher","unstructured":"Gamal, A., Emad, R., Mohamed, T., Mohamed, O., Hamdy, A., Ali, S.: Owl eye: An ai-driven visual testing tool. In: 2023 5th Novel Intelligent and Leading Emerging Sciences Conference (NILES), pp. 312\u2013315 (2023). https:\/\/doi.org\/10.1109\/NILES59815.2023.10296575","DOI":"10.1109\/NILES59815.2023.10296575"},{"key":"535_CR21","doi-asserted-by":"publisher","unstructured":"Gao, W., Du, K., Luo, Y., Shi, W., Yu, C., Shi, Y.: Easyask: An in-app contextual tutorial search assistant for older adults with voice and touch inputs. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 8(3) (2024). https:\/\/doi.org\/10.1145\/3678516","DOI":"10.1145\/3678516"},{"key":"535_CR22","unstructured":"Ge, J., Luo, H., Qian, S., Gan, Y., Fu, J., Zhang, S.: Chain of thought prompt tuning in vision language models (2023). https:\/\/arxiv.org\/abs\/2304.07919"},{"key":"535_CR23","doi-asserted-by":"crossref","unstructured":"Ge, Y., Xiao, Y., Xu, Z., Wang, X., Itti, L.: Contributions of shape, texture, and color in visual recognition. In: European Conference on Computer Vision, pp. 369\u2013386. Springer (2022)","DOI":"10.1007\/978-3-031-19775-8_22"},{"issue":"1","key":"535_CR24","doi-asserted-by":"publisher","first-page":"102","DOI":"10.1109\/83.817602","volume":"9","author":"T Gevers","year":"2000","unstructured":"Gevers, T., Smeulders, A.W.: Pictoseek: Combining color and shape invariant features for image retrieval. IEEE Trans. Image Process. 9(1), 102\u2013119 (2000)","journal-title":"IEEE Trans. Image Process."},{"issue":"9","key":"535_CR25","doi-asserted-by":"publisher","first-page":"10","DOI":"10.1167\/10.9.10","volume":"10","author":"M Giesel","year":"2010","unstructured":"Giesel, M., Gegenfurtner, K.R.: Color appearance of real objects varying in material, hue, and shape. J. Vis. 10(9), 10\u201310 (2010)","journal-title":"J. Vis."},{"key":"535_CR26","doi-asserted-by":"publisher","unstructured":"Gorisse, G., Christmann, O., Amato, E.A., Richir, S.: First- and third-person perspectives in immersive virtual environments: Presence and performance analysis of embodied users. Front. Robotics AI Volume 4 - 2017 (2017). https:\/\/doi.org\/10.3389\/frobt.2017.00033","DOI":"10.3389\/frobt.2017.00033"},{"key":"535_CR27","doi-asserted-by":"publisher","unstructured":"Harms, P.: Automated usability evaluation of virtual reality applications. ACM Trans. Comput.-Hum. Interact. 26(3) (2019). https:\/\/doi.org\/10.1145\/3301423","DOI":"10.1145\/3301423"},{"key":"535_CR28","unstructured":"Hertzmann, A., Seitz, S.M.: Shape and materials by example: A photometric stereo approach. In: 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2003. Proceedings., vol. 1. IEEE (2003)"},{"key":"535_CR29","unstructured":"Hu, Y., Wang, X., Wang, Y., Zhang, Y., Guo, S., Chen, C., Wang, X., Zhou, Y.: AUITestAgent: Automatic requirements oriented gui function testing (2024). https:\/\/arxiv.org\/abs\/2407.09018"},{"key":"535_CR30","doi-asserted-by":"publisher","unstructured":"Huang, Y., Wang, J., Liu, Z., Wang, Y., Wang, S., Chen, C., Hu, Y., Wang, Q.: Crashtranslator: Automatically reproducing mobile application crashes directly from stack trace. In: Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. ICSE \u201924. Association for Computing Machinery, New York, NY, USA (2024). https:\/\/doi.org\/10.1145\/3597503.3623298","DOI":"10.1145\/3597503.3623298"},{"key":"535_CR31","unstructured":"Huang, T., Yu, C., Shi, W., Peng, Z., Yang, D., Sun, W., Shi, Y.: PromptRPA: Generating robotic process automation on smartphones from textual prompts (2024). https:\/\/arxiv.org\/abs\/2404.02475"},{"key":"535_CR32","unstructured":"Jocher, G., Qiu, J., Chaurasia, A.: Ultralytics YOLO (2023). https:\/\/github.com\/ultralytics\/ultralytics"},{"key":"535_CR33","doi-asserted-by":"crossref","unstructured":"Jumi, J., Zaenuddin, A., Mulyono, T.: Performance analysis of shape, color and texture features on tracking information face based on cbir. In: IOP Conference Series: Materials Science and Engineering, vol. 1108, p. 012031. IOP Publishing (2021)","DOI":"10.1088\/1757-899X\/1108\/1\/012031"},{"key":"535_CR34","doi-asserted-by":"publisher","first-page":"111","DOI":"10.4028\/www.scientific.net\/AMM.761.111","volume":"761","author":"A Kadir","year":"2015","unstructured":"Kadir, A., Aziz, K., Irianto, I.: A new object recognition approach using combination of texture, color and shape features. Appl. Mech. Mater. 761, 111\u2013115 (2015)","journal-title":"Appl. Mech. Mater."},{"key":"535_CR35","doi-asserted-by":"publisher","unstructured":"Karakaya, K., Yigitbas, E., Engels, G.: Automated ux evaluation for user-centered design of vr interfaces. In: Bernhaupt, R., Ardito, C., Sauer, S. (eds.) Human-Centered Software Engineering, pp. 140\u2013149. Springer, Cham (2022). https:\/\/doi.org\/10.1007\/978-3-031-14785-2_9","DOI":"10.1007\/978-3-031-14785-2_9"},{"issue":"3","key":"535_CR36","doi-asserted-by":"publisher","first-page":"329","DOI":"10.1007\/s11370-021-00349-8","volume":"14","author":"SH Kasaei","year":"2021","unstructured":"Kasaei, S.H., Ghorbani, M., Schilperoort, J., Rest, W.: Investigating the importance of shape features, color constancy, color spaces, and similarity measures in open-ended 3d object recognition. Intel. Serv. Robot. 14(3), 329\u2013344 (2021)","journal-title":"Intel. Serv. Robot."},{"key":"535_CR37","unstructured":"Kolesnikov, A., Dosovitskiy, A., Weissenborn, D., Heigold, G., Uszkoreit, J., Beyer, L., Minderer, M., Dehghani, M., Houlsby, N., Gelly, S., Unterthiner, T., Zhai, X.: An image is worth 16x16 words: Transformers for image recognition at scale. (2021)"},{"issue":"4","key":"535_CR38","doi-asserted-by":"publisher","first-page":"1222","DOI":"10.1016\/j.patcog.2006.09.017","volume":"40","author":"A Leone","year":"2007","unstructured":"Leone, A., Distante, C.: Shadow detection for moving objects based on texture analysis. Pattern Recogn. 40(4), 1222\u20131233 (2007)","journal-title":"Pattern Recogn."},{"issue":"2","key":"535_CR39","doi-asserted-by":"publisher","first-page":"161","DOI":"10.1080\/095400997116676","volume":"9","author":"WK Leow","year":"1997","unstructured":"Leow, W.K., Miikkulainen, R.: Visual schemas in neural networks for object recognition and scene analysis. Connect. Sci. 9(2), 161\u2013200 (1997). https:\/\/doi.org\/10.1080\/095400997116676","journal-title":"Connect. Sci."},{"key":"535_CR40","doi-asserted-by":"publisher","unstructured":"Li, S., Gao, C., Zhang, J., Zhang, Y., Liu, Y., Gu, J., Peng, Y., Lyu, M.R.: Less cybersickness, please: Demystifying and detecting stereoscopic visual inconsistencies in virtual reality apps. Proc. ACM Softw. Eng. 1(FSE) (2024). https:\/\/doi.org\/10.1145\/3660803","DOI":"10.1145\/3660803"},{"key":"535_CR41","unstructured":"Li, S., Li, B., Liu, Y., Gao, C., Zhang, J., Cheung, S.-C., Lyu, M.R.: Grounded GUI understanding for vision based spatial intelligent agent: Exemplified by virtual reality apps (2024). https:\/\/arxiv.org\/abs\/2409.10811"},{"key":"535_CR42","doi-asserted-by":"crossref","unstructured":"Li, F., Zhang, H., Liu, S., Guo, J., Ni, L.M., Zhang, L.: Dn-detr: Accelerate detr training by introducing query denoising. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 13619\u201313627 (2022)","DOI":"10.1109\/CVPR52688.2022.01325"},{"key":"535_CR43","doi-asserted-by":"publisher","first-page":"6893","DOI":"10.1109\/TIP.2022.3216771","volume":"31","author":"T Liang","year":"2022","unstructured":"Liang, T., Chu, X., Liu, Y., Wang, Y., Tang, Z., Chu, W., Chen, J., Ling, H.: Cbnet: A composite backbone network architecture for object detection. IEEE Trans. Image Process. 31, 6893\u20136906 (2022). https:\/\/doi.org\/10.1109\/TIP.2022.3216771","journal-title":"IEEE Trans. Image Process."},{"key":"535_CR44","doi-asserted-by":"publisher","unstructured":"Liu, Z., Chen, C., Wang, J., Che, X., Huang, Y., Hu, J., Wang, Q.: Fill in the blank: Context-aware automated text input generation for mobile gui testing. In: Proceedings of the 45th International Conference on Software Engineering. ICSE \u201923, pp. 1355\u20131367. IEEE Press, Piscataway, NJ, USA (2023). https:\/\/doi.org\/10.1109\/ICSE48619.2023.00119","DOI":"10.1109\/ICSE48619.2023.00119"},{"key":"535_CR45","doi-asserted-by":"publisher","unstructured":"Liu, Z., Chen, C., Wang, J., Chen, M., Wu, B., Che, X., Wang, D., Wang, Q.: Make llm a testing expert: Bringing human-like interaction to mobile gui testing via functionality-aware decisions. In: Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. ICSE \u201924. Association for Computing Machinery, New York, NY, USA (2024). https:\/\/doi.org\/10.1145\/3597503.3639180","DOI":"10.1145\/3597503.3639180"},{"key":"535_CR46","unstructured":"Liu, S., Li, F., Zhang, H., Yang, X., Qi, X., Su, H., Zhu, J., Zhang, L.: DAB-DETR: Dynamic anchor boxes are better queries for DETR. In: International Conference on Learning Representations (2022). https:\/\/openreview.net\/forum?id=oMI9PjOb9Jl"},{"key":"535_CR47","unstructured":"Meta: Getting started with quest 2. Accessed 17 Dec 2024 (2024). https:\/\/www.meta.com\/help\/quest\/articles\/getting-started\/getting-started-with-quest-2\/"},{"issue":"5","key":"535_CR48","doi-asserted-by":"publisher","first-page":"13","DOI":"10.1167\/8.5.13","volume":"8","author":"M Olkkonen","year":"2008","unstructured":"Olkkonen, M., Hansen, T., Gegenfurtner, K.R.: Color appearance of familiar objects: Effects of object shape, texture, and illumination changes. J. Vis. 8(5), 13\u201313 (2008)","journal-title":"J. Vis."},{"key":"535_CR49","unstructured":"OpenAI: ChatGPT: OpenAI Conversational AI. Accessed 07 Jan 2025 (2025). https:\/\/chat.openai.com\/"},{"key":"535_CR50","unstructured":"OpenAI: Hello GPT-4o. Accessed 12 Dec 2024 (2024). https:\/\/openai.com\/index\/hello-gpt-4o\/"},{"key":"535_CR51","unstructured":"OpenAI: OpenAI Documentation: Structured Outputs Guide. Accessed 02 Jan 2025 (2025). https:\/\/platform.openai.com\/docs\/guides\/structured-outputs"},{"key":"535_CR52","unstructured":"OpenAI: Vision Capabilities: Low or High Fidelity Image Understanding. https:\/\/platform.openai.com\/docs\/guides\/vision#low-or-high-fidelity-image-understanding. Accessed 05 Jan 2025 (2025)"},{"key":"535_CR53","unstructured":"OpenAI: Vision Guide: Limitations. Accessed 20 Dec 2024 (2024). https:\/\/platform.openai.com\/docs\/guides\/vision#limitations"},{"key":"535_CR54","doi-asserted-by":"publisher","unstructured":"Prasetya, I.S.W.B., Pastor\u00a0Ric\u00f3s, F., Kifetew, F.M., Prandi, D., Shirzadehhajimahmood, S., Vos, T.E.J., Paska, P., Hovorka, K., Ferdous, R., Susi, A., Davidson, J.: An agent-based approach to automated game testing: An experience report. In: Proceedings of the 13th International Workshop on Automating Test Case Design, Selection and Evaluation. A-TEST 2022, pp. 1\u20138. Association for Computing Machinery, New York, NY, USA (2022). https:\/\/doi.org\/10.1145\/3548659.3561305","DOI":"10.1145\/3548659.3561305"},{"key":"535_CR55","doi-asserted-by":"publisher","unstructured":"Qin, X., Hassan, F.: Dytrec: A dynamic testing recommendation tool for unity-based virtual reality software. In: Proceedings of the 37th IEEE\/ACM International Conference on Automated Software Engineering. ASE \u201922. Association for Computing Machinery, New York, NY, USA (2023). https:\/\/doi.org\/10.1145\/3551349.3560510","DOI":"10.1145\/3551349.3560510"},{"key":"535_CR56","doi-asserted-by":"publisher","unstructured":"Qin, X., Weaver, G.: Utilizing generative ai for vr exploration testing: A case study. In: Proceedings of the 39th IEEE\/ACM International Conference on Automated Software Engineering Workshops. ASEW \u201924, pp. 228\u2013232. Association for Computing Machinery, New York, NY, USA (2024). https:\/\/doi.org\/10.1145\/3691621.3694955","DOI":"10.1145\/3691621.3694955"},{"key":"535_CR57","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748\u20138763. PMLR (2021)"},{"key":"535_CR58","unstructured":"Radulescu, A., Opheusden, B., Callaway, F., Griffiths, T., Hillis, J.: From heuristic to optimal models in naturalistic visual search. ICLR 2020 Workshop on \u201cBridging AI & Cognitive Science\u201d (BAICS) (2020)"},{"key":"535_CR59","doi-asserted-by":"publisher","unstructured":"Redmon, J., Divvala, S., Girshick, R., Farhadi, A.: You only look once: Unified, real-time object detection. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 779\u2013788. IEEE Computer Society, Los Alamitos, CA, USA (2016).https:\/\/doi.org\/10.1109\/CVPR.2016.91","DOI":"10.1109\/CVPR.2016.91"},{"issue":"6","key":"535_CR60","doi-asserted-by":"publisher","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","volume":"39","author":"S Ren","year":"2017","unstructured":"Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 39(6), 1137\u20131149 (2017). https:\/\/doi.org\/10.1109\/TPAMI.2016.2577031","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"535_CR61","doi-asserted-by":"publisher","unstructured":"Rzig, D.E., Iqbal, N., Attisano, I., Qin, X., Hassan, F.: Virtual reality (vr) automated testing in the wild: A case study on unity-based vr applications. In: Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis. ISSTA 2023, pp. 1269\u20131281. Association for Computing Machinery, New York, NY, USA (2023). https:\/\/doi.org\/10.1145\/3597926.3598134","DOI":"10.1145\/3597926.3598134"},{"key":"535_CR62","unstructured":"Sharan, L.: The perception of material qualities in real-world images. PhD thesis, Massachusetts Institute of Technology (2009)"},{"key":"535_CR63","doi-asserted-by":"publisher","unstructured":"Taeb, M., Swearngin, A., Schoop, E., Cheng, R., Jiang, Y., Nichols, J.: Axnav: Replaying accessibility tests from natural language. In: Proceedings of the CHI Conference on Human Factors in Computing Systems. CHI \u201824, pp. 1\u201316. ACM, New York, NY, USA (2024). https:\/\/doi.org\/10.1145\/3613904.3642777","DOI":"10.1145\/3613904.3642777"},{"key":"535_CR64","doi-asserted-by":"publisher","unstructured":"Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., Manning, C.: Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In: Bouamor, H., Pino, J., Bali, K. (eds.) Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 5433\u20135442. Association for Computational Linguistics, Singapore (2023). https:\/\/doi.org\/10.18653\/v1\/2023.emnlp-main.330","DOI":"10.18653\/v1\/2023.emnlp-main.330"},{"issue":"1","key":"535_CR65","doi-asserted-by":"publisher","first-page":"97","DOI":"10.1016\/0010-0285(80)90005-5","volume":"12","author":"AM Treisman","year":"1980","unstructured":"Treisman, A.M., Gelade, G.: A feature-integration theory of attention. Cogn. Psychol. 12(1), 97\u2013136 (1980). https:\/\/doi.org\/10.1016\/0010-0285(80)90005-5","journal-title":"Cogn. Psychol."},{"key":"535_CR66","unstructured":"Unity Technologies: Position Constraint Component - Unity Manual. https:\/\/docs.unity3d.com\/6000.0\/Documentation\/Manual\/class-PositionConstraint.html. Accessed 10 Dec 2024"},{"issue":"9","key":"535_CR67","doi-asserted-by":"publisher","first-page":"1582","DOI":"10.1109\/TPAMI.2009.154","volume":"32","author":"K Van De Sande","year":"2009","unstructured":"Van De Sande, K., Gevers, T., Snoek, C.: Evaluating color descriptors for object and scene recognition. IEEE Trans. Pattern Anal. Mach. Intell. 32(9), 1582\u20131596 (2009)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"535_CR68","unstructured":"Virtual Reality Market Size: https:\/\/www.fortunebusinessinsights.com\/industry-reports\/virtual-reality-market-101378. Accessed 08 Aug 2024 (2024)"},{"key":"535_CR69","doi-asserted-by":"publisher","unstructured":"Vu, M.D., Wang, H., Chen, J., Li, Z., Zhao, S., Xing, Z., Chen, C.: Gptvoicetasker: Advancing multi-step mobile task efficiency through dynamic interface exploration and learning. In: Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. UIST \u201924. Association for Computing Machinery, New York, NY, USA (2024). https:\/\/doi.org\/10.1145\/3654777.3676356","DOI":"10.1145\/3654777.3676356"},{"key":"535_CR70","doi-asserted-by":"publisher","unstructured":"Wang, X., Rafi, T., Meng, N.: Vrguide: Efficient testing of virtual reality scenes via dynamic cut coverage. In: 2023 38th IEEE\/ACM International Conference on Automated Software Engineering (ASE), pp. 951\u2013962 (2023). https:\/\/doi.org\/10.1109\/ASE56229.2023.00197","DOI":"10.1109\/ASE56229.2023.00197"},{"key":"535_CR71","unstructured":"Wang, X., Wei, J., Schuurmans, D., Le, Q.V., Chi, E.H., Narang, S., Chowdhery, A., Zhou, D.: Self-consistency improves chain of thought reasoning in language models. In: The Eleventh International Conference on Learning Representations (2023). https:\/\/openreview.net\/forum?id=1PL1NIMMrw"},{"key":"535_CR72","doi-asserted-by":"crossref","unstructured":"Wang, Y., Xu, Z., Wang, X., Shen, C., Cheng, B., Shen, H., Xia, H.: End-to-end video instance segmentation with transformers. In: Proceedings on IEEE Conference Computer Vision and Pattern Recognition (CVPR) (2021)","DOI":"10.1109\/CVPR46437.2021.00863"},{"key":"535_CR73","doi-asserted-by":"publisher","unstructured":"Wang, X.: Vrtest: An extensible framework for automatic testing of virtual reality scenes. In: Proceedings of the ACM\/IEEE 44th International Conference on Software Engineering: Companion Proceedings. ICSE \u201922, pp. 232\u2013236. Association for Computing Machinery, New York, NY, USA (2022). https:\/\/doi.org\/10.1145\/3510454.3516870","DOI":"10.1145\/3510454.3516870"},{"key":"535_CR74","doi-asserted-by":"crossref","unstructured":"Xiao, B., Wu, H., Xu, W., Dai, X., Hu, H., Lu, Y., Zeng, M., Liu, C., Yuan, L.: Florence-2: Advancing a unified representation for a variety of vision tasks. arXiv:2311.06242 (2023)","DOI":"10.1109\/CVPR52733.2024.00461"},{"key":"535_CR75","unstructured":"Xiong, M., Hu, Z., Lu, X., LI, Y., Fu, J., He, J., Hooi, B.: Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs. In: The Twelfth International Conference on Learning Representations (2024). https:\/\/openreview.net\/forum?id=gjeQKFxFpZ"},{"key":"535_CR76","unstructured":"Yang, D., Tsai, Y.-H.H., Yamada, M.: On verbalized confidence scores for LLMs (2024). https:\/\/arxiv.org\/abs\/2412.14737"},{"key":"535_CR77","doi-asserted-by":"publisher","unstructured":"Yu, S., Fang, C., Du, M., Ling, Y., Chen, Z., Su, Z.: Practical non-intrusive gui exploration testing with visual-based robotic arms. In: Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. ICSE \u201924. Association for Computing Machinery, New York, NY, USA (2024). https:\/\/doi.org\/10.1145\/3597503.3639161","DOI":"10.1145\/3597503.3639161"},{"key":"535_CR78","unstructured":"Yuan, L., Chen, D., Chen, Y.-L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., Liu, C., Liu, M., Liu, Z., Lu, Y., Shi, Y., Wang, L., Wang, J., Xiao, B., Xiao, Z., Yang, J., Zeng, M., Zhou, L., Zhang, P.: Florence: A new foundation model for computer vision (2021). https:\/\/arxiv.org\/abs\/2111.11432"},{"key":"535_CR79","unstructured":"Zang, Y., Li, W., Han, J., Zhou, K., Loy, C.C.: Contextual object detection with multimodal large language models. arXiv:2305.18279 (2023)"},{"key":"535_CR80","unstructured":"Zhang, C., He, S., Qian, J., Li, B., Li, L., Qin, S., Kang, Y., Ma, M., Liu, G., Lin, Q., Rajmohan, S., Zhang, D., Zhang, Q.: Large Language model-brained GUI agents: A survey (2024). https:\/\/arxiv.org\/abs\/2411.18279"},{"key":"535_CR81","unstructured":"Zhang, H., Li, F., Liu, S., Zhang, L., Su, H., Zhu, J., Ni, L.M., Shum, H.-Y.: DINO: DETR with improved DeNoising anchor boxes for end-to-end object detection (2022)"}],"container-title":["Automated Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10515-025-00535-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10515-025-00535-3","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10515-025-00535-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T11:14:41Z","timestamp":1768821281000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10515-025-00535-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,18]]},"references-count":81,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,5]]}},"alternative-id":["535"],"URL":"https:\/\/doi.org\/10.1007\/s10515-025-00535-3","relation":{},"ISSN":["0928-8910","1573-7535"],"issn-type":[{"value":"0928-8910","type":"print"},{"value":"1573-7535","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,18]]},"assertion":[{"value":"8 January 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 July 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 September 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"7"}}