{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T05:05:53Z","timestamp":1765343153285,"version":"3.46.0"},"publisher-location":"New York, NY, USA","reference-count":33,"publisher":"ACM","funder":[{"name":"National Research Foundation, Singapore","award":["NRFF Award NRF-NRFF13-2021-0008"],"award-info":[{"award-number":["NRFF Award NRF-NRFF13-2021-0008"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,27]]},"DOI":"10.1145\/3746027.3755618","type":"proceedings-article","created":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T07:27:39Z","timestamp":1761377259000},"page":"8692-8700","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Can I Trust You? Advancing GUI Task Automation with Action Trust Score"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3549-9684","authenticated-orcid":false,"given":"Haiyang","family":"Mei","sequence":"first","affiliation":[{"name":"Show Lab, National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8494-3492","authenticated-orcid":false,"given":"Difei","family":"Gao","sequence":"additional","affiliation":[{"name":"Show Lab, National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8497-611X","authenticated-orcid":false,"given":"Xiaopeng","family":"Wei","sequence":"additional","affiliation":[{"name":"Key Laboratory of Social Computing and Cognitive Intelligence, Dalian University of Technology, Dalian, Liaoning, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8046-722X","authenticated-orcid":false,"given":"Xin","family":"Yang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Social Computing and Cognitive Intelligence, Dalian University of Technology, Dalian, Liaoning, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7681-2166","authenticated-orcid":false,"given":"Mike Zheng","family":"Shou","sequence":"additional","affiliation":[{"name":"Show Lab, National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,27]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Exploring Human-AI Collaboration in Agile: Customised LLM Meeting Assistants. arXiv:2404","author":"Cabrero-Daniel Beatriz","year":"2024","unstructured":"Beatriz Cabrero-Daniel, Tomas Herda, Victoria Pichler, and Martin Eder. 2024. Exploring Human-AI Collaboration in Agile: Customised LLM Meeting Assistants. arXiv:2404.14871 (2024)."},{"key":"e_1_3_2_2_2_1","volume-title":"Optimising human-ai collaboration by learning convincing explanations. arXiv:2311.07426","author":"Chan Alex J","year":"2023","unstructured":"Alex J Chan, Alihan Huyuk, and Mihaela van der Schaar. 2023. Optimising human-ai collaboration by learning convincing explanations. arXiv:2311.07426 (2023)."},{"key":"e_1_3_2_2_3_1","unstructured":"Nina Corvelo Benz and Manuel Rodriguez. 2024. Human-aligned calibration for ai-assisted decision making. In NeurIPS."},{"key":"e_1_3_2_2_4_1","unstructured":"Xiang Deng Yu Gu Boyuan Zheng Shijie Chen Sam Stevens Boshi Wang Huan Sun and Yu Su. 2024. Mind2web: Towards a generalist agent for the web. In NeurIPS."},{"key":"e_1_3_2_2_5_1","volume-title":"Assistgui: Task-oriented desktop graphical user interface automation. In CVPR.","author":"Gao Difei","year":"2024","unstructured":"Difei Gao, Lei Ji, Zechen Bai, Mingyu Ouyang, Peiran Li, Dongxing Mao, Qinchen Wu, Weichen Zhang, Peiyi Wang, Xiangwu Guo, et al., 2024. Assistgui: Task-oriented desktop graphical user interface automation. In CVPR."},{"key":"e_1_3_2_2_6_1","volume-title":"Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. arXiv:2111.09543","author":"He Pengcheng","year":"2021","unstructured":"Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. arXiv:2111.09543 (2021)."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"crossref","unstructured":"Alexander Kirillov Eric Mintun Nikhila Ravi Hanzi Mao Chloe Rolland Laura Gustafson Tete Xiao Spencer Whitehead Alexander C Berg Wan-Yen Lo et al. 2023. Segment anything. In ICCV.","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_3_2_2_8_1","volume-title":"Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents. arXiv:2405.02957","author":"Li Junkai","year":"2024","unstructured":"Junkai Li, Siyu Wang, Meng Zhang, Weitao Li, Yunghwei Lai, Xinhui Kang, Weizhi Ma, and Yang Liu. 2024. Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents. arXiv:2405.02957 (2024)."},{"key":"e_1_3_2_2_9_1","volume-title":"Teaching models to express their uncertainty in words. arXiv:2205.14334","author":"Lin Stephanie","year":"2022","unstructured":"Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Teaching models to express their uncertainty in words. arXiv:2205.14334 (2022)."},{"key":"e_1_3_2_2_10_1","unstructured":"Haotian Liu Chunyuan Li Yuheng Li Bo Li Yuanhan Zhang Sheng Shen and Yong Jae Lee. 2024. LLaVA-NeXT: Improved reasoning OCR and world knowledge. https:\/\/llava-vl.github.io\/blog\/2024-01-30-llava-next\/"},{"key":"e_1_3_2_2_11_1","unstructured":"Shilong Liu Zhaoyang Zeng Tianhe Ren Feng Li Hao Zhang Jie Yang Chunyuan Li Jianwei Yang Hang Su Jun Zhu et al. 2023. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv:2303.05499 (2023)."},{"key":"e_1_3_2_2_12_1","volume-title":"Decoupled weight decay regularization. arXiv:1711.05101","author":"Loshchilov Ilya","year":"2017","unstructured":"Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv:1711.05101 (2017)."},{"key":"e_1_3_2_2_13_1","volume-title":"Exploring Dense Context for Salient Object Detection","author":"Mei Haiyang","year":"2021","unstructured":"Haiyang Mei, Yuanyuan Liu, Ziqi Wei, Dongsheng Zhou, Xiaopeng Wei, Qiang Zhang, and Xin Yang. 2021. Exploring Dense Context for Salient Object Detection. IEEE TCSVT (2021)."},{"key":"e_1_3_2_2_14_1","unstructured":"MetaAI. 2024. Meta Llama3. https:\/\/llama.meta.com\/llama3"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00494"},{"key":"e_1_3_2_2_16_1","unstructured":"OpenAI. 2024. ChatGPT. https:\/\/openai.com\/chatgpt"},{"key":"e_1_3_2_2_17_1","unstructured":"Adam Paszke Sam Gross Francisco Massa Adam Lerer James Bradbury Gregory Chanan Trevor Killeen Zeming Lin Natalia Gimelshein Luca Antiga et al. 2019. PyTorch: An imperative style high-performance deep learning library. In NeurIPS."},{"key":"e_1_3_2_2_18_1","volume-title":"Nicholas King, Harsha Nori, and Saleema Amershi.","author":"Rastogi Charvi","year":"2023","unstructured":"Charvi Rastogi, Marco Tulio Ribeiro, Nicholas King, Harsha Nori, and Saleema Amershi. 2023. Supporting human-ai collaboration in auditing llms with llms. In AIES."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"crossref","unstructured":"Joseph Redmon Santosh Divvala Ross Girshick and Ali Farhadi. 2016. You only look once: Unified real-time object detection. In CVPR.","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"crossref","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In EMNLP.","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_2_2_21_1","volume-title":"Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv:2305.14975","author":"Tian Katherine","year":"2023","unstructured":"Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv:2305.14975 (2023)."},{"key":"e_1_3_2_2_22_1","volume-title":"Androidenv: A reinforcement learning platform for android. arXiv:2105.13231","author":"Toyama Daniel","year":"2021","unstructured":"Daniel Toyama, Philippe Hamel, Anita Gergely, Gheorghe Comanici, Amelia Glaese, Zafarali Ahmed, Tyler Jackson, Shibl Mourad, and Doina Precup. 2021. Androidenv: A reinforcement learning platform for android. arXiv:2105.13231 (2021)."},{"key":"e_1_3_2_2_23_1","unstructured":"Kailas Vodrahalli Tobias Gerstenberg and James Y Zou. 2022. Uncalibrated models can improve human-ai collaboration. In NeurIPS."},{"key":"e_1_3_2_2_24_1","volume-title":"GOLF: Goal-Oriented Long-term liFe tasks supported by human-AI collaboration. arXiv:2403.17089","author":"Wang Ben","year":"2024","unstructured":"Ben Wang. 2024. GOLF: Goal-Oriented Long-term liFe tasks supported by human-AI collaboration. arXiv:2403.17089 (2024)."},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-024-40231-1"},{"key":"e_1_3_2_2_26_1","unstructured":"Peng Wang Shuai Bai Sinan Tan Shijie Wang Zhihao Fan Jinze Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge et al. 2024a. Qwen2-vl: Enhancing vision-language model's perception of the world at any resolution. arXiv:2409.12191 (2024)."},{"key":"e_1_3_2_2_27_1","volume-title":"Denny Zhou, et al.","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al., 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS."},{"key":"e_1_3_2_2_28_1","volume-title":"Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu.","author":"Wen Hao","year":"2023","unstructured":"Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2023. Empowering llm to use smartphone for intelligent task automation. arXiv:2308.15272 (2023)."},{"key":"e_1_3_2_2_29_1","volume-title":"Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al.","author":"Xie Tianbao","year":"2024","unstructured":"Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al., 2024. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. arXiv:2404.07972 (2024)."},{"key":"e_1_3_2_2_30_1","volume-title":"Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. arXiv:2306.13063","author":"Xiong Miao","year":"2023","unstructured":"Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2023. Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. arXiv:2306.13063 (2023)."},{"key":"e_1_3_2_2_31_1","volume-title":"Webshop: Towards scalable real-world web interaction with grounded language agents. In NeurIPS.","author":"Yao Shunyu","year":"2022","unstructured":"Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. Webshop: Towards scalable real-world web interaction with grounded language agents. In NeurIPS."},{"key":"e_1_3_2_2_32_1","unstructured":"Shunyu Yao Jeffrey Zhao Dian Yu Nan Du Izhak Shafran Karthik Narasimhan and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In ICLR."},{"key":"e_1_3_2_2_33_1","volume-title":"Can large language models transform computational social science? Computational Linguistics","author":"Ziems Caleb","year":"2024","unstructured":"Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024. Can large language models transform computational social science? Computational Linguistics (2024), 1-55."}],"event":{"name":"MM '25: The 33rd ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Dublin Ireland","acronym":"MM '25"},"container-title":["Proceedings of the 33rd ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746027.3755618","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T05:02:30Z","timestamp":1765342950000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746027.3755618"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,27]]},"references-count":33,"alternative-id":["10.1145\/3746027.3755618","10.1145\/3746027"],"URL":"https:\/\/doi.org\/10.1145\/3746027.3755618","relation":{},"subject":[],"published":{"date-parts":[[2025,10,27]]},"assertion":[{"value":"2025-10-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}