{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:26:03Z","timestamp":1781533563063,"version":"3.54.5"},"reference-count":33,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,5,30]],"date-time":"2026-05-30T00:00:00Z","timestamp":1780099200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100018537","name":"National Science and Technology Major Project","doi-asserted-by":"publisher","award":["2025ZD1602303"],"award-info":[{"award-number":["2025ZD1602303"]}],"id":[{"id":"10.13039\/501100018537","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U25A20433"],"award-info":[{"award-number":["U25A20433"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["92567203"],"award-info":[{"award-number":["92567203"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["42401521"],"award-info":[{"award-number":["42401521"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"award":["U25A20433"],"award-info":[{"award-number":["U25A20433"]}],"id":[{"id":"https:\/\/ror.org\/01h0zpd94","id-type":"ROR","asserted-by":"publisher"}]},{"award":["92567203"],"award-info":[{"award-number":["92567203"]}],"id":[{"id":"https:\/\/ror.org\/01h0zpd94","id-type":"ROR","asserted-by":"publisher"}]},{"award":["42401521"],"award-info":[{"award-number":["42401521"]}],"id":[{"id":"https:\/\/ror.org\/01h0zpd94","id-type":"ROR","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"Joint Research Fund for Beijing Natural Science Foundation and Haidian Original Innovation","doi-asserted-by":"publisher","award":["L232001"],"award-info":[{"award-number":["L232001"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Henan Key Research and Development Program","award":["241111320700"],"award-info":[{"award-number":["241111320700"]}]},{"DOI":"10.13039\/501100021171","name":"GuangDong Basic and Applied Basic Research Foundation","doi-asserted-by":"crossref","award":["2024A1515011866"],"award-info":[{"award-number":["2024A1515011866"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100021171","name":"GuangDong Basic and Applied Basic Research Foundation","doi-asserted-by":"crossref","award":["2024A1515011480"],"award-info":[{"award-number":["2024A1515011480"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100021171","name":"GuangDong Basic and Applied Basic Research Foundation","doi-asserted-by":"crossref","award":["2025A1515011300"],"award-info":[{"award-number":["2025A1515011300"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Central Guidance on Local Science and Technology Development Fund of ShanXi Province","award":["YDZJSX20231D005"],"award-info":[{"award-number":["YDZJSX20231D005"]}]},{"name":"Central Guidance on Local Science and Technology Development Fund of ShanXi Province","award":["YDZJSX2024B017"],"award-info":[{"award-number":["YDZJSX2024B017"]}]},{"name":"Science and Technology Innovation Program of Xiongan New Area","award":["2025XAGG0028"],"award-info":[{"award-number":["2025XAGG0028"]}]},{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2023YFF0905903"],"award-info":[{"award-number":["2023YFF0905903"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>The of personal computer (PC) tasks represents a systems-level challenge that integrates natural language processing, visual perception and mouse\u2013keyboard action control. While existing approaches mainly focus on the application programming interface (API)-based or terminal-based automation, which are incompatible with the majority of applications for the lack of accessible interface. In this article, we propose PCLLM, a novel end-to-end system that automates PC operations by integrating large language models (LLMs) with computer vision techniques to directly control the mouse and keyboard. First, a software knowledge-based prompt engineering method is developed to comprehend software architecture and operational sequences. Second, template matching techniques are integrated for precise element localization, allowing the system to accurately identify and interact. Third, a dual-LLM pipeline is designed to automatically generate the test data, where a questioner LLM generates diverse task commands and the PCLLM executes these tasks, the corresponding process data are recorded automatically for performance evaluation. Finally, PCLLM is further validated through three typically PC applications (Notepad, Wordpad and Calculator), demonstrating its flexible and robust performance towards intelligent PC automation. To evaluate the proposed system, we adopt task completion rate as the primary metric. Experimental results show that PCLLM achieves the highest completion rates of 98.59%, 95.77%, and 52.11% on Notepad for basic, intermediate, and advanced tasks respectively when powered by GPT-4o, outperforming the CogAgent baseline. These results demonstrate the effectiveness of our approach for PC task automation.<\/jats:p>","DOI":"10.3390\/computers15060351","type":"journal-article","created":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T17:11:01Z","timestamp":1780333861000},"page":"351","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["PCLLM: An Integrated LLM-Driven System for Automating Desktop Operations via Direct Mouse and Keyboard Control"],"prefix":"10.3390","volume":"15","author":[{"given":"Zhenqian","family":"Wang","sequence":"first","affiliation":[{"name":"School of Computer and Communication Engineering, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Dong","sequence":"additional","affiliation":[{"name":"Shanxi Taihang Laboratory Co., Ltd., Xi\u2019an 030006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2036-2784","authenticated-orcid":false,"given":"Meixia","family":"Fu","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Industrial Deterministic Networks and Intelligent Collaborative Control, School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing 100083, China"},{"name":"Shunde Graduate School, University of Science and Technology Beijing, Foshan 528399, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianquan","family":"Wang","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Industrial Deterministic Networks and Intelligent Collaborative Control, School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jie","family":"Sun","sequence":"additional","affiliation":[{"name":"Shanxi Taihang Laboratory Co., Ltd., Xi\u2019an 030006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6551-6807","authenticated-orcid":false,"given":"Qu","family":"Wang","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Industrial Deterministic Networks and Intelligent Collaborative Control, School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing 100083, China"},{"name":"Shunde Graduate School, University of Science and Technology Beijing, Foshan 528399, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yifan","family":"Lu","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Industrial Deterministic Networks and Intelligent Collaborative Control, School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8382-7729","authenticated-orcid":false,"given":"Na","family":"Chen","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Industrial Deterministic Networks and Intelligent Collaborative Control, School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ronghui","family":"Zhang","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Industrial Deterministic Networks and Intelligent Collaborative Control, School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wen","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer and Communication Engineering, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,5,30]]},"reference":[{"key":"ref_1","unstructured":"Zhang, C., He, S., Qian, J., Li, B., Li, L., Qin, S., Kang, Y., Ma, M., Liu, G., and Lin, Q. (2024). Large language model-brained gui agents: A survey. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zimmermann, D., and Koziolek, A. (2023, January 16\u201320). Automating gui-based software testing with gpt-3. Proceedings of the 2023 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW), Dublin, Ireland.","DOI":"10.1109\/ICSTW58534.2023.00022"},{"key":"ref_3","first-page":"532659","article-title":"End User Development: Survey of an Emerging Field for Empowering People","volume":"2013","author":"Fabio","year":"2013","journal-title":"Isrn Softw. Eng."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Schneider, S., Werner, S., Khalili, R., Hecker, A., and Karl, H. (2022, January 25\u201329). mobile-env: An open platform for reinforcement learning in wireless mobile networks. Proceedings of the NOMS 2022\u20132022 IEEE\/IFIP Network Operations and Management Symposium, Budapest, Hungary.","DOI":"10.1109\/NOMS54207.2022.9789886"},{"key":"ref_5","unstructured":"Collins, E., Neto, A., Vincenzi, A., and Maldonado, J. (October, January 27). Deep reinforcement learning based android application gui testing. Proceedings of the XXXV Brazilian Symposium on Software Engineering, Joinville, Brazil."},{"key":"ref_6","first-page":"28091","article-title":"Mind2web: Towards a generalist agent for the web","volume":"36","author":"Deng","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Pasupat, P., Jiang, T.S., Liu, E., Guu, K., and Liang, P. (November, January 31). Mapping natural language commands to web elements. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium.","DOI":"10.18653\/v1\/D18-1540"},{"key":"ref_8","unstructured":"Shi, T., Karpathy, A., Fan, L., Hernandez, J., and Liang, P. (2017, January 6\u201311). World of bits: An open-domain platform for web-based agents. Proceedings of the International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_9","first-page":"20744","article-title":"Webshop: Towards scalable real-world web interaction with grounded language agents","volume":"35","author":"Yao","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_10","unstructured":"Zhou, S., Xu, F.F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., and Fried, D. (2024, January 7\u201311). Webarena: A realistic web environment for building autonomous agents. Proceedings of the International Conference on Learning Representations, Vienna, Austria."},{"key":"ref_11","first-page":"38154","article-title":"Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face","volume":"36","author":"Shen","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_12","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_13","unstructured":"Pan, Y., Kong, D., Zhou, S., Cui, C., Leng, Y., Jiang, B., Liu, H., Shang, Y., Zhou, S., and Wu, T. (2024). Webcanvas: Benchmarking web agents in online environments. arXiv."},{"key":"ref_14","unstructured":"Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J. (2023). Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V. arXiv."},{"key":"ref_15","unstructured":"Yan, A., Yang, Z., Zhu, W., Lin, K., Li, L., Wang, J., Yang, J., Zhong, Y., McAuley, J., and Gao, J. (2023). Gpt-4v in wonderland: Large multimodal models for zero-shot smartphone gui navigation. arXiv."},{"key":"ref_16","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., and Azhar, F. (2023). Llama: Open and efficient foundation language models. arXiv."},{"key":"ref_17","unstructured":"Delaflor, M., Gendron, C., Toxtli, C., Li, W., and Delgado-Sol\u00f3rzano, C.T. (2023, January 20\u201324). ReActIn: Infusing Human Feedback into Intermediate Prompting Steps of Large Language Model. Proceedings of the AHFE International, San Francisco, CA, USA."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Burns, A., Arsan, D., Agrawal, S., Kumar, R., Saenko, K., and Plummer, B.A. (2022). A dataset for interactive vision-language navigation with unknown command feasibility. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-031-20074-8_18"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Li, Y., He, J., Zhou, X., Zhang, Y., and Baldridge, J. (2020, January 5\u201310). Mapping natural language instructions to mobile UI action sequences. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.acl-main.729"},{"key":"ref_20","first-page":"59708","article-title":"Androidinthewild: A large-scale dataset for android device control","volume":"36","author":"Rawles","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wen, H., Li, Y., Liu, G., Zhao, S., Yu, T., Li, T.J.J., Jiang, S., Liu, Y., Zhang, Y., and Liu, Y. (October, January 30). AutoDroid: LLM-powered Task Automation in Android. Proceedings of the 30th Annual International Conference on Mobile Computing and Networking (MobiCom 2024), Washington, DC, USA.","DOI":"10.1145\/3636534.3649379"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Gao, D., Ji, L., Bai, Z., Ouyang, M., Li, P., Mao, D., Wu, Q., Zhang, W., Wang, P., and Guo, X. (2024, January 16\u201322). ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01262"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Lin, X.V., Wang, C., Zettlemoyer, L., and Ernst, M.D. (2018, January 7\u201312). NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System. Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan.","DOI":"10.63317\/42er6x4bacdf"},{"key":"ref_24","unstructured":"Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., and Yang, K. (2024, January 7\u201311). AgentBench: Evaluating LLMs as Agents. Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria."},{"key":"ref_25","unstructured":"Wen, H., Wang, H., Liu, J., and Li, Y. (2023). Droidbot-gpt: Gpt-powered ui automation for android. arXiv."},{"key":"ref_26","unstructured":"Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Dong, Y., and Ding, M. (, January 16\u201322). Cogagent: A visual language model for gui agents. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA."},{"key":"ref_27","first-page":"68539","article-title":"Toolformer: Language models can teach themselves to use tools","volume":"36","author":"Schick","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_28","unstructured":"Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., and Qian, B. (2024, January 7\u201311). ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria."},{"key":"ref_29","unstructured":"Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., and Saunders, W. (2021). Webgpt: Browser-assisted question-answering with human feedback. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhao, N. (2024, January 24\u201326). Enhancing object detection with yolov8 transfer learning: A voc2012 dataset study. Proceedings of the International Conference Pattern Recognition Applications and Methods (ICPRAM), Rome, Italy.","DOI":"10.5220\/0012939600004508"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"110","DOI":"10.1007\/s00530-025-02152-2","article-title":"YOLO-TCS: An enhanced multi-scale network for traffic sign detection integrating multi-level feature fusion and attention","volume":"32","author":"Yu","year":"2026","journal-title":"Multimed. Syst."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"114770","DOI":"10.1016\/j.asoc.2026.114770","article-title":"Feature-aware multi-head self-attention hashing for Chinese ancient document image retrieval","volume":"193","author":"Jiang","year":"2026","journal-title":"Appl. Soft Comput."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhu, V., Ji, Z., Guo, D., Wang, P., Xia, Y., Lu, L., Ye, X., Zhu, W., and Jin, D. (2024). Low-rank continual pyramid vision transformer: Incrementally segment whole-body organs in CT with light-weighted adaptation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer.","DOI":"10.1007\/978-3-031-72111-3_35"}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/6\/351\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T04:14:33Z","timestamp":1780373673000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/6\/351"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,30]]},"references-count":33,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["computers15060351"],"URL":"https:\/\/doi.org\/10.3390\/computers15060351","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,30]]}}}