{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T18:30:16Z","timestamp":1777487416195,"version":"3.51.4"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2026,8,31]]},"abstract":"<jats:p>Large language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation framework to assess LLMs\u2019 proficiency in following instructions on diverse real-world tasks. We construct a hierarchical task tree encompassing seven major areas covering over 200 categories and over 800 tasks, which covers diverse capabilities such as question answering, reasoning, multi-turn dialogue, and text generation, to evaluate LLMs in a comprehensive and in-depth manner. We also design detailed evaluation standards and processes to facilitate consistent, unbiased judgments from human evaluators. A test set of over 3,000 instances is released, spanning different difficulty levels and knowledge domains. Our work provides a standardized methodology to evaluate human alignment in LLMs for both English and Chinese. We also analyze the feasibility of automating parts of evaluation with a strong LLM (GPT-4). Our framework supports a thorough assessment of LLMs as they are integrated into real-world applications. We have made publicly available the task tree, TencentLLMEval dataset, and evaluation methodology which have been demonstrated as effective in assessing the performance of Tencent Hunyuan LLMs. By doing so, we aim to facilitate the benchmarking of advances in the development of safe and human-aligned LLMs.<\/jats:p>","DOI":"10.1145\/3732784","type":"journal-article","created":{"date-parts":[[2025,4,29]],"date-time":"2025-04-29T12:54:44Z","timestamp":1745931284000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMs"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-8037-4700","authenticated-orcid":false,"given":"Shuyi","family":"Xie","sequence":"first","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4502-0350","authenticated-orcid":false,"given":"Wenlin","family":"Yao","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Seattle, Washington, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3041-5851","authenticated-orcid":false,"given":"Yong","family":"Dai","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2106-693X","authenticated-orcid":false,"given":"Shaobo","family":"Wang","sequence":"additional","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-4661-8918","authenticated-orcid":false,"given":"Zishan","family":"Xu","sequence":"additional","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-4461-1954","authenticated-orcid":false,"given":"Fan","family":"Lin","sequence":"additional","affiliation":[{"name":"Southeast University, Nanjing, China and Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5987-646X","authenticated-orcid":false,"given":"Donglin","family":"Zhou","sequence":"additional","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6754-7014","authenticated-orcid":false,"given":"Lifeng","family":"Jin","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Seattle, Washington, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-2061-0175","authenticated-orcid":false,"given":"Xinhua","family":"Feng","sequence":"additional","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-3709-378X","authenticated-orcid":false,"given":"Pengzhi","family":"Wei","sequence":"additional","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-9232-0506","authenticated-orcid":false,"given":"Yujie","family":"Lin","sequence":"additional","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-6626-7337","authenticated-orcid":false,"given":"Zhichao","family":"Hu","sequence":"additional","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0520-6844","authenticated-orcid":false,"given":"Dong","family":"Yu","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Seattle, Washington, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6606-2525","authenticated-orcid":false,"given":"Zhengyou","family":"Zhang","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Seattle, Washington, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-4334-1937","authenticated-orcid":false,"given":"Jing","family":"Nie","sequence":"additional","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-9083-4977","authenticated-orcid":false,"given":"Yuhong","family":"Liu","sequence":"additional","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,28]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Yuntao Bai Andy Jones Kamal Ndousse Amanda Askell Anna Chen Nova DasSarma Dawn Drain Stanislav Fort Deep Ganguli Tom Henighan et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv:2204.05862. Retrieved from https:\/\/arxiv.org\/abs\/2204.05862"},{"key":"e_1_3_3_3_2","unstructured":"Francesco Bombassei De Bona Gabriele Dominici Tim Miller Marc Langheinrich and Martin Gjoreski. 2024. Evaluating explanations through LLMs: Beyond traditional user studies. arXiv:2410.17781. Retrieved from https:\/\/arxiv.org\/abs\/2410.17781"},{"key":"e_1_3_3_4_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Chan Chi-Min","year":"2023","unstructured":"Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. ChatEval: Towards better LLM-based evaluators through multi-agent debate. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_3_5_2","unstructured":"Yupeng Chang Xu Wang Jindong Wang Yuan Wu Kaijie Zhu Hao Chen Linyi Yang Xiaoyuan Yi Cunxiang Wang Yidong Wang et al. 2023. A survey on evaluation of large language models. arXiv:2307.03109. Retrieved from https:\/\/arxiv.org\/abs\/2307.03109"},{"key":"e_1_3_3_6_2","article-title":"Deep reinforcement learning from human preferences","volume":"30","author":"Christiano Paul F.","year":"2017","unstructured":"Paul F. Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_7_2","unstructured":"Zhumin Chu Qingyao Ai Yiteng Tu Haitao Li and Yiqun Liu. 2024. PRE: A peer review based large language model evaluator. arXiv:2401.15641. Retrieved from https:\/\/arxiv.org\/abs\/2401.15641"},{"key":"e_1_3_3_8_2","unstructured":"OpenCompass Contributors. 2023. OpenCompass: A universal evaluation platform for foundation models. Retrieved from https:\/\/github.com\/open-compass\/opencompass"},{"key":"e_1_3_3_9_2","unstructured":"Dan Hendrycks Collin Burns Steven Basart Andy Zou Mantas Mazeika Dawn Song and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv:2009.03300. Retrieved from https:\/\/arxiv.org\/abs\/2009.03300"},{"key":"e_1_3_3_10_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR","author":"Hendrycks Dan","year":"2021","unstructured":"Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_3_11_2","unstructured":"Yuzhen Huang Yuzhuo Bai Zhihao Zhu Junlei Zhang Jinghan Zhang Tangjun Su Junteng Liu Chuancheng Lv Yikai Zhang Jiayi Lei et al. 2023. C-eval: A multi-level multi-discipline Chinese evaluation suite for foundation models. arXiv:2305.08322. Retrieved from https:\/\/arxiv.org\/abs\/2305.08322"},{"key":"e_1_3_3_12_2","unstructured":"Yuzhen Huang Yuzhuo Bai Zhihao Zhu Junlei Zhang Jinghan Zhang Tangjun Su Junteng Liu Chuancheng Lv Yikai Zhang Jiayi Lei et al. 2023. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_3_13_2","unstructured":"Haonan Li Yixuan Zhang Fajri Koto Yifei Yang Hai Zhao Yeyun Gong Nan Duan and Timothy Baldwin. 2023. CMMLU: Measuring massive multitask language understanding in Chinese. arXiv:2306.09212. Retrieved from https:\/\/arxiv.org\/abs\/2306.09212"},{"key":"e_1_3_3_14_2","unstructured":"Yanyang Li Jianqiao Zhao Duo Zheng Zi-Yuan Hu Zhi Chen Xiaohui Su Yongfeng Huang Shijia Huang Dahua Lin Michael R Lyu et al. 2023. CLEVA: Chinese Language Models EVAluation Platform. arXiv:2308.04813. Retrieved from https:\/\/arxiv.org\/abs\/2308.04813"},{"key":"e_1_3_3_15_2","unstructured":"Percy Liang Rishi Bommasani Tony Lee Dimitris Tsipras Dilara Soylu Michihiro Yasunaga Yian Zhang Deepak Narayanan Yuhuai Wu Ananya Kumar et al. 2022. Holistic evaluation of language models. arXiv:2211.09110. Retrieved from https:\/\/arxiv.org\/abs\/2211.09110"},{"key":"e_1_3_3_16_2","unstructured":"Fan Lin Shuyi Xie Yong Dai Wenlin Yao Tianjiao Lang Zishan Xu Zhichao Hu Xiao Xiao Yuhong Liu and Yu Zhang. 2024. IDGen: Item discrimination induced prompt generation for LLM evaluation. arXiv:2409.18892. Retrieved from https:\/\/arxiv.org\/abs\/2409.18892"},{"key":"e_1_3_3_17_2","unstructured":"Chuang Liu Renren Jin Yuqi Ren Linhao Yu Tianyu Dong Xiaohan Peng Shuting Zhang Jianxiang Peng Peiyi Zhang Qingqing Lyu et al. 2023. M3KE: A massive multi-level multi-subject knowledge evaluation benchmark for Chinese large language models. arXiv:2305.10263. Retrieved from https:\/\/arxiv.org\/abs\/2305.10263"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_3_3_19_2","volume-title":"An Overview of Bard: An Early Experiment with Generative AI","author":"Manyika James","year":"2023","unstructured":"James Manyika. 2023. An Overview of Bard: An Early Experiment with Generative AI. Technical Report, Google AI."},{"key":"e_1_3_3_20_2","unstructured":"OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11431-020-1647-3"},{"key":"e_1_3_3_22_2","unstructured":"Aarohi Srivastava Abhinav Rastogi Abhishek Rao Abu Awal Md Shoeb Abubakar Abid Adam Fisch Adam R. Brown Adam Santoro Aditya Gupta Adri\u00e0 Garriga-Alonso et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv:2206.04615. Retrieved from https:\/\/arxiv.org\/abs\/2206.04615"},{"key":"e_1_3_3_23_2","article-title":"Superglue: A stickier benchmark for general-purpose language understanding systems","volume":"32","author":"Wang Alex","year":"2019","unstructured":"Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems, Vol. 32.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_24_2","doi-asserted-by":"crossref","unstructured":"Alex Wang Amanpreet Singh Julian Michael Felix Hill Omer Levy and Samuel R. Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv:1804.07461. Retrieved from https:\/\/arxiv.org\/abs\/1804.07461","DOI":"10.18653\/v1\/W18-5446"},{"key":"e_1_3_3_25_2","unstructured":"Peiyi Wang Lei Li Liang Chen Dawei Zhu Binghuai Lin Yunbo Cao Qi Liu Tianyu Liu and Zhifang Sui. 2023. Large language models are not fair evaluators. arXiv:2305.17926. Retrieved from https:\/\/arxiv.org\/abs\/2305.17926"},{"key":"e_1_3_3_26_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Wang Yidong","year":"2023","unstructured":"Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Wenjin Yao, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, et al. 2023. PandaLM: An automatic evaluation benchmark for LLM instruction tuning optimization. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_3_27_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, Vol. 35, 24824\u201324837.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_28_2","unstructured":"Chengxing Xie Canyu Chen Feiran Jia Ziyu Ye Shiyang Lai Kai Shu Jindong Gu Adel Bibi Ziniu Hu David Jurgens et al. 2024. Can large language model agents simulate human trust behavior? arXiv:2402.04559. Retrieved from https:\/\/arxiv.org\/abs\/2402.04559"},{"key":"e_1_3_3_29_2","unstructured":"Liang Xu Anqi Li Lei Zhu Hang Xue Changtai Zhu Kangkang Zhao Haonan He Xuanwei Zhang Qiyue Kang and Zhenzhong Lan. 2023. SuperCLUE: A comprehensive Chinese large language model benchmark. arXiv:2307.15020. Retrieved from https:\/\/arxiv.org\/abs\/2307.15020"},{"key":"e_1_3_3_30_2","unstructured":"Aiyuan Yang Bin Xiao Bingning Wang Borong Zhang Chao Yin Chenxu Lv Da Pan Dian Wang Dong Yan Fan Yang et al. 2023. Baichuan 2: Open large-scale language models. arXiv:2309.10305. Retrieved from https:\/\/arxiv.org\/abs\/2309.10305"},{"key":"e_1_3_3_31_2","unstructured":"Xiaotian Zhang Chunyang Li Yi Zong Zhengyu Ying Liang He and Xipeng Qiu. 2023. Evaluating the performance of large language models on GAOKAO benchmark. arXiv:2305.12474. Retrieved from https:\/\/arxiv.org\/abs\/2305.12474"},{"key":"e_1_3_3_32_2","unstructured":"Lianmin Zheng Wei-Lin Chiang Ying Sheng Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zi Lin Zhuohan Li Dacheng Li Eric Xing et al. 2023. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv:2306.05685. Retrieved from https:\/\/arxiv.org\/abs\/2306.05685"},{"key":"e_1_3_3_33_2","unstructured":"Wanjun Zhong Ruixiang Cui Yiduo Guo Yaobo Liang Shuai Lu Yanlin Wang Amin Saied Weizhu Chen and Nan Duan. 2023. Agieval: A human-centric benchmark for evaluating foundation models. arXiv:2304.06364. Retrieved from https:\/\/arxiv.org\/abs\/2304.06364"},{"key":"e_1_3_3_34_2","unstructured":"Gu Zhouhong Zhu Xiaoxuan Ye Haoning Zhang Lin Wang Jianchen Jiang Sihang Xiong Zhuozhi Li Zihan He Qianyu Xu Rui et al. 2023. Xiezhi: An ever-updating benchmark for holistic domain knowledge evaluation. arXiv:2304.11679. Retrieved from https:\/\/arxiv.org\/abs\/2304.11679"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3732784","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T14:38:05Z","timestamp":1777387085000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3732784"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,28]]},"references-count":33,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,8,31]]}},"alternative-id":["10.1145\/3732784"],"URL":"https:\/\/doi.org\/10.1145\/3732784","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,28]]},"assertion":[{"value":"2024-01-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-26","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}