{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T05:43:14Z","timestamp":1780638194866,"version":"3.54.1"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"5-6","license":[{"start":{"date-parts":[[2024,11,11]],"date-time":"2024-11-11T00:00:00Z","timestamp":1731283200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2024,11,11]],"date-time":"2024-11-11T00:00:00Z","timestamp":1731283200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int. J. Mach. Learn. &amp; Cyber."],"published-print":{"date-parts":[[2025,6]]},"DOI":"10.1007\/s13042-024-02432-9","type":"journal-article","created":{"date-parts":[[2024,11,11]],"date-time":"2024-11-11T03:31:50Z","timestamp":1731295910000},"page":"2997-3017","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Performance evaluations of large language models for customer service"],"prefix":"10.1007","volume":"16","author":[{"given":"Fei","family":"Li","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanyan","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yin","family":"Xu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shiling","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Junli","family":"Liang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhengyi","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenrui","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qiangzhong","family":"Feng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ticheng","family":"Duan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Youzhi","family":"Huang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qi","family":"Song","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiangyang","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,11,11]]},"reference":[{"key":"2432_CR1","unstructured":"Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman FL, Almeida D, Altenschmidt J, Altman S, Anadkat S (2023) Gpt-4 technical report. arXiv:2303.08774"},{"key":"2432_CR2","unstructured":"Touvron H, Lavril T, Izacard G, Martinet X, Lachaux M-A, Lacroix T, Rozi\u00e8re B, Goyal N, Hambro E, Azhar F (2023) Llama: open and efficient foundation language models. arXiv:2302.13971"},{"key":"2432_CR3","unstructured":"Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, Bashlykov N, Batra S, Bhargava P, Bhosale S (2023) Llama 2: open foundation and fine-tuned chat models. arXiv:2307.09288"},{"key":"2432_CR4","doi-asserted-by":"publisher","unstructured":"Du Z, Qian Y, Liu X, Ding M, Qiu J, Yang Z, Tang J (2022) Glm: general language model pretraining with autoregressive blank infilling. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: long papers), pp 320\u2013335. https:\/\/doi.org\/10.18653\/v1\/2022.acl-long.26","DOI":"10.18653\/v1\/2022.acl-long.26"},{"key":"2432_CR5","unstructured":"Zeng A, Liu X, Du Z, Wang Z, Lai H, Ding M, Yang Z, Xu Y, Zheng W, Xia X (2022) Glm-130b: an open bilingual pre-trained model. arXiv:2210.02414"},{"key":"2432_CR6","unstructured":"Yang A, Xiao B, Wang B, Zhang B, Bian C, Yin C, Lv C, Pan D, Wang D, Yan D (2023) Baichuan 2: open large-scale language models. arXiv:2309.10305"},{"key":"2432_CR7","unstructured":"Yue S, Chen W, Wang S, Li B, Shen C, Liu S, Zhou Y, Xiao Y, Yun S, Lin W (2023) Disc-lawllm: fine-tuning large language models for intelligent legal services. arXiv:2309.11325"},{"key":"2432_CR8","unstructured":"Cui J, Li Z, Yan Y, Chen B, Yuan L (2023) Chatlaw: open-source legal large language model with integrated external knowledge bases. arXiv:2306.16092"},{"issue":"6","key":"2432_CR9","first-page":"e40895","volume":"15","author":"Y Li","year":"2023","unstructured":"Li Y, Li Z, Zhang K, Dan R, Jiang S, Zhang Y (2023) Chatdoctor: a medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge. Cureus 15(6):e40895","journal-title":"Cureus"},{"key":"2432_CR10","unstructured":"Xiong H, Wang S, Zhu Y, Zhao Z, Liu Y, Huang L, Wang Q, Shen D (2023) Doctorglm: fine-tuning your Chinese doctor is not a herculean task. arXiv:2304.01097"},{"key":"2432_CR11","unstructured":"Wu S, Irsoy O, Lu S, Dabravolski V, Dredze M, Gehrmann S, Kambadur P, Rosenberg D, Mann G (2023) Bloomberggpt: a large language model for finance. arXiv:2303.17564"},{"key":"2432_CR12","unstructured":"Chen W, Wang Q, Long Z, Zhang X, Lu Z, Li B, Wang S, Xu J, Bai X, Huang X (2023) Disc-finllm: a Chinese financial large language model based on multiple experts fine-tuning. arXiv:2310.15205"},{"key":"2432_CR13","unstructured":"Hendrycks D, Burns C, Basart S, Zou A, Mazeika M, Song D, Steinhardt J (2020) Measuring massive multitask language understanding. arXiv:2009.03300"},{"key":"2432_CR14","unstructured":"Srivastava A, Rastogi A, Rao A, Shoeb AAM, Abid A, Fisch A, Brown AR, Santoro A, Gupta A, Garriga-Alonso A (2022) Beyond the imitation game: quantifying and extrapolating the capabilities of language models. arXiv:2206.04615"},{"key":"2432_CR15","doi-asserted-by":"crossref","unstructured":"Li H, Zhang Y, Koto F, Yang Y, Zhao H, Gong Y, Duan N, Baldwin T (2023) Cmmlu: measuring massive multitask language understanding in Chinese. arXiv:2306.09212","DOI":"10.18653\/v1\/2024.findings-acl.671"},{"key":"2432_CR16","unstructured":"Huang Y, Bai Y, Zhu Z, Zhang J, Zhang J, Su T, Liu J, Lv C, Lei J, Fu Y et al (2023) C-eval: a multi-level multi-discipline chinese evaluation suite for foundation models. In: Advances in Neural Information Processing Systems, pp  62991\u201363010"},{"key":"2432_CR17","doi-asserted-by":"publisher","unstructured":"Zhong W, Cui R, Guo Y, Liang Y, Lu S, Wang Y, Saied A, Chen W, Duan N (2024) Agieval: a human-centric benchmark for evaluating foundation models. In: Findings of the Association for Computational Linguistics: NAACL 2024, pp 2299\u20132314 . https:\/\/doi.org\/10.18653\/v1\/2024.findings-naacl.149","DOI":"10.18653\/v1\/2024.findings-naacl.149"},{"key":"2432_CR18","unstructured":"Dai Y, Feng D, Huang J, Jia H, Xie Q, Zhang Y, Han W, Tian W, Wang H (2023) Laiw: a Chinese legal large language models benchmark (a technical report). arXiv:2310.05620"},{"key":"2432_CR19","doi-asserted-by":"crossref","unstructured":"Zhu W, Wang X, Zheng H, Chen M, Tang B (2023) Promptcblue: a Chinese prompt tuning benchmark for the medical domain. arXiv:2310.14151","DOI":"10.2139\/ssrn.4685921"},{"key":"2432_CR20","unstructured":"Xie Q, Han W, Zhang X, Lai Y, Peng M, Lopez-Lira A, Huang J (2023) Pixiu: a large language model, instruction data and evaluation benchmark for finance. In: Proceedings of the 37th international conference on neural information processing systems, pp 33469\u201333484"},{"key":"2432_CR21","doi-asserted-by":"publisher","DOI":"10.1109\/MCOM.001.2300364","author":"L Bariah","year":"2024","unstructured":"Bariah L, Zhao Q, Zou H, Tian Y, Bader F, Debbah M (2024) Large generative AI models for telecom: the next big thing? IEEE Commun Mag. https:\/\/doi.org\/10.1109\/MCOM.001.2300364","journal-title":"IEEE Commun Mag"},{"key":"2432_CR22","doi-asserted-by":"publisher","DOI":"10.1109\/MCOM.001.2300473","author":"A Maatouk","year":"2024","unstructured":"Maatouk A, Piovesan N, Ayed F, De Domenico A, Debbah M (2024) Large language models for telecom: forthcoming impact on the industry. IEEE Commun Mag. https:\/\/doi.org\/10.1109\/MCOM.001.2300473","journal-title":"IEEE Commun Mag"},{"key":"2432_CR23","unstructured":"Wang Z, Liu X, Liu S, Yao Y, Huang Y, He Z, Li X, Li Y, Che Z, Zhang Z (2024) Telechat technical report. arXiv:2401.03804"},{"key":"2432_CR24","unstructured":"Le\u00a0Scao T, Fan A, Akiki C, Pavlick E, Ili\u0107 S, Hesslow D, Castagn\u00e9 R, Luccioni AS, Yvon F, Gall\u00e9 M (2022) Bloom: a 176b-parameter open-access multilingual language model. arXiv:2211.05100"},{"key":"2432_CR25","unstructured":"Bai J, Bai S, Chu Y, Cui Z, Dang K, Deng X, Fan Y, Ge W, Han Y, Huang F (2023) Qwen technical report. arXiv:2309.16609"},{"key":"2432_CR26","doi-asserted-by":"publisher","unstructured":"Wang A, Singh A, Michael J, Hill F, Levy O, Bowman SR (2019) Glue: a multi-task benchmark and analysis platform for natural language understanding. In: 7th international conference on learning representations, ICLR 2019. https:\/\/doi.org\/10.18653\/v1\/W18-5446","DOI":"10.18653\/v1\/W18-5446"},{"key":"2432_CR27","first-page":"3261","volume":"32","author":"A Wang","year":"2019","unstructured":"Wang A, Pruksachatkun Y, Nangia N, Singh A, Michael J, Hill F, Levy O, Bowman S (2019) Superglue: a stickier benchmark for general-purpose language understanding systems. Adv Neural Inf Process Syst 32:3261\u20133275","journal-title":"Adv Neural Inf Process Syst"},{"key":"2432_CR28","doi-asserted-by":"publisher","unstructured":"Xu L, Hu H, Zhang X, Li L, Cao C, Li Y, Xu Y, Sun K, Yu D, Yu C (2020) Clue: a Chinese language understanding evaluation benchmark. In: Proceedings of the 28th international conference on Computational Linguistics. International Committee on Computational Linguistics. https:\/\/doi.org\/10.18653\/v1\/2020.coling-main.419","DOI":"10.18653\/v1\/2020.coling-main.419"},{"key":"2432_CR29","unstructured":"Kenton JDM-WC, Toutanova LK (2019) Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of naacL-HLT, vol 1, p 2"},{"key":"2432_CR30","unstructured":"Liu C, Jin R, Ren Y, Yu L, Dong T, Peng X, Zhang S, Peng J, Zhang P, Lyu Q (2023) M3ke: a massive multi-level multi-subject knowledge evaluation benchmark for Chinese large language models. arXiv:2305.10263"},{"key":"2432_CR31","first-page":"46595","volume":"36","author":"L Zheng","year":"2023","unstructured":"Zheng L, Chiang W-L, Sheng Y, Zhuang S, Wu Z, Zhuang Y, Lin Z, Li Z, Li D, Xing E et al (2023) Judging llm-as-a-judge with mt-bench and chat-bot arena. Adv Neural Inf Proc Syst 36:46595\u201346623","journal-title":"Adv Neural Inf Proc Syst"},{"issue":"6","key":"2432_CR32","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3419368","volume":"53","author":"SM Preum","year":"2021","unstructured":"Preum SM, Munir S, Ma M, Yasar MS, Stone DJ, Williams R, Alemzadeh H, Stankovic JA (2021) A review of cognitive assistants for healthcare: trends, prospects, and future directions. ACM Comput Surv (CSUR) 53(6):1\u201337","journal-title":"ACM Comput Surv (CSUR)"},{"issue":"8","key":"2432_CR33","doi-asserted-by":"publisher","first-page":"1930","DOI":"10.1038\/s41591-023-02448-8","volume":"29","author":"AJ Thirunavukarasu","year":"2023","unstructured":"Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW (2023) Large language models in medicine. Nat Med 29(8):1930\u20131940","journal-title":"Nat Med"},{"key":"2432_CR34","doi-asserted-by":"publisher","unstructured":"Juglan KC, Sharma B, Gehlot A, Singh SP, Hussein A, Alazzam MB (2023) Exploring the effectiveness of natural language processing in customer service. In: 2023 3rd international conference on advance computing and innovative technologies in engineering (ICACITE), pp 814\u2013818. IEEE. https:\/\/doi.org\/10.1109\/ICACITE57410.2023.10183107","DOI":"10.1109\/ICACITE57410.2023.10183107"},{"key":"2432_CR35","unstructured":"Mashaabi M, Alotaibi A, Qudaih H, Alnashwan R, Al-Khalifa H (2022) Natural language processing in customer service: a systematic review. 2212"},{"key":"2432_CR36","unstructured":"Yijing W (2018) Intelligent customer service system design based on natural language processing. In: Proceedings of 2018 5th international conference on electrical & electronics engineering and computer science, ICEEECS"},{"key":"2432_CR37","doi-asserted-by":"crossref","unstructured":"Darwish MF (2021) A survey of predictive maintenance and self-optimization in telecom field based on machine learning. In: 2021 8th international conference on soft computing & machine intelligence (ISCMI). IEEE, pp 9\u201314","DOI":"10.1109\/ISCMI53840.2021.9654819"},{"key":"2432_CR38","doi-asserted-by":"crossref","unstructured":"Zhou H, Hu C, Yuan Y, Cui Y, Jin Y, Chen C, Wu H, Yuan D, Jiang L, Wu D (2024) Large language model (llm) for telecommunications: a comprehensive survey on principles, key techniques, and opportunities. arXiv e-prints, 2405","DOI":"10.1109\/COMST.2024.3465447"},{"key":"2432_CR39","unstructured":"Maatouk A, Ayed F, Piovesan N, De\u00a0Domenico A, Debbah M, Luo Z-Q (2023) Teleqna: a benchmark dataset to assess large language models telecommunications knowledge. arXiv:2310.15051"},{"key":"2432_CR40","doi-asserted-by":"publisher","DOI":"10.1145\/3660522","author":"H Peng","year":"2024","unstructured":"Peng H, Zhang J, Huang X, Hao Z, Li A, Yu Z, Yu PS (2024) Unsupervised social bot detection via structural information theory. ACM Trans Inf Syst. https:\/\/doi.org\/10.1145\/3660522","journal-title":"ACM Trans Inf Syst"},{"key":"2432_CR41","doi-asserted-by":"crossref","unstructured":"Lin H, Ma L, Zhu J, Xiang L, Zhou Y, Zhang J, Zong C (2021) Csds: a fine-grained Chinese dataset for customer service dialogue summarization. In: Proceedings of the 2021 conference on empirical methods in natural language processing, pp 4436\u20134451","DOI":"10.18653\/v1\/2021.emnlp-main.365"},{"key":"2432_CR42","doi-asserted-by":"publisher","unstructured":"Xi X, Lv J, Liu S, Ye W, Yang F, Wan G (2022) Musied: a benchmark for event detection from multi-source heterogeneous informal texts. In: Proceedings of the 2022 conference on empirical methods in natural language processing, pp 2947\u20132964. https:\/\/doi.org\/10.18653\/v1\/2022.emnlp-main.191","DOI":"10.18653\/v1\/2022.emnlp-main.191"},{"key":"2432_CR43","doi-asserted-by":"crossref","unstructured":"Lee K, Ippolito D, Nystrom A, Zhang C, Eck D, Callison-Burch C, Carlini N (2022) Deduplicating training data makes language models better. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp 8424\u20138445","DOI":"10.18653\/v1\/2022.acl-long.577"},{"key":"2432_CR44","doi-asserted-by":"publisher","unstructured":"Longpre S, Yauney G, Reif E, Lee K, Roberts A, Zoph B, Zhou D, Wei J, Robinson K, Mimno D (2024) A pretrainer\u2019s guide to training data: measuring the effects of data age, domain coverage, quality, & toxicity. In: Proceedings of the 2024 conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp 3245\u20133276. https:\/\/doi.org\/10.18653\/v1\/2024.naacl-long.179","DOI":"10.18653\/v1\/2024.naacl-long.179"},{"key":"2432_CR45","unstructured":"Hu E.J, Shen Y, Wallis P, Allen-Zhu Z, Li Y, Wang S, Wang L, Chen W (2021) Lora: low-rank adaptation of large language models. arXiv:2106.09685"},{"key":"2432_CR46","doi-asserted-by":"publisher","unstructured":"Duan H, Wei J, Wang C, Liu H, Fang Y, Zhang S, Lin D, Chen K (2024) Botchat: evaluating llms\u2019 capabilities of having multi-turn dialogues. In: Findings of the Association for Computational Linguistics: NAACL 2024, pp 3184\u20133200. https:\/\/doi.org\/10.18653\/v1\/2024.findings-naacl.201","DOI":"10.18653\/v1\/2024.findings-naacl.201"},{"issue":"8","key":"2432_CR47","first-page":"242","volume":"22","author":"AE Elo","year":"1967","unstructured":"Elo AE (1967) The proposed uscf rating system, its development, theory, and applications. Chess Life 22(8):242\u2013247","journal-title":"Chess Life"},{"key":"2432_CR48","first-page":"27730","volume":"35","author":"L Ouyang","year":"2022","unstructured":"Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A (2022) Training language models to follow instructions with human feedback. Adv Neural Inf Process Syst 35:27730\u201327744","journal-title":"Adv Neural Inf Process Syst"},{"key":"2432_CR49","unstructured":"Xiao S, Liu Z, Zhang P, Muennighof N (2023) C-pack: packaged resources to advance general chinese embedding. arXiv:2309.07597"},{"key":"2432_CR50","unstructured":"Li Z, Zhang X, Zhang Y, Long D, Xie P, Zhang M (2023) Towards general text embeddings with multi-stage contrastive learning. arXiv:2308.03281"}],"container-title":["International Journal of Machine Learning and Cybernetics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13042-024-02432-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s13042-024-02432-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13042-024-02432-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,7]],"date-time":"2025-06-07T04:29:52Z","timestamp":1749270592000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s13042-024-02432-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,11]]},"references-count":50,"journal-issue":{"issue":"5-6","published-print":{"date-parts":[[2025,6]]}},"alternative-id":["2432"],"URL":"https:\/\/doi.org\/10.1007\/s13042-024-02432-9","relation":{},"ISSN":["1868-8071","1868-808X"],"issn-type":[{"value":"1868-8071","type":"print"},{"value":"1868-808X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,11]]},"assertion":[{"value":"27 June 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 October 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 November 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}