{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T13:24:26Z","timestamp":1783344266547,"version":"3.54.6"},"reference-count":67,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2025,8,11]],"date-time":"2025-08-11T00:00:00Z","timestamp":1754870400000},"content-version":"vor","delay-in-days":222,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2025,1,1]],"date-time":"2025-01-01T00:00:00Z","timestamp":1735689600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["International Journal of Intelligent Systems"],"published-print":{"date-parts":[[2025,1]]},"abstract":"<jats:p>The emergence of large language models (LLMs) has substantially changed the artificial intelligence field, enabling its wide use over different domains. As various LLM alternatives have been developed, the current study proposes a novel decision\u2010support framework for evaluating and benchmarking LLMs based on multicriteria decision\u2010making (MCDM) techniques. In the proposed framework, an improved version of the best\u2010worst method (BWM) is proposed to effectively reduce the computational complexity of assigning a critical weight for the evaluation criteria of LLMs. Then, the improved BWM is integrated with the combined compromise solution (CoCoSo) method for ranking LLM alternatives. Findings show that the improved BWM successfully computes the criteria weights with low computational complexity compared to the original BWM. According to the enhanced BWM, the \u2018factual errors\u2019 criterion received the highest significant weight (0.2681), while the \u2018logical inconsistencies\u2019 criteria obtained the lowest (0.0827). The rest of the criteria were distributed in between that range. Subsequently, CoCoSo ranked the involved LLM alternatives in two different runs based on the extracted weights. Sensitivity analysis was employed to evaluate the effect of the assessment criteria on LLMs\u2019 evaluation.<\/jats:p>","DOI":"10.1155\/int\/2376097","type":"journal-article","created":{"date-parts":[[2025,8,11]],"date-time":"2025-08-11T10:29:30Z","timestamp":1754908170000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["An Improved Best\u2010Worst Method Integrated With Combined Compromise Solution for Evaluating Large Language Models"],"prefix":"10.1155","volume":"2025","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7844-3990","authenticated-orcid":false,"given":"O. S.","family":"Albahri","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7286-0892","authenticated-orcid":false,"given":"M. A.","family":"Alsalem","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3335-457X","authenticated-orcid":false,"given":"A. S.","family":"Albahri","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8333-5575","authenticated-orcid":false,"given":"Moamin","family":"A. Mahmoud","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7296-5413","authenticated-orcid":false,"given":"Laith","family":"Alzubaidi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4393-5570","authenticated-orcid":false,"given":"A. H.","family":"Alamoodi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0878-5696","authenticated-orcid":false,"given":"Iman","family":"Mohamad Sharaf","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2025,8,11]]},"reference":[{"key":"e_1_2_13_1_2","doi-asserted-by":"publisher","DOI":"10.1055\/s-0043-1774399"},{"key":"e_1_2_13_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.nlp.2023.100024"},{"key":"e_1_2_13_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imu.2023.101304"},{"key":"e_1_2_13_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpi.2023.100338"},{"key":"e_1_2_13_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664930"},{"key":"e_1_2_13_6_2","first-page":"2023","article-title":"The Utility of Chatgpt as an Example of Large Language Models in Healthcare Education, Research and Practice: Systematic Review on the Future Perspectives and Potential Limitations","author":"Sallam M.","year":"2023","journal-title":"medRxiv"},{"key":"e_1_2_13_7_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2024.103809"},{"key":"e_1_2_13_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.metrad.2023.100017"},{"key":"e_1_2_13_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.iswa.2024.200431"},{"key":"e_1_2_13_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12599-023-00795-x"},{"key":"e_1_2_13_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10916-024-02090-y"},{"key":"e_1_2_13_12_2","doi-asserted-by":"publisher","DOI":"10.58496\/adsa\/2025\/003"},{"key":"e_1_2_13_13_2","unstructured":"LiangP. BommasaniR. LeeT.et al. Holistic Evaluation of Language Models 2022 https:\/\/arxiv.org\/abs\/2211.09110."},{"key":"e_1_2_13_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10489-019-01532-2"},{"key":"e_1_2_13_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12553-020-00451-4"},{"key":"e_1_2_13_16_2","doi-asserted-by":"publisher","DOI":"10.1142\/s021962202050042x"},{"key":"e_1_2_13_17_2","doi-asserted-by":"publisher","DOI":"10.1186\/s40854-021-00256-y"},{"key":"e_1_2_13_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/s40747-022-00689-7"},{"key":"e_1_2_13_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/access.2020.2994746"},{"key":"e_1_2_13_20_2","doi-asserted-by":"publisher","DOI":"10.58496\/adsa\/2023\/002"},{"key":"e_1_2_13_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/s40815-021-01246-z"},{"key":"e_1_2_13_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12652-021-03325-3"},{"key":"e_1_2_13_23_2","doi-asserted-by":"publisher","DOI":"10.1142\/s0219622021500127"},{"key":"e_1_2_13_24_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.120708"},{"key":"e_1_2_13_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12652-021-02897-4"},{"key":"e_1_2_13_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10916-019-1338-x"},{"key":"e_1_2_13_27_2","doi-asserted-by":"publisher","DOI":"10.3390\/math10010071"},{"key":"e_1_2_13_28_2","doi-asserted-by":"publisher","DOI":"10.1155\/2020\/1761893"},{"key":"e_1_2_13_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11356-021-13832-7"},{"key":"e_1_2_13_30_2","doi-asserted-by":"publisher","DOI":"10.3390\/su142416921"},{"key":"e_1_2_13_31_2","doi-asserted-by":"publisher","DOI":"10.31181\/dmame12012023b"},{"key":"e_1_2_13_32_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0283655"},{"key":"e_1_2_13_33_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.orp.2022.100263"},{"key":"e_1_2_13_34_2","doi-asserted-by":"publisher","DOI":"10.3390\/buildings12050655"},{"key":"e_1_2_13_35_2","doi-asserted-by":"publisher","DOI":"10.1108\/md-05-2017-0458"},{"key":"e_1_2_13_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.sftr.2025.100822"},{"key":"e_1_2_13_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksuci.2023.101714"},{"key":"e_1_2_13_38_2","doi-asserted-by":"publisher","DOI":"10.3390\/axioms12080729"},{"key":"e_1_2_13_39_2","doi-asserted-by":"publisher","DOI":"10.17270\/j.log.2023.857"},{"key":"e_1_2_13_40_2","doi-asserted-by":"publisher","DOI":"10.3389\/fenvs.2022.1088064"},{"key":"e_1_2_13_41_2","unstructured":"PengC. YangX. ChenA.et al. A Study of Generative Large Language Model for Medical Research and Healthcare 2023 https:\/\/arxiv.org\/abs\/2305.13523."},{"key":"e_1_2_13_42_2","article-title":"Performance Analysis of Banking by Using the Multi Criteria Decision Making Method: SECA. International Theory, Research and Reviews in Social","volume":"2023","author":"Yilmaz N.","year":"2023","journal-title":"Human and Administrative Sciences"},{"key":"e_1_2_13_43_2","unstructured":"WorkshopB. BigScienceT. TayY. ScaoT. L. PavlickE. andShusterK. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model 2022 https:\/\/arxiv.org\/abs\/2211.05100."},{"key":"e_1_2_13_44_2","doi-asserted-by":"crossref","unstructured":"ParkG. KimH. ZhouH. andLeeK. Automated Extraction of Molecular Interactions and Pathway Knowledge Using Large Language Model Galactica: Opportunities and Challenges Proceedings of the 22nd Workshop on Biomedical Natural Language Processing and BioNLP Shared Tasks 2023 Toronto Canada.","DOI":"10.18653\/v1\/2023.bionlp-1.22"},{"key":"e_1_2_13_45_2","doi-asserted-by":"crossref","unstructured":"SamsiS. GuptaU. SeznecM. andRagan-KelleyJ. From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference Proceedings of the 2023 IEEE High Performance Extreme Computing Conference (HPEC) 2023 Boston IEEE.","DOI":"10.1109\/HPEC58863.2023.10363447"},{"key":"e_1_2_13_46_2","unstructured":"ZhangS. RollerS. GoyalN. ArtetxeM. ChenM. andChenS. OPT: Open Pre-Trained Transformer Language Models 2022 https:\/\/arxiv.org\/abs\/2205.01068."},{"key":"e_1_2_13_47_2","doi-asserted-by":"crossref","unstructured":"GilbertT. K. KruegerD. RaeJ. andOsbandI. Reward Reports for Reinforcement Learning Proceedings of the 2023 AAAI\/ACM Conference on AI Ethics and Society 2023 Montr\u00e9al Canada.","DOI":"10.1145\/3600211.3604698"},{"key":"e_1_2_13_48_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.tourman.2017.11.009"},{"key":"e_1_2_13_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.seps.2019.02.004"},{"key":"e_1_2_13_50_2","doi-asserted-by":"publisher","DOI":"10.3390\/su10082817"},{"key":"e_1_2_13_51_2","doi-asserted-by":"publisher","DOI":"10.3390\/su10072371"},{"key":"e_1_2_13_52_2","doi-asserted-by":"publisher","DOI":"10.3390\/su10051626"},{"key":"e_1_2_13_53_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.omega.2014.11.009"},{"key":"e_1_2_13_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/bf01585569"},{"key":"e_1_2_13_55_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.omega.2015.12.001"},{"key":"e_1_2_13_56_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.trip.2020.100240"},{"key":"e_1_2_13_57_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.rser.2018.12.035"},{"key":"e_1_2_13_58_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ifacol.2018.08.218"},{"key":"e_1_2_13_59_2","doi-asserted-by":"publisher","DOI":"10.3390\/math8081342"},{"key":"e_1_2_13_60_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cie.2018.09.011"},{"key":"e_1_2_13_61_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dajour.2023.100378"},{"key":"e_1_2_13_62_2","article-title":"A Barrier Evaluation Framework for Forest Carbon Sink Project Implementation in China Using an Integrated BWM-IT2F-PROMETHEE II Method","volume":"2023","author":"Wei Q.","year":"2023","journal-title":"Expert Systems with Applications"},{"key":"e_1_2_13_63_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2024.108023"},{"key":"e_1_2_13_64_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.123498"},{"key":"e_1_2_13_65_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.121420"},{"key":"e_1_2_13_66_2","doi-asserted-by":"publisher","DOI":"10.1007\/s40815-023-01597-9"},{"key":"e_1_2_13_67_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.aei.2023.102191"}],"container-title":["International Journal of Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1155\/int\/2376097","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1155\/int\/2376097","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1155\/int\/2376097","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,8]],"date-time":"2026-03-08T18:05:29Z","timestamp":1772993129000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1155\/int\/2376097"}},"subtitle":[],"editor":[{"given":"Richard","family":"Murray","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2025,1]]},"references-count":67,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1]]}},"alternative-id":["10.1155\/int\/2376097"],"URL":"https:\/\/doi.org\/10.1155\/int\/2376097","archive":["Portico"],"relation":{},"ISSN":["0884-8173","1098-111X"],"issn-type":[{"value":"0884-8173","type":"print"},{"value":"1098-111X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1]]},"assertion":[{"value":"2024-03-20","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-09","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"2376097"}}