{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T10:52:22Z","timestamp":1781693542468,"version":"3.54.5"},"reference-count":27,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T00:00:00Z","timestamp":1781136000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MAKE"],"abstract":"<jats:p>Large language models (LLMs) increasingly require robust evaluation under realistic instruction-following conditions, particularly for fine-tuned task-specific adapters operating in multilingual environments. This study proposes a scenario-adaptive evaluation framework for assessing the reliability of fine-tuned text models across two application regimes: misinformation detection (disinfo) and knowledge-grounded factual biography generation (heroes). The framework integrates automated generation of balanced risk-oriented scenarios, bilingual evaluation in English and Ukrainian, the LLM-as-a-Judge paradigm, and multidimensional robustness analysis through the Alignment Robustness Index (ARI). Six LoRA-adapted models based on Qwen2.5-3B-Instruct, SmolLM2-1.7B-Instruct, and TinyLlama-1.1B-Chat-v1.0 were evaluated. The implemented pipeline generated 2052 scenarios and 6156 model responses, producing a final bilingual analytical subset of 4104 judged records. Experimental results show that task-specific adaptation produces task-dependent robustness profiles. In the disinfo case, Qwen2.5-3B achieved the strongest overall performance, combining the highest safety and classification accuracy. In contrast, the heroes case revealed a more compressed and multidimensional vulnerability space without a single dominant model. The results further demonstrate the importance of multilingual evaluation, as weaker adapters exhibited more pronounced cross-lingual safety gaps. Overall, the framework provides a reproducible and practically applicable methodology for evaluating fine-tuned language models under imperfect instruction conditions.<\/jats:p>","DOI":"10.3390\/make8060161","type":"journal-article","created":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T01:52:08Z","timestamp":1781229128000},"page":"161","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Scenario-Adaptive Evaluation of Trustworthy Fine-Tuned Text Models Across Knowledge-Grounded Generation and Misinformation Detection"],"prefix":"10.3390","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2441-6292","authenticated-orcid":false,"given":"Khrystyna","family":"Lipianina-Honcharenko","sequence":"first","affiliation":[{"name":"Department of Information Computer Systems and Control, West Ukrainian National University, 11 Lvivska Str., 46009 Ternopil, Ukraine"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5705-5702","authenticated-orcid":false,"given":"Pavlo","family":"Bykovyy","sequence":"additional","affiliation":[{"name":"Department of Information Computer Systems and Control, West Ukrainian National University, 11 Lvivska Str., 46009 Ternopil, Ukraine"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andriy","family":"Krysovatyy","sequence":"additional","affiliation":[{"name":"S. I. Yuriy Department of Finance, West Ukrainian National University, 11 Lvivska Str., 46009 Ternopil, Ukraine"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6541-0359","authenticated-orcid":false,"given":"Myroslav","family":"Komar","sequence":"additional","affiliation":[{"name":"Department of Information Computer Systems and Control, West Ukrainian National University, 11 Lvivska Str., 46009 Ternopil, Ukraine"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Borys","family":"Yazlyuk","sequence":"additional","affiliation":[{"name":"Department of Economic Expertise and Land Management, West Ukrainian National University, 11 Lvivska Str., 46009 Ternopil, Ukraine"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,6,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Huang, H., Bu, X., Zhou, H., Qu, Y., Liu, J., Yang, M., Xu, B., and Zhao, T. (2024). An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-Tuned Judge Model Is Not a General Substitute for GPT-4. arXiv.","DOI":"10.18653\/v1\/2025.findings-acl.306"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Shi, L., Ma, C., Liang, W., Diao, X., Ma, W., and Vosoughi, S. (2024). Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. arXiv.","DOI":"10.18653\/v1\/2025.ijcnlp-long.18"},{"key":"ref_3","unstructured":"Bolukbasi, T., Chang, K.W., Zou, J., Saligrama, V., and Kalai, A. (2016). Man is to computer programmer as woman is to homemaker? Debiasing word embeddings. Adv. Neural Inf. Process. Syst., 29."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1126\/science.aal4230","article-title":"Semantics derived automatically from language corpora contain human-like biases","volume":"356","author":"Caliskan","year":"2017","journal-title":"Science"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhao, J., Wang, T., Yatskar, M., Ordonez, V., and Chang, K.W. (2017, January 7\u201311). Men also like shopping: Reducing gender bias amplification using corpus-level constraints. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Copenhagen, Denmark.","DOI":"10.18653\/v1\/D17-1323"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Sheng, E., Chang, K.W., Natarajan, P., and Peng, N. (2019, January 3\u20137). The woman worked as a babysitter: On biases in language generation. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China.","DOI":"10.18653\/v1\/D19-1339"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N.A. (2020). RealToxicityPrompts: Evaluating neural toxic degeneration in language models. Findings of the Association for Computational Linguistics: ACL 2020, Online.","DOI":"10.18653\/v1\/2020.findings-emnlp.301"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lin, S., Hilton, J., and Evans, O. (2022, January 22\u201327). TruthfulQA: Measuring how models mimic human falsehoods. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), Dublin, Ireland.","DOI":"10.18653\/v1\/2022.acl-long.229"},{"key":"ref_9","unstructured":"Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., and Kumar, A. (2022). Holistic evaluation of language models (HELM). Trans. Mach. Learn. Res."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Lei, L., Wu, L., Sun, R., Huang, Y., Long, C., Liu, X., Lei, X., Tang, J., and Huang, M. (2023). SafetyBench: Evaluating the Safety of Large Language Models. arXiv.","DOI":"10.18653\/v1\/2024.acl-long.830"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zheng, L., Chiang, W.L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., and Xing, E. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems (NeurIPS), NeurIPS.","DOI":"10.52202\/075280-2020"},{"key":"ref_12","unstructured":"Li, H., Dong, Q., Chen, J., Su, H., Zhou, Y., Ai, Q., Ye, Z., and Liu, Y. (2024). LLMs-as-Judges: A Comprehensive Survey on LLM-Based Evaluation Methods. arXiv."},{"key":"ref_13","unstructured":"Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., and McKinnon, C. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv."},{"key":"ref_14","unstructured":"Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., and Ndousse, K. (2022). Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Nozza, D., Bianchi, F., and Hovy, D. (2021). HONEST: Measuring Hurtful Sentence Completion in Language Models. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2021.naacl-main.191"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wang, W., Tu, Z., Chen, C., Yuan, Y., Huang, J., Jiao, W., and Lyu, M. (2024). All Languages Matter: On the Multilingual Safety of LLMs. Findings of the Association for Computational Linguistics: ACL 2024, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.findings-acl.349"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"7659","DOI":"10.1609\/aaai.v34i05.6267","article-title":"On Measuring and Mitigating Biased Inferences of Word Embeddings","volume":"Volume 34","author":"Dev","year":"2020","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Nadeem, M., Bethke, A., and Reddy, S. (2021). StereoSet: Measuring stereotypical bias in pretrained language models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Association for Computational Linguistics.","DOI":"10.18653\/v1\/2021.acl-long.416"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Parrish, A., Chen, A., Nangia, N., Padmakumar, V., Phang, J., Thompson, J., Htut, P.M., and Bowman, S. (2022). BBQ: A hand-built bias benchmark for question answering. Findings of the Association for Computational Linguistics: ACL 2022, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2022.findings-acl.165"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Tan, X., Hansanti, P., Turkatenko, A., Chuang, J., Wood, C., Yu, B., Ropers, C., and Costa-juss\u00e0, M.R. (2025). Towards massive multilingual holistic bias. Proceedings of the 6th Workshop on Gender Bias in Natural Language Processing (GeBNLP), Association for Computational Linguistics.","DOI":"10.18653\/v1\/2025.gebnlp-1.35"},{"key":"ref_21","unstructured":"Fisher, R.A. (1925). Statistical Methods for Research Workers, Oliver and Boyd. Available online: https:\/\/psychclassics.yorku.ca\/Fisher\/Methods\/."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Hwang, C.L., and Yoon, K. (1981). Multiple Attribute Decision Making: Methods and Applications: A State-of-the-Art Survey, Springer. Lecture Notes in Economics and Mathematical Systems.","DOI":"10.1007\/978-3-642-48318-9"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"101","DOI":"10.1016\/j.trac.2009.09.009","article-title":"Sum of Ranking Differences Compares Methods or Models Fairly","volume":"29","year":"2010","journal-title":"TrAC Trends Anal. Chem."},{"key":"ref_24","first-page":"344649","article-title":"Sum of Euclidean Distance Differences and Sum of Absolute Manhattan Distance Differences: Multicriteria Decision Making Tools for Small Data Tables","volume":"1381","year":"2025","journal-title":"Anal. Chim. Acta"},{"key":"ref_25","unstructured":"(2026, May 07). Fake and Real News Dataset. Available online: https:\/\/www.kaggle.com\/datasets\/clmentbisaillon\/fake-and-real-news-dataset."},{"key":"ref_26","unstructured":"Lipianina-Honcharenko, K., Komar, M., Bykovyy, P., and Osolinskyi, O. (2026). Task-Specific LoRA-Adapted Language Models for Disinformation Detection and Factual Biography Generation. figshare."},{"key":"ref_27","unstructured":"(2026, May 07). Wikipedia Biographies Text Generation Dataset. Available online: https:\/\/www.kaggle.com\/datasets\/thedevastator\/wikipedia-biographies-text-generation-dataset."}],"container-title":["Machine Learning and Knowledge Extraction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-4990\/8\/6\/161\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T10:00:33Z","timestamp":1781690433000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-4990\/8\/6\/161"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,11]]},"references-count":27,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["make8060161"],"URL":"https:\/\/doi.org\/10.3390\/make8060161","relation":{},"ISSN":["2504-4990"],"issn-type":[{"value":"2504-4990","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,11]]}}}