{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T08:07:41Z","timestamp":1782202061572,"version":"3.54.5"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T00:00:00Z","timestamp":1764892800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T00:00:00Z","timestamp":1764892800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100006359","name":"Blekinge Institute of Technology","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100006359","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Empir Software Eng"],"published-print":{"date-parts":[[2026,3]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Context<\/jats:title>\n                    <jats:p>Generative AI (GenAI) is increasingly adopted in software development for tasks such as document generation, data analysis, and code generation. However, evaluating the quality of GenAI applications becomes challenging, as traditional quality measurements may not be fully applicable.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Objective<\/jats:title>\n                    <jats:p>In this study, we explore how practitioners evaluate the quality of GenAI applications and investigate quality evaluation techniques.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Method<\/jats:title>\n                    <jats:p>We conducted a multi-case study in three industrial projects from software development companies. We examined four GenAI application domains: document generation, data analysis and insight generation, customer service, and code generation. Data were collected through three workshops and 23 semi-structured interviews with industrial practitioners.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>We identified fourteen GenAI use cases and 28 metrics currently used to evaluate the quality of GenAI applications\u2019 outputs. We synthesized the identified metrics\u2019 usage patterns and challenges based on the collected data.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions<\/jats:title>\n                    <jats:p>This study presents practical insights into using metrics to measure GenAI-based system qualities in real industrial settings. Our findings indicate that practitioners use custom-built and context-specific metrics; combining these with academic metrics can strengthen GenAI system quality evaluation.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1007\/s10664-025-10759-2","type":"journal-article","created":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T06:49:21Z","timestamp":1764917361000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Evaluating the quality of GenAI applications in software engineering: a multi-case study"],"prefix":"10.1007","volume":"31","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5949-1375","authenticated-orcid":false,"given":"Liang","family":"Yu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Emil","family":"Al\u00e9groth","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Panagiota","family":"Chatzipetrou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tony","family":"Gorschek","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,12,5]]},"reference":[{"key":"10759_CR1","unstructured":"Allow J, Neustadt I (1999) Uml 2 and the unified process. Reading: Addison Wesley"},{"key":"10759_CR2","unstructured":"Banerjee S, Lavie A (2005) Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and\/or summarization, pp 65\u201372"},{"key":"10759_CR3","unstructured":"Bommasani R, Hudson DA, Adeli E, Altman R, Arora S, von Arx S, Bernstein MS, Bohg J, Bosselut A, Brunskill E et\u00a0al (2021) On the opportunities and risks of foundation models. arXiv:2108.07258"},{"key":"10759_CR4","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A et al (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"10759_CR5","unstructured":"Celikyilmaz A, Clark E, Gao J (2020) Evaluation of text generation: a survey. arXiv:2006.14799"},{"key":"10759_CR6","unstructured":"Chen M, Tworek J, Jun H, Yuan Q, Pinto HPDO, Kaplan J, Edwards H, Burda Y, Joseph N, Brockman G et\u00a0al (2021) Evaluating large language models trained on code. arXiv:2107.03374"},{"issue":"240","key":"10759_CR7","first-page":"1","volume":"24","author":"A Chowdhery","year":"2023","unstructured":"Chowdhery A, Narang S, Devlin J, Bosma M, Mishra G, Roberts A, Barham P, Chung HW, Sutton C, Gehrmann S et al (2023) Palm: Scaling language modeling with pathways. J Mach Learn Res 24(240):1\u2013113","journal-title":"J Mach Learn Res"},{"key":"10759_CR8","doi-asserted-by":"crossref","unstructured":"Cruzes DS, Dyba T (2011) Recommended steps for thematic synthesis in software engineering. In: 2011 International symposium on empirical software engineering and measurement. IEEE, pp 275\u2013284","DOI":"10.1109\/ESEM.2011.36"},{"key":"10759_CR9","unstructured":"Eddine MK, Shang G, Tixier AJP, Vazirgiannis M (2021) Frugalscore: Learning cheaper, lighter and faster evaluation metricsfor automatic text generation. arXiv:2110.08559"},{"key":"10759_CR10","doi-asserted-by":"crossref","unstructured":"Esposito M, Li X, Moreschini S, Ahmad N, Cerny T, Vaidhyanathan K, Lenarduzzi V, Taibi D (2025) Generative ai for software architecture. applications, trends, challenges, and future directions. arXiv:2503.13310","DOI":"10.2139\/ssrn.5196419"},{"key":"10759_CR11","doi-asserted-by":"publisher","first-page":"391","DOI":"10.1162\/tacl_a_00373","volume":"9","author":"AR Fabbri","year":"2021","unstructured":"Fabbri AR, Kry\u015bci\u0144ski W, McCann B, Xiong C, Socher R, Radev D (2021) Summeval: Re-evaluating summarization evaluation. Trans Assoc Comput Linguist 9:391\u2013409","journal-title":"Trans Assoc Comput Linguist"},{"key":"10759_CR12","doi-asserted-by":"crossref","unstructured":"Feldt R, de\u00a0Oliveira\u00a0Neto FG, Torkar R (2018) Ways of applying artificial intelligence in software engineering. In: Proceedings of the 6th International workshop on realizing artificial intelligence synergies in software engineering, pp 35\u201341","DOI":"10.1145\/3194104.3194109"},{"key":"10759_CR13","doi-asserted-by":"crossref","unstructured":"Honovich O, Aharoni R, Herzig J, Taitelbaum H, Kukliansy D, Cohen V, Scialom T, Szpektor I, Hassidim A, Matias Y (2022) True: Re-evaluating factual consistency evaluation. arXiv:2204.04991","DOI":"10.18653\/v1\/2022.naacl-main.287"},{"key":"10759_CR14","doi-asserted-by":"crossref","unstructured":"Hou X, Zhao Y, Liu Y, Yang Z, Wang K, Li L, Luo X, Lo D, Grundy J, Wang H (2023) Large language models for software engineering: a systematic literature review. ACM Trans Softw Eng Methodol","DOI":"10.1145\/3695988"},{"key":"10759_CR15","unstructured":"Huang D, Bu Q, Zhang JM, Luck M, Cui H (2023) Agentcoder: Multi-agent-based code generation with iterative testing and optimisation. arXiv:2312.13010"},{"key":"10759_CR16","unstructured":"ISO (2016) ISO\/IEC-25023: systems and software engineering: systems and software quality requirements and evaluation (SQuaRE): measurement of system and software product quality"},{"key":"10759_CR17","unstructured":"Jiang D, Ku M, Li T, Ni Y, Sun S, Fan R, Chen W (2024) Genai arena: An open evaluation platform for generative models. arXiv:2406.04485"},{"key":"10759_CR18","doi-asserted-by":"crossref","unstructured":"Kenthapadi K, Sameki M, Taly A (2024) Grounding and evaluation for large language models: practical challenges and lessons learned (survey). In: Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining, pp 6523\u20136533","DOI":"10.1145\/3637528.3671467"},{"key":"10759_CR19","unstructured":"Kynk\u00e4\u00e4nniemi T, Karras T, Laine S, Lehtinen J, Aila T (2019) Improved precision and recall metric for assessing generative models. Adv Neural Inf Process Syst 32"},{"key":"10759_CR20","doi-asserted-by":"crossref","unstructured":"Lago P, Runeson P, Song Q, Verdecchia R (2024) Threats to validity in software engineering\u2013hypocritical paper section or essential analysis? In: Proceedings of the 18th ACM\/IEEE International symposium on empirical software engineering and measurement, pp 314\u2013324","DOI":"10.1145\/3674805.3686691"},{"key":"10759_CR21","unstructured":"Lin, CY (2004) Rouge: A package for automatic evaluation of summaries. In: Text summarization branches out, pp 74\u201381"},{"key":"10759_CR22","doi-asserted-by":"crossref","unstructured":"Lo CK (2019) Yisi-a unified semantic mt quality evaluation and estimation metric for languages with different levels of available resources. In: Proceedings of the fourth conference on machine translation (Volume 2: Shared Task Papers, Day 1), pp 507\u2013513","DOI":"10.18653\/v1\/W19-5358"},{"key":"10759_CR23","doi-asserted-by":"crossref","unstructured":"Min S, Krishna K, Lyu X, Lewis M, Yih WT, Koh PW, Iyyer M, Zettlemoyer L, Hajishirzi H (2023) Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. arXiv:2305.14251","DOI":"10.18653\/v1\/2023.emnlp-main.741"},{"key":"10759_CR24","doi-asserted-by":"crossref","unstructured":"Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B, Spitzer E, Raji ID, Gebru T (2019) Model cards for model reporting. In: Proceedings of the conference on fairness, accountability, and transparency, pp 220\u2013229","DOI":"10.1145\/3287560.3287596"},{"key":"10759_CR25","first-page":"27730","volume":"35","author":"L Ouyang","year":"2022","unstructured":"Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A et al (2022) Training language models to follow instructions with human feedback. Adv Neural Inf Process Syst 35:27730\u201327744","journal-title":"Adv Neural Inf Process Syst"},{"key":"10759_CR26","doi-asserted-by":"crossref","unstructured":"Papineni K, Roukos S, Ward T, Zhu WJ (2002) Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp 311\u2013318","DOI":"10.3115\/1073083.1073135"},{"key":"10759_CR27","unstructured":"Patton\u00a0Quinn M (2002) Qualitative research & evaluation methods"},{"key":"10759_CR28","doi-asserted-by":"crossref","unstructured":"Petersen K, Wohlin C (2009) Context in industrial software engineering research. In: 2009 3rd International symposium on empirical software engineering and measurement. IEEE, pp 401\u2013404","DOI":"10.1109\/ESEM.2009.5316010"},{"key":"10759_CR29","doi-asserted-by":"publisher","first-page":"654","DOI":"10.1007\/s10664-010-9136-6","volume":"15","author":"K Petersen","year":"2010","unstructured":"Petersen K, Wohlin C (2010) The effect of moving from a plan-driven to an incremental software development approach with agile practices: An industrial case study. Empir Softw Eng 15:654\u2013693","journal-title":"Empir Softw Eng"},{"key":"10759_CR30","doi-asserted-by":"crossref","unstructured":"Raji ID, Smart A, White RN, Mitchell M, Gebru T, Hutchinson B, Smith-Loud J, Theron D, Barnes P (2020) Closing the ai accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In: Proceedings of the 2020 conference on fairness, accountability, and transparency, pp 33\u201344","DOI":"10.1145\/3351095.3372873"},{"key":"10759_CR31","doi-asserted-by":"publisher","first-page":"103609","DOI":"10.1016\/j.cad.2023.103609","volume":"165","author":"L Regenwetter","year":"2023","unstructured":"Regenwetter L, Srivastava A, Gutfreund D, Ahmed F (2023) Beyond statistical similarity: rethinking metrics for deep generative models in engineering design. Comput Aided Des 165:103609","journal-title":"Comput Aided Des"},{"key":"10759_CR32","doi-asserted-by":"crossref","unstructured":"Rei R, Stewart C, Farinha AC, Lavie A (2020) Comet: A neural framework for mt evaluation. arXiv:2009.09025","DOI":"10.18653\/v1\/2020.emnlp-main.213"},{"key":"10759_CR33","unstructured":"Ren S, Guo D, Lu S, Zhou L, Liu S, Tang D, Sundaresan N, Zhou M, Blanco A, Ma S (2020) Codebleu: a method for automatic evaluation of code synthesis. arXiv:2009.10297"},{"issue":"2","key":"10759_CR34","doi-asserted-by":"publisher","first-page":"131","DOI":"10.1007\/s10664-008-9102-8","volume":"14","author":"P Runeson","year":"2009","unstructured":"Runeson P, H\u00f6st M (2009) Guidelines for conducting and reporting case study research in software engineering. Empir Softw Eng 14(2):131","journal-title":"Empir Softw Eng"},{"key":"10759_CR35","doi-asserted-by":"crossref","unstructured":"Sellam T, Das D, Parikh AP (2020) Bleurt: Learning robust metrics for text generation. arXiv:2004.04696","DOI":"10.18653\/v1\/2020.acl-main.704"},{"key":"10759_CR36","doi-asserted-by":"crossref","unstructured":"Wang A, Cho K, Lewis M (2020) Asking and answering questions to evaluate the factual consistency of summaries. arXiv:2004.04228","DOI":"10.18653\/v1\/2020.acl-main.450"},{"key":"10759_CR37","doi-asserted-by":"crossref","unstructured":"Wohlin C, Runeson P, H\u00f6st M, Ohlsson MC, Regnell B, Wessl\u00e9n A et\u00a0al (2012) Experimentation in software engineering, vol 236. Springer","DOI":"10.1007\/978-3-642-29044-2"},{"key":"10759_CR38","doi-asserted-by":"crossref","unstructured":"Wohlin C (2014) Guidelines for snowballing in systematic literature studies and a replication in software engineering. In: Proceedings of the 18th international conference on evaluation and assessment in software engineering, pp 1\u201310","DOI":"10.1145\/2601248.2601268"},{"key":"10759_CR39","doi-asserted-by":"publisher","first-page":"401","DOI":"10.1162\/tacl_a_00107","volume":"4","author":"W Xu","year":"2016","unstructured":"Xu W, Napoles C, Pavlick E, Chen Q, Callison-Burch C (2016) Optimizing statistical machine translation for text simplification. Trans Assoc Comput Linguist 4:401\u2013415","journal-title":"Trans Assoc Comput Linguist"},{"key":"10759_CR40","doi-asserted-by":"crossref","unstructured":"Yu L, Al\u00e9groth E, Chatzipetrou P, Gorschek T (2025) Measuring the quality of generative ai systems: mapping metrics to quality characteristics - snowballing literature review. Inf Softw Technol 107802","DOI":"10.1016\/j.infsof.2025.107802"},{"key":"10759_CR41","unstructured":"Zhang Z, Chen C, Liu B, Liao C, Gong Z, Yu H, Li J, Wang R (2023) A survey on language models for code. arXiv:2311.07989"},{"key":"10759_CR42","unstructured":"Zhang T, Kishore V, Wu F, Weinberger KQ, Artzi Y (2019) Bertscore: Evaluating text generation with bert. arXiv:1904.09675"},{"key":"10759_CR43","doi-asserted-by":"crossref","unstructured":"Zhao W, Peyrard M, Liu F, Gao Y, Meyer CM, Eger S (2019) Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance. arXiv:1909.02622","DOI":"10.18653\/v1\/D19-1053"}],"container-title":["Empirical Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-025-10759-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10664-025-10759-2","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-025-10759-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,14]],"date-time":"2026-01-14T04:33:08Z","timestamp":1768365188000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10664-025-10759-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,5]]},"references-count":43,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,3]]}},"alternative-id":["10759"],"URL":"https:\/\/doi.org\/10.1007\/s10664-025-10759-2","relation":{},"ISSN":["1382-3256","1573-7616"],"issn-type":[{"value":"1382-3256","type":"print"},{"value":"1573-7616","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,5]]},"assertion":[{"value":"10 January 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 October 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 December 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declared that they have no competing interests that are relevant to the content of this article. I would like to declare on behalf of my co-authors that the work described was original research that has not been published previously and is not under consideration for publication elsewhere. All the authors listed have approved the manuscript.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of Interest"}}],"article-number":"29"}}