{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T05:43:16Z","timestamp":1784698996647,"version":"3.55.0"},"reference-count":21,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,7,21]],"date-time":"2025-07-21T00:00:00Z","timestamp":1753056000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Big Data"],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>The adoption of Large Language Models (LLMs) in search systems necessitates new evaluation methodologies beyond traditional rule-based or manual approaches.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>We propose a general framework for evaluating structured outputs using LLMs, focusing on search query parsing within an online classified platform. Our approach leverages LLMs' contextual reasoning capabilities through three evaluation methodologies: Pointwise, Pairwise, and Pass\/Fail assessments. Additionally, we introduce a Contextual Evaluation Prompt Routing strategy to improve reliability and reduce hallucinations.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>Experiments conducted on both small- and large-scale datasets demonstrate that LLM-based evaluation achieves approximately 90% agreement with human judgments.<\/jats:p><\/jats:sec><jats:sec><jats:title>Discussion<\/jats:title><jats:p>These results validate LLM-driven evaluation as a scalable, interpretable, and effective alternative to traditional evaluation methods, providing robust query parsing for real-world search systems.<\/jats:p><\/jats:sec>","DOI":"10.3389\/fdata.2025.1611389","type":"journal-article","created":{"date-parts":[[2025,7,21]],"date-time":"2025-07-21T11:47:36Z","timestamp":1753098456000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["LLM-as-a-Judge: automated evaluation of search query parsing using large language models"],"prefix":"10.3389","volume":"8","author":[{"given":"Mehmet Selman","family":"Baysan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Serkan","family":"Uysal","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"\u0130rem","family":"\u0130\u015flek","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"\u00c7a\u011fla","family":"\u00c7\u0131\u011f Karaman","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tunga","family":"G\u00fcng\u00f6r","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2025,7,21]]},"reference":[{"key":"B1","first-page":"748","article-title":"\u201cSmatch: an evaluation metric for semantic feature structures,\u201d","volume-title":"Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Cai","year":"2013"},{"key":"B2","first-page":"8928","article-title":"\u201cA closer look into using large language models for automatic evaluation,\u201d","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023","author":"Chiang","year":"2023"},{"key":"B3","volume-title":"Statistics Without Maths for Psychology","author":"Dancey","year":"2007"},{"key":"B4","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2404.04475","article-title":"Length-controlled alpacaeval: a simple way to debias automatic evaluators","author":"Dubois","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B5","unstructured":"Sweeping heterogeneity with Smart MoPs: mixture of prompts for LLM task adaptation\n          \n          \n            \n              Dun\n              C.\n            \n            \n              Garcia\n              M. H.\n            \n            \n              Zheng\n              G.\n            \n            \n              Awadallah\n              A. H.\n            \n            \n              Kyrillidis\n              A.\n            \n            \n              Sim\n              R.\n            \n          \n          arXiv [Preprint].\n          \n          2025"},{"key":"B6","unstructured":"A survey on LLM-as-a-Judge\n          \n          \n            \n              Gu\n              J.\n            \n            \n              Jiang\n              X.\n            \n            \n              Shi\n              Z.\n            \n            \n              Tan\n              H.\n            \n            \n              Zhai\n              X.\n            \n            \n              Xu\n              C.\n            \n          \n          arXiv [Preprint].\n          \n          2025"},{"key":"B7","doi-asserted-by":"publisher","first-page":"1201","DOI":"10.3390\/sym16091201","article-title":"A survey of semantic parsing techniques","volume":"16","author":"Jiang","year":"2024","journal-title":"Symmetry"},{"key":"B8","unstructured":"\u201cDecomposed prompting: a modular approach for solving complex tasks,\u201d\n          \n          \n            \n              Khot\n              T.\n            \n            \n              Trivedi\n              H.\n            \n            \n              Finlayson\n              M.\n            \n            \n              Fu\n              Y.\n            \n            \n              Richardson\n              K.\n            \n            \n              Clark\n              P.\n            \n          \n          The Eleventh International Conference on Learning Representations\n          \n          2023"},{"key":"B9","doi-asserted-by":"publisher","DOI":"10.48500\/arXiv.2105.03317","article-title":"Toward code generation: a survey and lessons from semantic parsing","author":"Lee","year":"2021","journal-title":"arXiv [Preprint]."},{"key":"B10","doi-asserted-by":"crossref","first-page":"181","DOI":"10.18653\/v1\/2024.findings-emnlp.10","article-title":"\u201cUnleashing large language models' proficiency in zero-shot essay scoring,\u201d","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2024","author":"Lee","year":"2024"},{"key":"B11","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2412.05579","article-title":"Llms-as-judges: a comprehensive survey on LLM-based evaluation methods","author":"Li","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B12","author":"Li","year":"2024","journal-title":"From Crowdsourced Data to High-quality Benchmarks: Arena-hard and Benchbuilder Pipeline"},{"key":"B13","doi-asserted-by":"crossref","DOI":"10.1145\/3539618.3591858","article-title":"\u201cImplicit query parsing for product search,\u201d","volume-title":"SIGIR '23: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information","author":"Luo","year":"2023"},{"key":"B14","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2411.00142","article-title":"Judgerank: leveraging large language models for reasoning-intensive reranking","author":"Niu","year":"2024","journal-title":"arXiv [Preprint]."},{"key":"B15","first-page":"311","article-title":"\u201cBleu: a method for automatic evaluation of machine translation,\u201d","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni","year":"2002"},{"key":"B16","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2408.08808","article-title":"Constructing domain-specific evaluation sets for llm-as-a-judge","author":"Raju","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B17","doi-asserted-by":"publisher","first-page":"807","DOI":"10.5220\/0012394300003636","article-title":"\u201cEvaluating large language models in semantic parsing for conversational question answering over knowledge graphs,\u201d","author":"Schneider","year":"2024","journal-title":"Proceedings of the 16th International Conference on Agents and Artificial Intelligence, ICAART 2024, Volume 3, Rome, Italy, February 24\u201326, 2024"},{"key":"B18","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2412.12509","article-title":"Can you trust LLM judgments? Reliability of Llm-as-a-judge","author":"Schroeder","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B19","first-page":"1","article-title":"\u201cWho validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences,\u201d","volume-title":"Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, UIST 2024, Pittsburgh, PA, USA, October 13-16, 2024","author":"Shankar","year":"2024"},{"key":"B20","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv:2407.10999","article-title":"TALEC: teach your LLM to evaluate in specific domain with in-house criteria by criteria division and zero-shot plus few-shot","author":"Zhang","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B21","article-title":"\u201cJudging LLM-as-a-judge with mt-bench and chatbot arena,\u201d","volume-title":"Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023","author":"Zheng","year":"2023"}],"container-title":["Frontiers in Big Data"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdata.2025.1611389\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,21]],"date-time":"2025-07-21T11:47:39Z","timestamp":1753098459000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdata.2025.1611389\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,21]]},"references-count":21,"alternative-id":["10.3389\/fdata.2025.1611389"],"URL":"https:\/\/doi.org\/10.3389\/fdata.2025.1611389","relation":{},"ISSN":["2624-909X"],"issn-type":[{"value":"2624-909X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,21]]},"article-number":"1611389"}}