{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T15:58:49Z","timestamp":1784390329826,"version":"3.55.0"},"reference-count":207,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,1,31]]},"abstract":"<jats:p>Question answering (QA) systems are a leading and rapidly advancing field of natural language processing (NLP) research. One of their key advantages is that they enable more natural interactions between humans and machines, such as in virtual assistants or search engines. Over the past few decades, many QA systems have been developed to handle diverse QA tasks. However, the evaluation of these systems is intricate, as many of the available evaluation scores are not task-agnostic. Furthermore, translating human judgment into measurable metrics continues to be an open issue. These complexities add challenges to their assessment. This survey provides a systematic overview of evaluation scores and introduces a taxonomy with two main branches: Human-Centric Evaluation Scores (HCES) and Automatic Evaluation Scores (AES). Since many of these scores were originally designed for specific tasks but have been applied more generally, we also cover the basics of QA frameworks and core paradigms to provide a deeper understanding of their capabilities and limitations. Lastly, we discuss benchmark datasets that are critical for conducting systematic evaluations across various QA tasks.<\/jats:p>","DOI":"10.1145\/3744663","type":"journal-article","created":{"date-parts":[[2025,6,12]],"date-time":"2025-06-12T07:32:36Z","timestamp":1749713556000},"page":"1-43","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Evaluation of Question Answering Systems: Complexity of Judging a Natural Language"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7145-7137","authenticated-orcid":false,"given":"Amer","family":"Farea","sequence":"first","affiliation":[{"name":"Information Technology and Communication Sciences, Tampere University","place":["Tampere, Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6258-4783","authenticated-orcid":false,"given":"Zhen","family":"Yang","sequence":"additional","affiliation":[{"name":"Information Technology and Communication Sciences, Tampere University","place":["Tampere, Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-3508-0972","authenticated-orcid":false,"given":"Kien","family":"Duong","sequence":"additional","affiliation":[{"name":"Information Technology and Communication Sciences, Tampere University","place":["Tampere, Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9907-5939","authenticated-orcid":false,"given":"Nadeesha","family":"Perera","sequence":"additional","affiliation":[{"name":"Information Technology and Communication Sciences, Tampere University","place":["Tampere, Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0745-5641","authenticated-orcid":false,"given":"Frank","family":"Emmert-Streib","sequence":"additional","affiliation":[{"name":"Information Technology and Communication Sciences, Tampere University","place":["Tampere, Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,8,30]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1412"},{"key":"e_1_3_1_3_2","first-page":"213","volume-title":"Proceedings of the Actes de la 17e conf\u00e9rence sur le Traitement Automatique des Langues Naturelles. Articles courts","author":"Ali Husam","year":"2010","unstructured":"Husam Ali, Yllias Chali, and Sadid A Hasan. 2010. Automatic question generation from sentences. In Proceedings of the Actes de la 17e conf\u00e9rence sur le Traitement Automatique des Langues Naturelles. Articles courts. 213\u2013218."},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Tahani H. Alwaneen Aqil M. Azmi Hatim A. Aboalsamh Erik Cambria and Amir Hussain. 2022. Arabic question answering system: A survey. Artificial Intelligence Review 55 1 (2022) 207\u2013253.","DOI":"10.1007\/s10462-021-10031-1"},{"key":"e_1_3_1_5_2","volume-title":"Evaluating the Evaluators: Subjective Bias and Consistency in Human Evaluation of Natural Language Generation","author":"Amidei Jacopo","year":"2021","unstructured":"Jacopo Amidei. 2021. Evaluating the Evaluators: Subjective Bias and Consistency in Human Evaluation of Natural Language Generation. Open University (United Kingdom)."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1017\/s135132490000005x"},{"key":"e_1_3_1_7_2","unstructured":"Amanda Askell Yuntao Bai Anna Chen Dawn Drain Deep Ganguli Tom Henighan Andy Jones Nicholas Joseph Ben Mann Nova DasSarma et\u00a0al. 2021. A general language assistant as a laboratory for alignment. arXiv:2112.00861. Retrieved from https:\/\/arxiv.org\/abs\/2112.00861"},{"key":"e_1_3_1_8_2","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations ICLR 2015 San Diego CA USA May 7-9 2015 Conference Track Proceedings. http:\/\/arxiv.org\/abs\/1409.0473"},{"key":"e_1_3_1_9_2","unstructured":"Yushi Bai Jiahao Ying Yixin Cao Xin Lv Yuze He Xiaozhi Wang Jifan Yu Kaisheng Zeng Yijia Xiao Haozhe Lyu Jiayin Zhang Juanzi Li and Lei Hou. 2023. Benchmarking foundation models with language-model-as-an-examiner. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans LA USA) (NIPS\u201923). 1\u201326."},{"key":"e_1_3_1_10_2","unstructured":"Payal Bajaj Daniel Campos Nick Craswell Li Deng Jianfeng Gao Xiaodong Liu Rangan Majumder Andrew McNamara Bhaskar Mitra Tri Nguyen Mir Rosenberg Xia Song Alina Stoica Saurabh Tiwary and Tong Wang. 2018. MS MARCO: A human generated MAchine Reading COmprehension Dataset. arXiv:1611.09268. https:\/\/arxiv.org\/abs\/1611.09268"},{"key":"e_1_3_1_11_2","first-page":"65","volume-title":"Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization. 65\u201372."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","unstructured":"Yejin Bang Samuel Cahyawijaya Nayeon Lee Wenliang Dai Dan Su Bryan Wilie Holy Lovenia Ziwei Ji Tiezheng Yu Willy Chung Quyet V. Do Yan Xu and Pascale Fung. 2023. A multitask multilingual multimodal evaluation of ChatGPT on reasoning hallucination and interactivity. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics Nusa Dua Bali 675\u2013718. 10.18653\/v1\/2023.ijcnlp-main.45","DOI":"10.18653\/v1\/2023.ijcnlp-main.45"},{"issue":"22","key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"9977","DOI":"10.1073\/pnas.92.22.9977","article-title":"Models of natural language understanding","volume":"92","author":"Bates Madeleine","year":"1995","unstructured":"Madeleine Bates. 1995. Models of natural language understanding. Proceedings of the National Academy of Sciences 92, 22 (1995), 9977\u20139982.","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.5555\/1873738.1873743"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1186\/s12859-019-3119-4"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D13-1160"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","unstructured":"Jonathan Berant and Percy Liang. 2014. Semantic parsing via paraphrasing. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1415\u20131425. 10.3115\/v1\/p14-1133","DOI":"10.3115\/v1\/p14-1133"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","unstructured":"Yonatan Bisk Siva Reddy John Blitzer Julia Hockenmaier and Mark Steedman. 2016. Evaluating induced CCG parsers on grounded semantic parsing. In 2016 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics 2022\u20132027.","DOI":"10.18653\/v1\/D16-1214"},{"key":"e_1_3_1_19_2","unstructured":"Antoine Bordes Nicolas Usunier Sumit Chopra and Jason Weston. 2015. Large-scale simple question answering with memory networks. arXiv preprint arXiv:1506.02075 (2015)."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.5555\/176313.176316"},{"key":"e_1_3_1_21_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems 33 (2020), 1877\u20131901.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_22_2","first-page":"249","volume-title":"11th Conference of the European Chapter of the Association for Computational Linguistics","author":"Callison-Burch Chris","year":"2006","unstructured":"Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006. Re-evaluating the role of BLEU in machine translation research. In 11th Conference of the European Chapter of the Association for Computational Linguistics. 249\u2013256."},{"key":"e_1_3_1_23_2","unstructured":"Asli Celikyilmaz Elizabeth Clark and Jianfeng Gao. 2021. Evaluation of text generation: A survey. arXiv preprint arXiv:2006.14799 (2021)."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","unstructured":"Daniel Cer Yinfei Yang Sheng-yi Kong Nan Hua Nicole Limtiaco Rhomni St. John Noah Constant Mario Guajardo-Cespedes Steve Yuan Chris Tar Yun-Hsuan Sung Brian Strope and Ray Kurzweil. 2018. Universal sentence encoder. arXiv e-prints (2018). 10.48550\/arXiv.1803.11175","DOI":"10.48550\/arXiv.1803.11175"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-5817"},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Danqi Chen Adam Fisch Jason Weston and Antoine Bordes. 2017. Reading wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1870\u20131879.","DOI":"10.18653\/v1\/P17-1171"},{"key":"e_1_3_1_27_2","volume-title":"Proceedings of the 12th International AAAI Conference on Web and Social Media","author":"Chen Guanliang","year":"2018","unstructured":"Guanliang Chen, Jie Yang, Claudia Hauff, and Geert-Jan Houben. 2018. LearningQ: A large-scale dataset for educational question generation. In Proceedings of the 12th International AAAI Conference on Web and Social Media."},{"key":"e_1_3_1_28_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et\u00a0al. 2021. Evaluating large language models trained on code. https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_1_29_2","unstructured":"Kyunghyun Cho Bart van Merri\u00ebnboer \u00c7a\u011flar Gul\u00e7ehre Dzmitry Bahdanau Fethi Bougares Holger Schwenk and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder\u2013decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 1724\u20131734."},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"Woon Sang Cho Yizhe Zhang Sudha Rao Chris Brockett and Sungjin Lee. 2019. Generating a common question from multiple documents using multi-source encoder-decoder models. In Proceedings of the 3rd Workshop on Neural Generation and Translation. 32\u201343.","DOI":"10.18653\/v1\/D19-5604"},{"key":"e_1_3_1_31_2","doi-asserted-by":"crossref","unstructured":"Eunsol Choi He He Mohit Iyyer Mark Yatskar Wen-tau Yih Yejin Choi Percy Liang and Luke Zettlemoyer. 2018. QuAC: Question answering in context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2174\u20132184.","DOI":"10.18653\/v1\/D18-1241"},{"key":"e_1_3_1_32_2","unstructured":"Paul F. Christiano Jan Leike Tom B. Brown Miljan Martic Shane Legg and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Proceedings of the 31st International Conference on Neural Information Processing Systems. 4302\u20134310."},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","unstructured":"Yung-Sung Chuang Chi-Liang Liu Hung-yi Lee and Lin-shan Lee. 2020. SpeechBERT: An audio-and-text jointly learned language model for end-to-end spoken question answering. In Proc. Interspeech 2020. 4168\u20134172.","DOI":"10.21437\/Interspeech.2020-1570"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1264"},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","unstructured":"Alexis Conneau Douwe Kiela Holger Schwenk Lo\u00efc Barrault and Antoine Bordes. 2017. Supervised learning of universal sentence representations from natural language inference data. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 670\u2013680.","DOI":"10.18653\/v1\/D17-1070"},{"key":"e_1_3_1_36_2","first-page":"5408","volume-title":"Proceedings of the 12th Language Resources and Evaluation Conference","author":"Cortes Eduardo","year":"2020","unstructured":"Eduardo Cortes, Vinicius Woloszyn, Arne Binder, Tilo Himmelsbach, Dante Barone, and Sebastian M\u00f6ller. 2020. An empirical comparison of question classification methods for question answering systems. In Proceedings of the 12th Language Resources and Evaluation Conference. 5408\u20135416."},{"key":"e_1_3_1_37_2","unstructured":"Zihang Dai Lei Li and Wei Xu. 2016. CFO: Conditional focused neural question answering with large-scale knowledge bases. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 800\u2013810."},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","unstructured":"Hoa Dang Jimmy Lin and Diane Kelly. 2007. Overview of the TREC 2007 question answering track. In Trec 7 (2007) 63.","DOI":"10.6028\/NIST.SP.500-274.qa-overview"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-020-09866-x"},{"key":"e_1_3_1_40_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Volume 1 (Long and Short Papers). 4171\u20134186."},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","unstructured":"Kaustubh Dhole and Christopher D. Manning. 2020. [RETRACTED] Syn-QG: Syntactic and shallow semantic rules for question generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 752\u2013765.","DOI":"10.18653\/v1\/2020.acl-main.69"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-017-1100-y"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10844-019-00584-7"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2018.2867638"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.5555\/1289189.1289273"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","unstructured":"Li Dong Furu Wei Ming Zhou and Ke Xu. 2015. Question answering over freebase with multi-column convolutional neural networks. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 260\u2013269. DOI:10.3115\/v1\/p15-1026","DOI":"10.3115\/v1\/p15-1026"},{"key":"e_1_3_1_47_2","unstructured":"Li Dong Nan Yang Wenhui Wang Furu Wei Xiaodong Liu Yu Wang Jianfeng Gao Ming Zhou and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. In Proceedings of the 33rd International Conference on Neural Information Processing Systems 32 (2019) 1\u201313."},{"key":"e_1_3_1_48_2","unstructured":"Xinya Du Junru Shao and Claire Cardie. 2017. Learning to Ask: Neural question generation for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1342\u20131352."},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/d17-1090"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","unstructured":"Mohnish Dubey Debayan Banerjee Abdelrahman Abdelkawi and Jens Lehmann. 2019. Lc-quad 2.0: A large dataset for complex question answering over wikidata and dbpedia. In International Semantic Web Conference. Springer 69\u201378. 10.1007\/978-3-030-30796-7_5","DOI":"10.1007\/978-3-030-30796-7_5"},{"key":"e_1_3_1_51_2","unstructured":"Matthew Dunn Levent Sagun Mike Higgins V. Ugur Guney Volkan Cirik and Kyunghyun Cho. 2017. Searchqa: A new q&a dataset augmented with context from a search engine. arXiv preprint arXiv:1704.05179 (2017)."},{"issue":"1","key":"e_1_3_1_52_2","doi-asserted-by":"crossref","first-page":"521","DOI":"10.3390\/make1010032","article-title":"Evaluation of regression models: Model assessment, model selection and generalization error","volume":"1","author":"Emmert-Streib Frank","year":"2019","unstructured":"Frank Emmert-Streib and Matthias Dehmer. 2019. Evaluation of regression models: Model assessment, model selection and generalization error. Machine Learning and Knowledge Extraction 1, 1 (2019), 521\u2013551.","journal-title":"Machine Learning and Knowledge Extraction"},{"issue":"5","key":"e_1_3_1_53_2","first-page":"e1303","article-title":"A comprehensive survey of error measures for evaluating binary decision making in data science","volume":"9","author":"Emmert-Streib Frank","year":"2019","unstructured":"Frank Emmert-Streib, Salisou Moutari, and Matthias Dehmer. 2019. A comprehensive survey of error measures for evaluating binary decision making in data science. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 9, 5 (2019), e1303.","journal-title":"Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery"},{"key":"e_1_3_1_54_2","doi-asserted-by":"crossref","unstructured":"Frank Emmert-Streib Zhen Yang Han Feng Shailesh Tripathi and Matthias Dehmer. 2020. An introductory review of deep learning for prediction models with big data: Frank Emmert-Streib Zhen Yang Han Feng Shailesh Tripathi Matthias Dehmer:[Ressource \u00e9lectronique]. Frontiers in Artificial Intelligence 3 4 (2020) 23.","DOI":"10.3389\/frai.2020.00004"},{"key":"e_1_3_1_55_2","unstructured":"Alexander Richard Fabbri Patrick Ng Zhiguo Wang Ramesh Nallapati and Bing Xiang. 2020. Template-based question generation from retrieved sentences for improved unsupervised question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 4508\u20134513."},{"key":"e_1_3_1_56_2","doi-asserted-by":"crossref","unstructured":"Amer Farea and Frank Emmert-Streib. 2024. Experimental design of extractive question-answering systems: Influence of error scores and answer length. Journal of Artificial Intelligence Research 80 (2024) 87\u2013125.","DOI":"10.1613\/jair.1.15642"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2020.3010650"},{"issue":"3","key":"e_1_3_1_58_2","first-page":"1","article-title":"Introduction to \u201cthis is watson\u201d","volume":"56","author":"Ferrucci David A","year":"2012","unstructured":"David A Ferrucci. 2012. Introduction to \u201cthis is watson\u201d. IBM Journal of Research and Development 56, 3.4 (2012), 1\u20131.","journal-title":"IBM Journal of Research and Development"},{"key":"e_1_3_1_59_2","doi-asserted-by":"crossref","unstructured":"Luciano Floridi and Massimo Chiriatti. 2020. GPT-3: Its nature scope limits and consequences. Minds and Machines 30 4 (2020) 681\u2013694.","DOI":"10.1007\/s11023-020-09548-1"},{"key":"e_1_3_1_60_2","unstructured":"Jinlan Fu See Kiong Ng Zhengbao Jiang and Pengfei Liu. 2024. GPTScore: Evaluate as you desire. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 6556\u20136576."},{"key":"e_1_3_1_61_2","doi-asserted-by":"crossref","unstructured":"Michel Galley Chris Brockett Alessandro Sordoni Yangfeng Ji Michael Auli Chris Quirk Margaret Mitchell Jianfeng Gao and William B. Dolan. 2015. deltaBLEU: A discriminative metric for generation tasks with intrinsically diverse targets. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). 445\u2013450.","DOI":"10.3115\/v1\/P15-2073"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","unstructured":"Yifan Gao Lidong Bing Wang Chen Michael R Lyu and Irwin King. 2018. Difficulty controllable generation of reading comprehension questions. arXiv preprint arXiv:1807.03586 (2018). 10.24963\/ijcai.2019\/690","DOI":"10.24963\/ijcai.2019\/690"},{"key":"e_1_3_1_63_2","doi-asserted-by":"crossref","unstructured":"Albert Gatt and Emiel Krahmer. 2018. Survey of the state of the art in natural language generation: Core tasks applications and evaluation. Journal of Artificial Intelligence Research 61 1 (2018) 65\u2013170.","DOI":"10.1613\/jair.5477"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1162\/089976600300015015"},{"key":"e_1_3_1_65_2","first-page":"219","volume-title":"Proceedings of the Papers Presented at the May 9-11, 1961, Western Joint IRE-AIEE-ACM Computer Conference","author":"Jr Bert F. Green","year":"1961","unstructured":"Bert F. Green Jr, Alice K. Wolf, Carol Chomsky, and Kenneth Laughery. 1961. Baseball: An automatic question-answerer. In Proceedings of the Papers Presented at the May 9-11, 1961, Western Joint IRE-AIEE-ACM Computer Conference. 219\u2013224."},{"key":"e_1_3_1_66_2","first-page":"297","volume-title":"Proceedings of the 13th International Conference on Artificial Intelligence and Statistics","author":"Gutmann Michael","year":"2010","unstructured":"Michael Gutmann and Aapo Hyv\u00e4rinen. 2010. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics. JMLR Workshop and Conference Proceedings, 297\u2013304."},{"key":"e_1_3_1_67_2","doi-asserted-by":"crossref","unstructured":"Tianyong Hao Xinxin Li Yulan He Fu Lee Wang and Yingying Qu. 2022. Recent progress in leveraging deep learning methods for question answering. Neural Computing & Applications 34 4 (2022) 2765\u20132783.","DOI":"10.1007\/s00521-021-06748-3"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.3115\/1219840.1219866"},{"issue":"2","key":"e_1_3_1_69_2","first-page":"91","article-title":"Applying deep matching networks to Chinese medical question answering: A study and a dataset","volume":"19","author":"He Junqing","year":"2019","unstructured":"Junqing He, Mingming Fu, and Manshu Tu. 2019. Applying deep matching networks to Chinese medical question answering: A study and a dataset. BMC Medical Informatics and Decision Making 19, 2 (2019), 91\u2013100.","journal-title":"BMC Medical Informatics and Decision Making"},{"key":"e_1_3_1_70_2","doi-asserted-by":"crossref","unstructured":"Wei He Kai Liu Jing Liu Yajuan Lyu Shiqi Zhao Xinyan Xiao Yuan Liu Yizhong Wang Hua Wu Qiaoqiao She et\u00a0al. 2018. DuReader: A chinese machine reading comprehension dataset from real-world applications. In Proceedings of the Workshop on Machine Reading for Question Answering. 37\u201346.","DOI":"10.18653\/v1\/W18-2605"},{"key":"e_1_3_1_71_2","first-page":"609","volume-title":"Proceedings of the Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics","author":"Heilman Michael","year":"2010","unstructured":"Michael Heilman and Noah A Smith. 2010. Good question! statistical ranking for question generation. In Proceedings of the Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics. 609\u2013617."},{"key":"e_1_3_1_72_2","unstructured":"Felix Hill Antoine Bordes Sumit Chopra and Jason Weston. 2016. The Goldilocks principle: Reading children\u2019s books with explicit memory representations. In 4th International Conference on Learning Representations ICLR 2016."},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1142\/S0218488598000094"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/tkde.2017.2766634"},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3233771"},{"key":"e_1_3_1_76_2","first-page":"376","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers","author":"Indurthi Sathish Reddy","year":"2017","unstructured":"Sathish Reddy Indurthi, Dinesh Raghu, Mitesh M. Khapra, and Sachindra Joshi. 2017. Generating natural language question-answer pairs from a knowledge graph using a RNN based question generation model. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 376\u2013385."},{"key":"e_1_3_1_77_2","first-page":"1148","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies","author":"Ji Heng","year":"2011","unstructured":"Heng Ji and Ralph Grishman. 2011. Knowledge base population: Successful approaches and challenges. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies. 1148\u20131158."},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco_a_01174"},{"key":"e_1_3_1_79_2","doi-asserted-by":"crossref","unstructured":"Mandar Joshi Eunsol Choi Daniel SWeld and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1601\u20131611.","DOI":"10.18653\/v1\/P17-1147"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.lindif.2023.102274"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.2196\/medinform.8751"},{"key":"e_1_3_1_82_2","first-page":"524","volume-title":"Proceedings of the Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Proceedings of the Main Conference","author":"Ko Jeongwoo","year":"2007","unstructured":"Jeongwoo Ko, Luo Si, and Eric Nyberg. 2007. A probabilistic framework for answer selection in question answering. In Proceedings of the Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Proceedings of the Main Conference. 524\u2013531."},{"key":"e_1_3_1_83_2","doi-asserted-by":"crossref","unstructured":"Tom\u00e1\u0161 Ko\u010disk\u1ef3 Jonathan Schwarz Phil Blunsom Chris Dyer Karl Moritz Hermann G\u00e1bor Melis and Edward Grefenstette. 2018. The narrativeqa reading comprehension challenge. Transactions of the Association for Computational Linguistics 6 (2018) 317\u2013328.","DOI":"10.1162\/tacl_a_00023"},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2011.07.047"},{"key":"e_1_3_1_85_2","first-page":"2877","volume-title":"Proceedings of the Interspeech","author":"Kombrink Stefan","year":"2011","unstructured":"Stefan Kombrink, Tomas Mikolov, Martin Karafi\u00e1t, and Luk\u00e1s Burget. 2011. Recurrent neural network based language modeling in meeting recognition. In Proceedings of the Interspeech 11 (2011), 2877\u20132880."},{"key":"e_1_3_1_86_2","unstructured":"Klaus Krippendorff. 2011. Computing Krippendorff\u2019s Alpha-Reliability. https:\/\/api.semanticscholar.org\/CorpusID:59901023"},{"key":"e_1_3_1_87_2","doi-asserted-by":"crossref","unstructured":"Kalpesh Krishna Aurko Roy and Mohit Iyyer. 2021. Hurdles to progress in long-form question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4940\u20134957.","DOI":"10.18653\/v1\/2021.naacl-main.393"},{"key":"e_1_3_1_88_2","first-page":"1378","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Kumar Ankit","year":"2016","unstructured":"Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher. 2016. Ask me anything: Dynamic memory networks for natural language processing. In Proceedings of the International Conference on Machine Learning. PMLR, 1378\u20131387."},{"key":"e_1_3_1_89_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00276"},{"key":"e_1_3_1_90_2","doi-asserted-by":"crossref","unstructured":"Guokun Lai Qizhe Xie Hanxiao Liu Yiming Yang and Eduard Hovy. 2017. RACE: Large-scale reading comprehension dataset from examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D17-1082"},{"key":"e_1_3_1_91_2","doi-asserted-by":"crossref","unstructured":"Yunshi Lan Gaole He Jinhao Jiang Jing Jiang Wayne Xin Zhao and Ji-Rong Wen. 2021. A survey on complex knowledge base question answering: methods challenges and solutions. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization 4483\u20134491.","DOI":"10.24963\/ijcai.2021\/611"},{"key":"e_1_3_1_92_2","doi-asserted-by":"publisher","unstructured":"Yunshi Lan and Jing Jiang. 2020. Query graph generation for answering multi-hop complex questions from knowledge bases. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics Online 969\u2013974. 10.18653\/v1\/2020.acl-main.91","DOI":"10.18653\/v1\/2020.acl-main.91"},{"key":"e_1_3_1_93_2","doi-asserted-by":"publisher","DOI":"10.1109\/taslp.2019.2926125"},{"key":"e_1_3_1_94_2","unstructured":"Zhenzhong Lan Mingda Chen Sebastian Goodman Kevin Gimpel Piyush Sharma and Radu Soricut. 2019. ALBERT: A lite BERT for self-supervised learning of language representations. In International Conference on Learning Representations."},{"key":"e_1_3_1_95_2","doi-asserted-by":"publisher","DOI":"10.1109\/taslp.2019.2913499"},{"key":"e_1_3_1_96_2","unstructured":"Kenton Lee Ming-Wei Chang and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 6086\u20136096."},{"key":"e_1_3_1_97_2","doi-asserted-by":"publisher","DOI":"10.3233\/sw-140134"},{"key":"e_1_3_1_98_2","unstructured":"Vladimir I. Levenshtein. 1966. Binary codes capable of correcting deletions insertions and reversals. In Soviet Physics Doklady Vol. 10. Soviet Union 707\u2013710."},{"key":"e_1_3_1_99_2","first-page":"2231","volume-title":"Proceedings of the LREC","author":"Levy Roger","year":"2006","unstructured":"Roger Levy and Galen Andrew. 2006. Tregex and tsurgeon: Tools for querying and manipulating tree data structures.. In Proceedings of the LREC. Citeseer, 2231\u20132234."},{"key":"e_1_3_1_100_2","doi-asserted-by":"crossref","unstructured":"Patrick Lewis Pontus Stenetorp and Sebastian Riedel. 2021. Question and answer test-train overlap in open- domain question answering datasets. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 1000\u20131008.","DOI":"10.18653\/v1\/2021.eacl-main.86"},{"key":"e_1_3_1_101_2","unstructured":"Percy Liang Rishi Bommasani Tony Lee Dimitris Tsipras Dilara Soylu Michihiro Yasunaga Yian Zhang Deepak Narayanan Yuhuai Wu Ananya Kumar et\u00a0al. 2023. Holistic evaluation of language models. Trans. Mach. Learn. Res. (2023)."},{"key":"e_1_3_1_102_2","article-title":"A technique for the measurement of attitudes.","author":"Likert Rensis","year":"1932","unstructured":"Rensis Likert. 1932. A technique for the measurement of attitudes. Archives of Psychology 140 (1932), 5\u201355.","journal-title":"Archives of Psychology"},{"key":"e_1_3_1_103_2","first-page":"74","volume-title":"Proceedings of the Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Proceedings of the Text Summarization Branches Out. 74\u201381."},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.3115\/1073445.1073465"},{"key":"e_1_3_1_105_2","first-page":"105","volume-title":"Proceedings of the 14th European Workshop on Natural Language Generation","author":"Lindberg David","year":"2013","unstructured":"David Lindberg, Fred Popowich, John Nesbit, and Phil Winne. 2013. Generating natural language questions to support learning on-line. In Proceedings of the 14th European Workshop on Natural Language Generation. 105\u2013114."},{"key":"e_1_3_1_106_2","first-page":"75","volume-title":"Proceedings of the National CCF Conference on Natural Language Processing and Chinese Computing","author":"Liu Tianyu","year":"2017","unstructured":"Tianyu Liu, Bingzhen Wei, Baobao Chang, and Zhifang Sui. 2017. Large-scale simple question generation by template-based seq2seq learning. In Proceedings of the National CCF Conference on Natural Language Processing and Chinese Computing. Springer, 75\u201387."},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-15-9323-9_15"},{"key":"e_1_3_1_108_2","doi-asserted-by":"crossref","unstructured":"Yang Liu Dan Iter Yichong Xu Shuohang Wang Ruochen Xu and Chenguang Zhu. 2023. G-Eval: NLG evaluation using Gpt-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2511\u20132522.","DOI":"10.18653\/v1\/2023.emnlp-main.153"},{"key":"e_1_3_1_109_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)."},{"key":"e_1_3_1_110_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-4767"},{"key":"e_1_3_1_111_2","unstructured":"Lajanugen Logeswaran and Honglak Lee. 2018. An efficient framework for learning sentence representations. In International Conference on Learning Representations."},{"key":"e_1_3_1_112_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-89363-7_25"},{"key":"e_1_3_1_113_2","doi-asserted-by":"crossref","unstructured":"Alejandro Lopez-Lira and Yuehua Tang. 2023. Can ChatGPT forecast stock price movements? Return predictability and large language models. SSRN Electronic Journal (2023).","DOI":"10.2139\/ssrn.4412788"},{"key":"e_1_3_1_114_2","doi-asserted-by":"crossref","unstructured":"Ryan Lowe Michael Noseworthy Iulian Vlad Serban Nicolas Angelard-Gontier Yoshua Bengio and Joelle Pineau. 2017. Towards an automatic turing test: learning to evaluate dialogue responses. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1116\u20131126.","DOI":"10.18653\/v1\/P17-1103"},{"key":"e_1_3_1_115_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.345"},{"key":"e_1_3_1_116_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/d18-1242"},{"key":"e_1_3_1_117_2","first-page":"598","volume-title":"Proceedings of the 2nd Conference on Machine Translation","author":"Ma Qingsong","year":"2017","unstructured":"Qingsong Ma, Yvette Graham, Shugen Wang, and Qun Liu. 2017. Blend: A novel combined MT metric based on direct assessment\u2014CASICT-DCU submission to WMT17 metrics task. In Proceedings of the 2nd Conference on Machine Translation. 598\u2013603."},{"key":"e_1_3_1_118_2","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Greg Corrado and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2 (Lake Tahoe Nevada) (NIPS\u201913). Curran Associates Inc. Red Hook NY USA 3111\u20133119."},{"key":"e_1_3_1_119_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksuci.2014.10.007"},{"key":"e_1_3_1_120_2","doi-asserted-by":"crossref","unstructured":"Jinjie Ni Tom Young Vlad Pandelea Fuzhao Xue and Erik Cambria. 2023. Recent advances in deep learning based dialogue systems: A systematic survey. Artificial Intelligence Review 56 4 (2023) 3055\u20133155.","DOI":"10.1007\/s10462-022-10248-8"},{"key":"e_1_3_1_121_2","doi-asserted-by":"crossref","unstructured":"Jekaterina Novikova Ond\u0159ej Du\u0161ek Amanda Cercas Curry and Verena Rieser. 2017. Why we need new evaluation metrics for NLG. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2241\u20132252.","DOI":"10.18653\/v1\/D17-1238"},{"key":"e_1_3_1_122_2","doi-asserted-by":"publisher","DOI":"10.13052\/jwe1540-9589.1785"},{"key":"e_1_3_1_123_2","unstructured":"OpenAI. 2023. OpenAI Chat. Retrieved October 20 2024 from https:\/\/chat.openai.com"},{"key":"e_1_3_1_124_2","unstructured":"Long Ouyang Jeff Wu Xu Jiang Diogo Almeida Carroll L. Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray et\u00a0al. 2022. Training language models to follow instructions with human feedback. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS\u201922). 27730\u201327744."},{"key":"e_1_3_1_125_2","unstructured":"Liangming Pan Wenqiang Lei Tat-Seng Chua and Min-Yen Kan. 2019. Recent advances in neural question generation. arXiv preprint arXiv:1905.08949 (2019)."},{"key":"e_1_3_1_126_2","unstructured":"Liangming Pan Yuxi Xie Yansong Feng Tat-Seng Chua and Min-Yen Kan. 2020. Semantic graphs for generating deep questions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 1463\u20131475."},{"key":"e_1_3_1_127_2","doi-asserted-by":"publisher","DOI":"10.1109\/icassp40776.2020.9053822"},{"key":"e_1_3_1_128_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_3_1_129_2","first-page":"20","volume-title":"Proceedings of the International Conference on Document Analysis and Recognition","author":"Pe\u00f1a Alejandro","year":"2023","unstructured":"Alejandro Pe\u00f1a, Aythami Morales, Julian Fierrez, Ignacio Serna, Javier Ortega-Garcia, I\u00f1igo Puente, Jorge Cordova, and Gonzalo Cordova. 2023. Leveraging large language models for topic classification in the domain of public affairs. In Proceedings of the International Conference on Document Analysis and Recognition. Springer, 20\u201333."},{"key":"e_1_3_1_130_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-012-9177-0"},{"key":"e_1_3_1_131_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W15-3049"},{"key":"e_1_3_1_132_2","unstructured":"Weizhen Qi Yu Yan Yeyun Gong Dayiheng Liu Nan Duan Jiusheng Chen Ruofei Zhang and Ming Zhou. 2020. ProphetNet: Predicting future N-gram for sequence-to-sequence Pre-training. In Findings of the Association for Computational Linguistics: EMNLP 2020. 2401\u20132410."},{"key":"e_1_3_1_133_2","unstructured":"Chengwei Qin Aston Zhang Zhuosheng Zhang Jiaao Chen Michihiro Yasunaga and Diyi Yang. 2023. Is ChatGPT a general-purpose natural language processing task solver? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 1339\u20131384."},{"key":"e_1_3_1_134_2","unstructured":"Yingqi Qu Jie Liu Liangyi Kang Qinfeng Shi and Dan Ye. 2018. Question answering over freebase via attentive RNN with similarity matrix based CNN. arXiv:1804.03317. Retrieved from https:\/\/arxiv.org\/abs\/1804.03317"},{"key":"e_1_3_1_135_2","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. (2018)."},{"key":"e_1_3_1_136_2","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1 8 (2019) 9."},{"key":"e_1_3_1_137_2","doi-asserted-by":"crossref","unstructured":"Pranav Rajpurkar Robin Jia and Percy Liang. 2018. Know what you don\u2019t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 784\u2013789.","DOI":"10.18653\/v1\/P18-2124"},{"key":"e_1_3_1_138_2","doi-asserted-by":"crossref","unstructured":"Pranav Rajpurkar Jian Zhang Konstantin Lopyrev and Percy Liang. 2016. SQuAD: 100 000+ Questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2383\u20132392.","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_3_1_139_2","article-title":"Learning from crowds.","author":"Raykar Vikas C.","year":"2010","unstructured":"Vikas C. Raykar, Shipeng Yu, Linda H Zhao, Gerardo Hermosillo Valadez, Charles Florin, Luca Bogoni, and Linda Moy. 2010. Learning from crowds. Journal of Machine Learning Research 11, 4 (2010), 1297\u20131322.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_140_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00266"},{"key":"e_1_3_1_141_2","doi-asserted-by":"publisher","DOI":"10.1162\/coli.2009.35.4.35405"},{"key":"e_1_3_1_142_2","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324997001502"},{"key":"e_1_3_1_143_2","doi-asserted-by":"crossref","unstructured":"Nicholas Riccardi Xuan Yang and Rutvik H. Desai. 2024. The Two Word Test as a semantic benchmark for large language models. Scientific Reports 14 1 (2024) 21593.","DOI":"10.1038\/s41598-024-72528-3"},{"key":"e_1_3_1_144_2","doi-asserted-by":"crossref","first-page":"193","DOI":"10.18653\/v1\/D13-1020","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing","author":"Richardson Matthew","year":"2013","unstructured":"Matthew Richardson, Christopher J. C. Burges, and Erin Renshaw. 2013. Mctest: A challenge dataset for the open-domain machine comprehension of text. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. 193\u2013203."},{"key":"e_1_3_1_145_2","doi-asserted-by":"crossref","unstructured":"Adam Roberts Colin Raffel and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 5418\u20135426.","DOI":"10.18653\/v1\/2020.emnlp-main.437"},{"key":"e_1_3_1_146_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-27036-4_8"},{"key":"e_1_3_1_147_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560260"},{"key":"e_1_3_1_148_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485766"},{"key":"e_1_3_1_149_2","first-page":"167","article-title":"Learning concepts by asking questions","volume":"2","author":"Sammut Claude","year":"1986","unstructured":"Claude Sammut and Ranan B. Banerji. 1986. Learning concepts by asking questions. Machine Learning: An Artificial Intelligence Approach 2 (1986), 167\u2013192.","journal-title":"Machine Learning: An Artificial Intelligence Approach"},{"key":"e_1_3_1_150_2","unstructured":"Victor Sanh Lysandre Debut Julien Chaumond and Thomas Wolf. 2019. DistilBERT a distilled version of BERT: Smaller faster cheaper and lighter. arXiv:1910.01108. Retrieved from https:\/\/arxiv.org\/abs\/1910.01108"},{"key":"e_1_3_1_151_2","doi-asserted-by":"crossref","unstructured":"Thibault Sellam Dipanjan Das and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 7881\u20137892.","DOI":"10.18653\/v1\/2020.acl-main.704"},{"key":"e_1_3_1_152_2","unstructured":"Minjoon Seo Aniruddha Kembhavi Ali Farhadi and Hananneh Hajishirzi. 2017. Bi-Directional attention flow for machine comprehension. In 5th International Conference on Learning Representations (ICLR 2017)."},{"key":"e_1_3_1_153_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.10983"},{"key":"e_1_3_1_154_2","doi-asserted-by":"crossref","unstructured":"Naeha Sharif Mohammed Bennamoun Lyndon Rhys White and Syed Afaq Ali Shah. 2018. Learning-based composite metrics for improved caption evaluation. In 56th Annual Meeting of Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-3003"},{"key":"e_1_3_1_155_2","doi-asserted-by":"crossref","first-page":"751","DOI":"10.18653\/v1\/W18-6456","volume-title":"Proceedings of the 3rd Conference on Machine Translation: Shared Task Papers","author":"Shimanaka Hiroki","year":"2018","unstructured":"Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2018. Ruse: Regressor using sentence embeddings for automatic machine translation evaluation. In Proceedings of the 3rd Conference on Machine Translation: Shared Task Papers. 751\u2013758."},{"key":"e_1_3_1_156_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06291-2"},{"key":"e_1_3_1_157_2","doi-asserted-by":"crossref","unstructured":"Koustuv Sinha Prasanna Parthasarathi Jasmine Wang Ryan Lowe William L. Hamilton and Joelle Pineau. 2020. Learning an unreferenced metric for online dialogue evaluation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2430\u20132441.","DOI":"10.18653\/v1\/2020.acl-main.220"},{"key":"e_1_3_1_158_2","first-page":"569","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)","author":"Song Linfeng","year":"2018","unstructured":"Linfeng Song, Zhiguo Wang, Wael Hamza, Yue Zhang, and Daniel Gildea. 2018. Leveraging context information for natural question generation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). 569\u2013574."},{"key":"e_1_3_1_159_2","first-page":"75","volume-title":"Proceedings of the International Conference on Computer Vision and Image Processing","author":"Srivastava Yash","year":"2020","unstructured":"Yash Srivastava, Vaishnav Murali, Shiv Ram Dubey, and Snehasis Mukherjee. 2020. Visual question answering using deep learning: A survey and performance analysis. In Proceedings of the International Conference on Computer Vision and Image Processing. Springer, 75\u201386."},{"key":"e_1_3_1_160_2","doi-asserted-by":"crossref","first-page":"414","DOI":"10.3115\/v1\/W14-3354","volume-title":"Proceedings of the 9th Workshop on Statistical Machine Translation","author":"Stanojevi\u0107 Milo\u0161","year":"2014","unstructured":"Milo\u0161 Stanojevi\u0107 and Khalil Sima\u2019an. 2014. Beer: Better evaluation as ranking. In Proceedings of the 9th Workshop on Statistical Machine Translation. 414\u2013419."},{"key":"e_1_3_1_161_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(99)00037-5"},{"key":"e_1_3_1_162_2","unstructured":"Sainbayar Sukhbaatar Arthur Szlam Jason Weston and Rob Fergus. 2015. End-To-end memory networks. In Advances in Neural Information Processing Systems 28 (2015)."},{"key":"e_1_3_1_163_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.websem.2021.100698"},{"key":"e_1_3_1_164_2","unstructured":"Ilya Sutskever Oriol Vinyals and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems-Volume 2. 3104\u20133112."},{"key":"e_1_3_1_165_2","doi-asserted-by":"crossref","unstructured":"Alon Talmor and Jonathan Berant. 2018. The web as a knowledge-base for answering complex questions. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Volume 1 (Long Papers). 641\u2013651.","DOI":"10.18653\/v1\/N18-1059"},{"key":"e_1_3_1_166_2","volume-title":"Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence","author":"Tao Chongyang","year":"2018","unstructured":"Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan. 2018. Ruber: An unsupervised method for automatic evaluation of open-domain dialog systems. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_1_167_2","unstructured":"Zhengwei Tao Zhi Jin Xiaoying Bai Haiyan Zhao Yanlin Feng Jia Li and Wenpeng Hu. 2023. Eveval: A comprehensive evaluation of event semantics for large language models. arXiv:2305.15268. Retrieved from https:\/\/arxiv.org\/abs\/2305.15268"},{"key":"e_1_3_1_168_2","doi-asserted-by":"crossref","unstructured":"Craig Thomson and Ehud Reiter. 2020. A gold standard methodology for evaluating accuracy in data-to-text systems. In Proceedings of the 13th International Conference on Natural Language Generation. 158\u2013168.","DOI":"10.18653\/v1\/2020.inlg-1.22"},{"key":"e_1_3_1_169_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et\u00a0al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)."},{"key":"e_1_3_1_170_2","doi-asserted-by":"crossref","unstructured":"Adam Trischler Tong Wang Xingdi Yuan Justin Harris Alessandro Sordoni Philip Bachman and Kaheer Suleman. 2017. NewsQA: A machine comprehension dataset. In Proceedings of the 2nd Workshop on Representation Learning for NLP. 191\u2013200.","DOI":"10.18653\/v1\/W17-2623"},{"key":"e_1_3_1_171_2","doi-asserted-by":"crossref","first-page":"355","DOI":"10.18653\/v1\/W19-8643","volume-title":"Proceedings of the 12th International Conference on Natural Language Generation","author":"Lee Chris Van Der","year":"2019","unstructured":"Chris Van Der Lee, Albert Gatt, Emiel Van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019. Best practices for the human evaluation of automatically generated text. In Proceedings of the 12th International Conference on Natural Language Generation. 355\u2013368."},{"key":"e_1_3_1_172_2","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems. 6000\u20136010."},{"key":"e_1_3_1_173_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_1_174_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_1_175_2","doi-asserted-by":"crossref","unstructured":"Ellen M. Voorhees. 1999. The trec-8 question answering track report. In Trec Vol. 99. 77\u201382.","DOI":"10.6028\/NIST.SP.500-246.qa-overview"},{"key":"e_1_3_1_176_2","doi-asserted-by":"crossref","unstructured":"Jiaan Wang Yunlong Liang Fandong Meng Zengkui Sun Haoxiang Shi Zhixu Li Jinan Xu Jianfeng Qu and Jie Zhou. 2023. Is ChatGPT a good NLG evaluator? A preliminary study. In Proceedings of the 4th New Frontiers in Summarization Workshop. 1\u201311.","DOI":"10.18653\/v1\/2023.newsum-1.1"},{"key":"e_1_3_1_177_2","first-page":"22","volume-title":"Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL)","author":"Wang Mengqiu","year":"2007","unstructured":"Mengqiu Wang, Noah A. Smith, and Teruko Mitamura. 2007. What is the jeopardy model? A quasi-synchronous grammar for QA. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL). 22\u201332."},{"key":"e_1_3_1_178_2","first-page":"441","volume-title":"Proceedings of the June 4-8, 1973, National Computer Conference and Exposition","author":"Woods William A.","year":"1973","unstructured":"William A. Woods. 1973. Progress in natural language understanding: An application to lunar geology. In Proceedings of the June 4-8, 1973, National Computer Conference and Exposition. 441\u2013450."},{"key":"e_1_3_1_179_2","doi-asserted-by":"crossref","unstructured":"Qi Wu Damien Teney Peng Wang Chunhua Shen Anthony Dick and Anton van den Hengel. 2017. Visual question answering. Computer Vision and Image Understanding 163 C (2017) 21\u201340.","DOI":"10.1016\/j.cviu.2017.05.001"},{"key":"e_1_3_1_180_2","doi-asserted-by":"crossref","unstructured":"Yu Wu Wei Wu Chen Xing Ming Zhou and Zhoujun Li. 2017. Sequential matching network: A new architecture for multi-turn response selection in retrieval-based chatbots. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 496\u2013505.","DOI":"10.18653\/v1\/P17-1046"},{"key":"e_1_3_1_181_2","doi-asserted-by":"crossref","unstructured":"Zhaofeng Wu Linlu Qiu Alexis Ross Ekin Aky\u00fcrek Boyuan Chen Bailin Wang Najoung Kim Jacob Andreas and Yoon Kim. 2024. Reasoning or reciting? Exploring the capabilities and limitations of language models through counterfactual tasks. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 1819\u20131862.","DOI":"10.18653\/v1\/2024.naacl-long.102"},{"key":"e_1_3_1_182_2","first-page":"2048","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of the International Conference on Machine Learning. PMLR, 2048\u20132057."},{"key":"e_1_3_1_183_2","doi-asserted-by":"publisher","unstructured":"Kun Xu Yansong Feng Songfang Huang and Dongyan Zhao. 2015. Question answering via phrasal semantic parsing. In International Conference of the Cross-Language Evaluation Forum for European Languages. Springer 414\u2013426. 10.1007\/978-3-319-24027-5_43","DOI":"10.1007\/978-3-319-24027-5_43"},{"key":"e_1_3_1_184_2","unstructured":"Kai-Cheng Yang and Filippo Menczer. 2023. Large language models can rate news outlet credibility. arXiv preprint arXiv:2304.00228 (2023)."},{"key":"e_1_3_1_185_2","doi-asserted-by":"crossref","unstructured":"Zhilin Yang Junjie Hu Ruslan Salakhutdinov and William Cohen. 2017. Semi-supervised QA with generative domain-adaptive nets. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1040\u20131050.","DOI":"10.18653\/v1\/P17-1096"},{"key":"e_1_3_1_186_2","doi-asserted-by":"crossref","unstructured":"Zhilin Yang Peng Qi Saizheng Zhang Yoshua Bengio William Cohen Ruslan Salakhutdinov and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2369\u20132380.","DOI":"10.18653\/v1\/D18-1259"},{"key":"e_1_3_1_187_2","first-page":"68","volume-title":"Proceedings of QG2010: The 3rd Workshop on Question Generation","author":"Yao Xuchen","year":"2010","unstructured":"Xuchen Yao and Yi Zhang. 2010. Question generation with minimal recursion semantics. In Proceedings of QG2010: The 3rd Workshop on Question Generation. Citeseer, 68\u201375."},{"key":"e_1_3_1_188_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/p15-1128"},{"key":"e_1_3_1_189_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/p16-2033"},{"key":"e_1_3_1_190_2","doi-asserted-by":"crossref","unstructured":"Xingdi Yuan Tong Wang \u00c7a\u011flar Gul\u00e7ehre Alessandro Sordoni Philip Bachman Saizheng Zhang Sandeep Subramanian and Adam Trischler. 2017. Machine comprehension by Text-to-Text neural question generation. In Proceedings of the 2nd Workshop on Representation Learning for NLP. 15\u201325.","DOI":"10.18653\/v1\/W17-2603"},{"key":"e_1_3_1_191_2","doi-asserted-by":"crossref","unstructured":"Chunyi Yue Hanqiang Cao Kun Xiong Anqi Cui Haocheng Qin and Ming Li. 2017. Enhanced question understanding with dynamic memory networks for textual question answering. Expert Systems With Applications: An International Journal 80 C (2017) 39\u201345.","DOI":"10.1016\/j.eswa.2017.03.006"},{"key":"e_1_3_1_192_2","unstructured":"Aohan Zeng Xiao Liu Zhengxiao Du Zihan Wang Hanyu Lai Ming Ding Zhuoyi Yang Yifan Xu Wendi Zheng Xiao Xia et\u00a0al. 2022. GLM-130B: An open bilingual pre-trained model. In The Eleventh International Conference on Learning Representations."},{"key":"e_1_3_1_193_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2022.108252"},{"key":"e_1_3_1_194_2","unstructured":"Liwen Zhang John Winn and Ryota Tomioka. 2016. Gaussian attention model and its application to knowledge base embedding and question answering. arXiv:1611.02266. Retrieved from https:\/\/arxiv.org\/abs\/1611.02266"},{"key":"e_1_3_1_195_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468889"},{"key":"e_1_3_1_196_2","doi-asserted-by":"crossref","unstructured":"Shiyue Zhang and Mohit Bansal. 2019. Addressing semantic drift in question generation for semi-supervised question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2495\u20132509.","DOI":"10.18653\/v1\/D19-1253"},{"key":"e_1_3_1_197_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-018-1312-9"},{"key":"e_1_3_1_198_2","doi-asserted-by":"publisher","DOI":"10.1109\/access.2018.2883637"},{"key":"e_1_3_1_199_2","unstructured":"Tianyi Zhang Varsha Kishore Felix Wu Kilian Q. Weinberger and Yoav Artzi. 2019. BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations."},{"key":"e_1_3_1_200_2","doi-asserted-by":"crossref","unstructured":"Wenxuan Zhang Yue Deng Bing Liu Sinno Jialin Pan and Lidong Bing. 2024. Sentiment analysis in the era of large language models: A reality check. In NAACL-HLT (Findings).","DOI":"10.18653\/v1\/2024.findings-naacl.246"},{"key":"e_1_3_1_201_2","doi-asserted-by":"crossref","unstructured":"Xu Zhang Wenpeng Lu Fangfang Li Xueping Peng and Ruoyu Zhang. 2019. Deep feature fusion model for sentence semantic matching. Computers Materials & Continua 61 2 (2019) 601\u2013616.","DOI":"10.32604\/cmc.2019.06045"},{"key":"e_1_3_1_202_2","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence","author":"Zhang Yuyu","year":"2018","unstructured":"Yuyu Zhang, Hanjun Dai, Zornitsa Kozareva, Alexander J. Smola, and Le Song. 2018. Variational reasoning for question answering with knowledge graph. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_1_203_2","unstructured":"Wayne Xin Zhao Kun Zhou Junyi Li Tianyi Tang Xiaolei Wang Yupeng Hou Yingqian Min Beichen Zhang Junjie Zhang Zican Dong et\u00a0al. 2023. A survey of large language models. arXiv:2303.18223. Retrieved from https:\/\/arxiv.org\/abs\/2303.18223"},{"key":"e_1_3_1_204_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/d18-1424"},{"key":"e_1_3_1_205_2","unstructured":"Mantong Zhou Minlie Huang and Xiaoyan Zhu. 2018. An interpretable reasoning network for multi-relation question answering. In Proceedings of the 27th International Conference on Computational Linguistics. 2010\u20132022."},{"key":"e_1_3_1_206_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-73618-1_56"},{"key":"e_1_3_1_207_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1036"},{"key":"e_1_3_1_208_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2019.09.003"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3744663","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,30]],"date-time":"2025-08-30T13:37:32Z","timestamp":1756561052000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3744663"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,30]]},"references-count":207,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1,31]]}},"alternative-id":["10.1145\/3744663"],"URL":"https:\/\/doi.org\/10.1145\/3744663","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,30]]},"assertion":[{"value":"2022-09-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-02","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-30","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}