{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,8]],"date-time":"2026-01-08T06:23:33Z","timestamp":1767853413317,"version":"3.49.0"},"reference-count":54,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2022,10,24]],"date-time":"2022-10-24T00:00:00Z","timestamp":1666569600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"SFI Research Centres Programme","award":["13\/RC\/2106_P2"],"award-info":[{"award-number":["13\/RC\/2106_P2"]}]},{"name":"SFI Research Centres Programme","award":["13\/RC\/2106"],"award-info":[{"award-number":["13\/RC\/2106"]}]},{"name":"SFI Research Centres Programme","award":["18\/CRT\/6183"],"award-info":[{"award-number":["18\/CRT\/6183"]}]},{"name":"Science Foundation Ireland through the SFI Centre for Research Training in Machine Learning","award":["13\/RC\/2106_P2"],"award-info":[{"award-number":["13\/RC\/2106_P2"]}]},{"name":"Science Foundation Ireland through the SFI Centre for Research Training in Machine Learning","award":["13\/RC\/2106"],"award-info":[{"award-number":["13\/RC\/2106"]}]},{"name":"Science Foundation Ireland through the SFI Centre for Research Training in Machine Learning","award":["18\/CRT\/6183"],"award-info":[{"award-number":["18\/CRT\/6183"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Question Generation (QG) aims to automate the task of composing questions for a passage with a set of chosen answers found within the passage. In recent years, the introduction of neural generation models has resulted in substantial improvements of automatically generated questions in terms of quality, especially compared to traditional approaches that employ manually crafted heuristics. However, current QG evaluation metrics solely rely on the comparison between the generated questions and references, ignoring the passages or answers. Meanwhile, these metrics are generally criticized because of their low agreement with human judgement. We therefore propose a new reference-free evaluation metric called QAScore, which is capable of providing a better mechanism for evaluating QG systems. QAScore evaluates a question by computing the cross entropy according to the probability that the language model can correctly generate the masked words in the answer to that question. Compared to existing metrics such as BLEU and BERTScore, QAScore can obtain a stronger correlation with human judgement according to our human evaluation experiment, meaning that applying QAScore in the QG task benefits to a higher level of evaluation accuracy.<\/jats:p>","DOI":"10.3390\/e24111514","type":"journal-article","created":{"date-parts":[[2022,10,24]],"date-time":"2022-10-24T06:37:45Z","timestamp":1666593465000},"page":"1514","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["QAScore\u2014An Unsupervised Unreferenced Metric for the Question Generation Evaluation"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0143-6220","authenticated-orcid":false,"given":"Tianbo","family":"Ji","sequence":"first","affiliation":[{"name":"ADAPT Centre, School of Computing, Dublin City University, 9 Dublin, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chenyang","family":"Lyu","sequence":"additional","affiliation":[{"name":"SFI Centre for Research Training in Machine Learning, School of Computing, Dublin City University, 9 Dublin, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gareth","family":"Jones","sequence":"additional","affiliation":[{"name":"ADAPT Centre, School of Computing, Dublin City University, 9 Dublin, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liting","family":"Zhou","sequence":"additional","affiliation":[{"name":"ADAPT Centre, School of Computing, Dublin City University, 9 Dublin, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yvette","family":"Graham","sequence":"additional","affiliation":[{"name":"ADAPT Centre, School of Computer Science and Statistics, Trinity College Dublin, 2 Dublin, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,10,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Du, X., Shao, J., and Cardie, C. (2017). Learning to Ask: Neural Question Generation for Reading Comprehension. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics.","DOI":"10.18653\/v1\/P17-1123"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Xie, Y., Pan, L., Wang, D., Kan, M.Y., and Feng, Y. (2020). Exploring Question-Specific Rewards for Generating Deep Questions. Proceedings of the 28th International Conference on Computational Linguistics, International Committee on Computational Linguistics.","DOI":"10.18653\/v1\/2020.coling-main.228"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Pan, L., Xie, Y., Feng, Y., Chua, T.S., and Kan, M.Y. (2020). Semantic Graphs for Generating Deep Questions. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.acl-main.135"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Puri, R., Spring, R., Shoeybi, M., Patwary, M., and Catanzaro, B. (2020). Training Question Answering Models From Synthetic Data. Proceedings of 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.emnlp-main.468"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Lyu, C., Shang, L., Graham, Y., Foster, J., Jiang, X., and Liu, Q. (2021). Improving Unsupervised Question Answering via Summarization-Informed Question Generation. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2021.emnlp-main.340"},{"key":"ref_6","unstructured":"Chen, Y., Wu, L., and Zaki, M.J. (2020, January 26\u201330). Reinforcement Learning Based Graph-to-Sequence Model for Natural Question Generation. Proceedings of the 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia."},{"key":"ref_7","unstructured":"Qiu, H., Zhang, C., Fei, Z., Qiu, M., and Kung, S.Y. (2021). TEBC-Net: An Effective Relation Extraction Approach for Simple Question Answering over Knowledge Graphs. Proceedings of the Knowledge Science, Engineering and Management, Springer International Publishing."},{"key":"ref_8","first-page":"6602","article-title":"Improving Neural Question Generation Using Answer Separation","volume":"33","author":"Kim","year":"2019","journal-title":"Proc. AAAI Conf. Artific. Intell."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Wang, L., Xu, Z., Lin, Z., Zheng, H., and Shen, Y. (2020). Answer-driven Deep Question Generation based on Reinforcement Learning. Proceedings of the 28th International Conference on Computational Linguistics, International Committee on Computational Linguistics.","DOI":"10.18653\/v1\/2020.coling-main.452"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Cho, W.S., Zhang, Y., Rao, S., Celikyilmaz, A., Xiong, C., Gao, J., Wang, M., and Dolan, B. (2021). Contrastive Multi-document Question Generation. Proceedings of 16th Conference of the European Chapter of the Association for Computational Linguistics: Mainv Volume, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2021.eacl-main.2"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002). Bleu: A Method for Automatic Evaluation of Machine Translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Nema, P., and Khapra, M.M. (2018). Towards a Better Metric for Evaluating Question Generation Systems. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1429"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Sellam, T., Das, D., and Parikh, A. (2020). BLEURT: Learning Robust Metrics for Text Generation. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.acl-main.704"},{"key":"ref_14","unstructured":"Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., and Artzi, Y. (2020, January 26\u201330). BERTScore: Evaluating Text Generation with BERT. Proceedings of the International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"393","DOI":"10.1162\/coli_a_00322","article-title":"A Structured Review of the Validity of BLEU","volume":"44","author":"Reiter","year":"2018","journal-title":"Comput. Linguist."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Graham, Y. (2015). Re-evaluating Automatic Summarization with BLEU and 192 Shades of ROUGE. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics.","DOI":"10.18653\/v1\/D15-1013"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Graham, Y., and Liu, Q. (2016, January 12\u201317). Achieving accurate conclusions in evaluation of automatic machine translation metrics. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego, CA, USA.","DOI":"10.18653\/v1\/N16-1001"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ji, T., Graham, Y., Jones, G.J., Lyu, C., and Liu, Q. (2022). Achieving Reliable Human Assessment of Open-Domain Dialogue Systems. arXiv.","DOI":"10.18653\/v1\/2022.acl-long.445"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ji, T., Lyu, C., Cao, Z., and Cheng, P. (2021). Multi-Hop Question Generation Using Hierarchical Encoding-Decoding and Context Switch Mechanism. Entropy, 23.","DOI":"10.3390\/e23111449"},{"key":"ref_20","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR, abs\/1907.11692, Available online: http:\/\/xxx.lanl.gov\/abs\/1907.11692."},{"key":"ref_21","unstructured":"Chen, D., Fisch, A., Weston, J., and Bordes, A. (August, January 30). Reading Wikipedia to Answer Open-Domain Questions. Proceedings of the Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vancouver, BC, Canada."},{"key":"ref_22","unstructured":"Zhu, F., Lei, W., Wang, C., Zheng, J., Poria, S., and Chua, T.S. (2021). Retrieving and reading: A comprehensive survey on open-domain question answering. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016, January 1\u20134). SQuAD: 100,000+ Questions for Machine Comprehension of Text. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, TX, USA.","DOI":"10.18653\/v1\/D16-1264"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Saha, A., Aralikatte, R., Khapra, M.M., and Sankaranarayanan, K. (2018). DuoRC: Towards Complex Language Understanding with Paraphrased Reading Comprehension. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-1156"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"317","DOI":"10.1162\/tacl_a_00023","article-title":"The NarrativeQA Reading Comprehension Challenge","volume":"6","author":"Schwarz","year":"2018","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Xu, Y., Wang, D., Yu, M., Ritchie, D., Yao, B., Wu, T., Zhang, Z., Li, T., Bradford, N., and Sun, B. (2022). Fantastic Questions and Where to Find Them: FairytaleQA\u2014An Authentic Dataset for Narrative Comprehension. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics.","DOI":"10.18653\/v1\/2022.acl-long.34"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Trischler, A., Wang, T., Yuan, X., Harris, J., Sordoni, A., Bachman, P., and Suleman, K. (2017). NewsQA: A Machine Comprehension Dataset. Proceedings of the 2nd Workshop on Representation Learning for NLP, Association for Computational Linguistics.","DOI":"10.18653\/v1\/W17-2623"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Lyu, C., Foster, J., and Graham, Y. (2022). Extending the Scope of Out-of-Domain: Examining QA models in multiple subdomains. Proceedings of the Third Workshop on Insights from Negative Results in NLP, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2022.insights-1.4"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Lewis, P., Wu, Y., Liu, L., Minervini, P., K\u00fcttler, H., Piktus, A., Stenetorp, P., and Riedel, S. (2021). PAQ: 65 Million Probably-Asked Questions and What You Can Do with Them. arXiv.","DOI":"10.1162\/tacl_a_00415"},{"key":"ref_30","unstructured":"Zhang, Z., Zhao, H., and Wang, R. (2020). Machine Reading Comprehension: The Role of Contextualized Language Models and Beyond. arXiv."},{"key":"ref_31","unstructured":"Pan, L., Lei, W., Chua, T., and Kan, M. (2019). Recent Advances in Neural Question Generation. CoRR, abs\/1905.08949, Available online: http:\/\/xxx.lanl.gov\/abs\/1905.08949."},{"key":"ref_32","unstructured":"Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N., and Weinberger, K.Q. (2014). Sequence to Sequence Learning with Neural Networks. Proceedings of the Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_33","unstructured":"Wu, Y., Schuster, M., Chen, Z., Le, Q.V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., and Macherey, K. (2016). Google\u2019s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. CoRR, abs\/1609.08144, Available online: http:\/\/xxx.lanl.gov\/abs\/1609.08144."},{"key":"ref_34","unstructured":"Lin, C.Y. (2004). ROUGE: A Package for Automatic Evaluation of Summaries. Proceedings of the Text Summarization Branches Out, Association for Computational Linguistics."},{"key":"ref_35","unstructured":"Banerjee, S., and Lavie, A. (2005). METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization, Association for Computational Linguistics."},{"key":"ref_36","first-page":"9459","article-title":"Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks","volume":"Volume 33","author":"Larochelle","year":"2020","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_37","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational Linguistics."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Yuan, X., Wang, T., Gulcehre, C., Sordoni, A., Bachman, P., Zhang, S., Subramanian, S., and Trischler, A. (2017). Machine Comprehension by Text-to-Text Neural Question Generation. Proceedings of the 2nd Workshop on Representation Learning for NLP, Association for Computational Linguistics.","DOI":"10.18653\/v1\/W17-2603"},{"key":"ref_39","unstructured":"Jia, X., Zhou, W., Sun, X., and Wu, Y. (2021, January 2\u20139). EQG-RACE: Examination-Type Question Generation. Proceedings of the AAAI, Palo Alto, CA, USA."},{"key":"ref_40","unstructured":"Ren, S., and Zhu, K.Q. (2020). Knowledge-Driven Distractor Generation for Cloze-style Multiple Choice Questions. CoRR, abs\/2004.09853, Available online: http:\/\/xxx.lanl.gov\/abs\/2004.09853."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Liu, B., Wei, H., Niu, D., Chen, H., and He, Y. (2020, January 20\u201324). Asking Questions the Human Way: Scalable Question-Answer Generation from Text Corpus. Proceedings of the Web Conference 2020, New York, NY, USA.","DOI":"10.1145\/3366423.3380270"},{"key":"ref_42","first-page":"8464","article-title":"Improving Question Generation with Sentence-Level Semantic Matching and Answer Position Inferring","volume":"34","author":"Ma","year":"2020","journal-title":"Proc. AAAI Conf. Artific. Intell."},{"key":"ref_43","unstructured":"Narayan, S., Sim\u00f5es, G., Ma, J., Craighead, H., and McDonald, R.T. (2020). QURIOUS: Question Generation Pretraining for Text Generation. CoRR, abs\/2004.11026, Available online: http:\/\/xxx.lanl.gov\/abs\/2004.11026."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zhou, S., and Zhang, Y. (2021). DATLMedQA: A Data Augmentation and Transfer Learning Based Solution for Medical Question Answering. Appl. Sci., 11.","DOI":"10.3390\/app112311251"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Shin, T., Razeghi, Y., Logan IV, R.L., Wallace, E., and Singh, S. (2020). AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.emnlp-main.346"},{"key":"ref_46","unstructured":"Lyu, C. (2022, October 20). Knowledge and Pre-Trained Language Models Inside and Out: A Deep-Dive into Datasets and External Knowledge 2022. Available online: https:\/\/scholar.google.co.jp\/scholar?hl=zh-TW&as_sdt=0%2C5&q=Knowledge+and+Pre-trained+Language+Models+Inside+and+Out%3A+A+deep-dive+++into+datasets+and+external+knowledge&btnG=."},{"key":"ref_47","unstructured":"Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W.W., Salakhutdinov, R., and Manning, C.D. (November, January 31). HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Brussels, Belgium."},{"key":"ref_48","first-page":"1","article-title":"Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. (2020). BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"ref_50","first-page":"9","article-title":"Language Models are Unsupervised Multitask Learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI Blog"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Cho, K., van Merri\u00ebnboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014, January 25\u201329). Learning Phrase Representations using RNN Encoder\u2013Decoder for Statistical Machine Translation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Ji, T., Graham, Y., and Jones, G.J. (2020). Contrasting Human Opinion of Non-Factoid Question Answering with Automatic Evaluation. Proceedings of the 2020 Conference on Human Information Interaction and Retrieval, Association for Computing Machinery.","DOI":"10.1145\/3343413.3377996"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1017\/S1351324915000339","article-title":"Can machine translation systems be evaluated by the crowd alone","volume":"23","author":"Graham","year":"2017","journal-title":"Nat. Lang. Eng."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Graham, Y., Haddow, B., and Koehn, P. (2020). Statistical Power and Translationese in Machine Translation Evaluation. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.emnlp-main.6"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/11\/1514\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:01:35Z","timestamp":1760144495000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/11\/1514"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,24]]},"references-count":54,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2022,11]]}},"alternative-id":["e24111514"],"URL":"https:\/\/doi.org\/10.3390\/e24111514","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,24]]}}}