{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,7]],"date-time":"2026-03-07T01:24:53Z","timestamp":1772846693810,"version":"3.50.1"},"reference-count":58,"publisher":"Cambridge University Press (CUP)","issue":"2","license":[{"start":{"date-parts":[[2016,4,14]],"date-time":"2016-04-14T00:00:00Z","timestamp":1460592000000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2017,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We describe the development, pilot-testing, refinement, and four evaluations of Diagnostic Question Generator (DQGen), which automatically generates multiple choice cloze (fill-in-the-blank) questions to test children's comprehension while reading a given text. Unlike previous methods, DQGen tests comprehension not only of an individual sentence but of the context preceding it. To test different aspects of comprehension, DQGen generates three types of distractors: ungrammatical distractors test syntax; nonsensical distractors test semantics; and locally plausible distractors test inter-sentential processing.<jats:list list-type=\"number\"><jats:list-item><jats:label>(1)<\/jats:label><jats:p>A pilot study of DQGen 2012 evaluated its overall questions and individual distractors, guiding its refinement into DQGen 2014.<\/jats:p><\/jats:list-item><jats:list-item><jats:label>(2)<\/jats:label><jats:p>Twenty-four elementary students generated 200 responses to multiple choice cloze questions that DQGen 2014 generated from forty-eight stories. In 130 of the responses, the child chose the correct answer. We define the<jats:italic>distractiveness<\/jats:italic>of a distractor as the frequency with which students choose it over the correct answer. The incorrect responses were consistent with expected distractiveness: twenty-seven were plausible, twenty-two were nonsensical, fourteen were ungrammatical, and seven were null.<\/jats:p><\/jats:list-item><jats:list-item><jats:label>(3)<\/jats:label><jats:p>To compare DQGen 2014 against DQGen 2012, five human judges categorized candidate choices without knowing their intended type or whether they were the correct answer or a distractor generated by DQGen 2012 or DQGen 2014. The percentage of distractors categorized as their intended type was significantly higher for DQGen 2014.<\/jats:p><\/jats:list-item><jats:list-item><jats:label>(4)<\/jats:label><jats:p>We evaluated DQGen 2014 against human performance based on 1,486 similarly blind categorizations by twenty-seven judges of sixteen correct answers, forty-eight distractors generated by DQGen 2014, and 504 distractors authored by twenty-one humans. Surprisingly, DQGen 2014 did significantly better than humans at generating ungrammatical distractors and marginally better than humans at generating nonsensical distractors, albeit slightly worse at generating plausible distractors. Moreover, vetting DQGen 2014's output and writing distractors only when necessary would halve the time to write them all, and produce higher quality distractors.<\/jats:p><\/jats:list-item><\/jats:list><\/jats:p>","DOI":"10.1017\/s1351324916000024","type":"journal-article","created":{"date-parts":[[2016,4,14]],"date-time":"2016-04-14T09:43:04Z","timestamp":1460626984000},"page":"245-294","source":"Crossref","is-referenced-by-count":12,"title":["Developing, evaluating, and refining an automatic generator of diagnostic multiple choice cloze questions to assess children's comprehension while reading"],"prefix":"10.1017","volume":"23","author":[{"given":"JACK","family":"MOSTOW","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"YI-TING","family":"HUANG","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"HYEJU","family":"JANG","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"ANDERS","family":"WEINSTEIN","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"JOE","family":"VALERI","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"DONNA","family":"GATES","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2016,4,14]]},"reference":[{"key":"S1351324916000024_ref058","unstructured":"Zhang X. , Mostow J. , and Beck J. E. 2007, July 9\u201313. Can a computer listen for fluctuations in reading comprehension?. In R. Luckin, K. R. Koedinger, and J. Greer (eds.), Proceedings of the 13th International Conference on Artificial Intelligence in Education, pp. 495\u2013502. Marina del Rey, CA: IOS Press."},{"key":"S1351324916000024_ref056","first-page":"131","volume-title":"The Psychology of Science Text Comprehension","author":"van den Broek","year":"2002"},{"key":"S1351324916000024_ref048","unstructured":"Rus V. , Wyse B. , Piwek P. , Lintean M. , Stoyanchev S. , and Moldovan C. 2010. The first question generation shared task evaluation challenge. In Proceedings of the 6th International Natural Language Generation Conference, pp. 251\u20137. Dublin, Ireland, Association for Computational Linguistics."},{"key":"S1351324916000024_ref021","unstructured":"Heilman M. , and Smith N. A. 2010, June. Good question! Statistical ranking for question generation. In Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the ACL, pp. 609\u201317. Los Angeles, CA, Association for Computational Linguistics."},{"key":"S1351324916000024_ref009","volume-title":"How to Prepare Better Multiple-Choice Test Items: Guidelines for University Faculty","author":"Burton","year":"1991"},{"key":"S1351324916000024_ref033","doi-asserted-by":"crossref","unstructured":"Li L. , and Sporleder C. 2009. Classifier combination for contextual idiom detection without labelled data, In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pp. 315\u201323. Singapore, Association for Computational Linguistics.","DOI":"10.3115\/1699510.1699552"},{"key":"S1351324916000024_ref049","doi-asserted-by":"publisher","DOI":"10.1037\/0033-2909.86.2.420"},{"key":"S1351324916000024_ref020","doi-asserted-by":"publisher","DOI":"10.21236\/ADA531042"},{"key":"S1351324916000024_ref040","unstructured":"Mostow J. , Beck J. E. , Bey J. , Cuneo A. , Sison J. , Tobin B. , and Valeri J. 2004. Using automated questions to assess reading comprehension, vocabulary, and effects of tutorial interventions. Technology, Instruction, Cognition and Learning 2 (1\u20132): 97\u2013134."},{"key":"S1351324916000024_ref032","unstructured":"Li L. , Roth B. , and Sporleder C. 2010. Topic models for word sense disambiguation and token-based idiom detection. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, pp. 1138\u201347. Uppsala, Sweden, Association for Computational Linguistics."},{"key":"S1351324916000024_ref015","volume-title":"The Encyclopedia of Applied Linguistics","author":"Fellbaum","year":"2012"},{"key":"S1351324916000024_ref013","doi-asserted-by":"crossref","unstructured":"Coniam D. 1997. A preliminary inquiry into using corpus word frequency data in the automatic generation of english language cloze tests. CALICO Journal 14 (2\u20134): 15\u201333.","DOI":"10.1558\/cj.v14i2-4.15-33"},{"key":"S1351324916000024_ref007","volume-title":"Words Worth Teaching: Closing the Vocabulary Gap","author":"Biemiller","year":"2009"},{"key":"S1351324916000024_ref003","first-page":"27","volume-title":"The 7th International Conference on NLP","author":"Aldabe","year":"2010"},{"key":"S1351324916000024_ref017","unstructured":"Goto T. , Kojiri T. , Watanabe T. , Iwata T. , and Yamada T. 2010. Automatic generation system of multiple-choice cloze questions and its evaluation. Knowledge Management & E-Learning: An International Journal (KM& EL) 2 (3): 210\u201324."},{"key":"S1351324916000024_ref023","doi-asserted-by":"crossref","unstructured":"Huang Y.-T. , Chen M. C. , and Sun Y. S. 2012, November 26\u201330. Personalized automatic quiz generation based on proficiency level estimation. In Proceedings of the 20th International Conference on Computers in Education (ICCE 2012), pp. 553\u201360. Singapore.","DOI":"10.58459\/icce.2012.880"},{"key":"S1351324916000024_ref052","doi-asserted-by":"crossref","unstructured":"Sumita E. , Sugaya F. , and Yamamoto S. 2005. Measuring non-native speakers\u2019 proficiency of english by using a test with automatically-generated fill-in-the-blank questions. In Proceedings of the Second Workshop on Building Educational Applications Using NLP, pp. 61\u20138. Ann Arbor, Michigan, Association for Computational Linguistics.","DOI":"10.3115\/1609829.1609839"},{"key":"S1351324916000024_ref022","doi-asserted-by":"crossref","unstructured":"Hensler B. S. , and Beck J. E. 2006, June 26\u201330. Better student assessing by finding difficulty factors in a fully automated comprehension measure [best paper nominee]. In K. Ashley and M. Ikeda (eds.), Proceedings of the 8th International Conference on Intelligent Tutoring Systems, pp. 21\u201330. Jhongli, Taiwan, Springer-Verlag.","DOI":"10.1007\/11774303_3"},{"key":"S1351324916000024_ref011","unstructured":"Chang K.-M. , Nelson J. , Pant U. , and Mostow J. 2013. Toward exploiting eeg input in a reading tutor. International Journal of Artificial Intelligence in Education 22(1, \u201cBest of AIED2011 Part 1\u201d): 29\u201341."},{"key":"S1351324916000024_ref050","volume-title":"Third International Workshop on Parsing Technologies","author":"Sleator","year":"1993"},{"key":"S1351324916000024_ref036","doi-asserted-by":"publisher","DOI":"10.1109\/TLT.2012.5"},{"key":"S1351324916000024_ref047","unstructured":"Raghunathan K. , Lee H. , Rangarajan S. , Chambers N. , Surdeanu M. , Jurafsky D. , and Manning C. 2010. A multi-pass sieve for coreference resolution. In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, pp. 492\u2013501. MIT, Cambridge, MA, Association for Computational Linguistics."},{"key":"S1351324916000024_ref012","unstructured":"Chen W. , Mostow J. , and Aist G. S. 2013. Recognizing young readers\u2019 spoken questions. International Journal of Artificial Intelligence in Education 21 (4): 255\u201369."},{"key":"S1351324916000024_ref053","doi-asserted-by":"crossref","unstructured":"Tapanainen P. , and J\u00e4rvinen T. 1997. A non-projective dependency parser. In Proceedings of the 5th Conference on Applied Natural Language Processing, pp. 64\u201371. Washington, DC, Association for Computational Linguistics.","DOI":"10.3115\/974557.974568"},{"key":"S1351324916000024_ref027","doi-asserted-by":"crossref","unstructured":"Klein D. , and Manning C. D. 2003, July 7\u201312. Accurate unlexicalized parsing. In E. W. Hinrichs and D. Roth (eds.), Proceedings of the 41st Annual Meeting of the Association for Computational Linguistics, pp. 423\u201330. Sapporo, Japan, Association for Computational Linguistics.","DOI":"10.3115\/1075096.1075150"},{"key":"S1351324916000024_ref014","unstructured":"Correia R. , Baptista J. , Mamede N. , Trancoso I. , and Eskenazi M. 2010, September 22\u201324. Automatic generation of cloze question distractors. In Proceedings of the Interspeech 2010 Satellite Workshop on Second Language Studies: Acquisition, Learning, Education and Technology, Waseda University, Tokyo, Japan."},{"key":"S1351324916000024_ref018","doi-asserted-by":"publisher","DOI":"10.1207\/s1532799xssr0203_4"},{"key":"S1351324916000024_ref051","unstructured":"Smith S. , Sommers S. , and Kilgarriff A. 2008. Learning words right with the sketch engine and webbootcat: automatic cloze generation from corpora and the web. In Proceedings of the 25th International Conference of English Teaching and Learning & 2008 International Conference on English Instruction and Assessment, pp. 1\u20138. Lisbon, Portugal."},{"key":"S1351324916000024_ref057","doi-asserted-by":"crossref","unstructured":"Zesch T. , and Melamud O. 2014. Automatic generation of challenging distractors using context-sensitive inference rules. In Workshop on Innovative Use of NLP for Building Educational Applications (BEA), pp. 143\u20138. Baltimore, MD.","DOI":"10.3115\/v1\/W14-1817"},{"key":"S1351324916000024_ref045","unstructured":"Pino J. , Heilman M. , and Eskenazi M. 2008. A selection strategy to improve cloze question quality. In Proceedings of the Workshop on Intelligent Tutoring Systems for Ill-Defined Domains. 9th International Conference on Intelligent Tutoring Systems, pp. 22\u201334. Montreal, Canada."},{"key":"S1351324916000024_ref038","doi-asserted-by":"crossref","unstructured":"Mitkov R. , Ha L. A. , Varga A. , and Rello L. 2009, March 31. Semantic similarity of distractors in multiple-choice tests: extrinsic evaluation. In R. Basili and M. Pennacchiotti (eds.), EACL 2009 Workshop on GEMS: GEometrical Models of Natural Language Semantics, pp. 49\u201356. Athens, Greece, Association for Computational Linguistics.","DOI":"10.3115\/1705415.1705422"},{"key":"S1351324916000024_ref026","doi-asserted-by":"crossref","unstructured":"Kintsch W. 2005. An overview of top-down and bottom-up effects in comprehension: the ci perspective. Discourse Processes 39 (2\u20133): 125\u20138.","DOI":"10.1080\/0163853X.2005.9651676"},{"key":"S1351324916000024_ref008","doi-asserted-by":"crossref","unstructured":"Brown J. C. , Frishkoff G. A. , and Eskenazi M. 2005, October 6\u20138. Automatic question generation for vocabulary assessment. In Proceedings of the Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing, pp. 819\u201326. Vancouver, BC, Canada. Stroudsburg, PA, USA: Association for Computational Linguistics.","DOI":"10.3115\/1220575.1220678"},{"key":"S1351324916000024_ref037","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324906004177"},{"key":"S1351324916000024_ref025","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177732186"},{"key":"S1351324916000024_ref002","author":"Agarwal","year":"2011"},{"key":"S1351324916000024_ref039","doi-asserted-by":"crossref","unstructured":"Mostow J. 2013, July. Lessons from project listen: what have we learned from a reading tutor that listens? (keynote). In H. C. Lane, K. Yacef, J. Mostow, and P. Pavlik (eds.), Proceedings of the 16th International Conference on Artificial Intelligence in Education, pp. 557\u20138. Memphis, TN, LNAI, vol. 7926. Springer.","DOI":"10.1007\/978-3-642-39112-5_58"},{"key":"S1351324916000024_ref005","first-page":"656","volume-title":"Proceedings of the 14th International Conference on Artificial Intelligence in Education (AIED2009)","author":"Aldabe","year":"2009"},{"key":"S1351324916000024_ref031","doi-asserted-by":"crossref","unstructured":"Lee J. , and Seneff S. 2007, August 27\u201331. Automatic generation of cloze items for prepositions. In Proceedings of INTERSPEECH, pp. 2173\u20136. Antwerp, Belgium,","DOI":"10.21437\/Interspeech.2007-592"},{"key":"S1351324916000024_ref030","doi-asserted-by":"publisher","DOI":"10.2307\/2529310"},{"key":"S1351324916000024_ref035","doi-asserted-by":"crossref","unstructured":"Liu C.-L. , Wang C.-H. , Gao Z.-M. , and Huang S.-M. 2005, June 29. Applications of lexical information for algorithmically composing multiple-choice cloze items. In Proceedings of the Second Workshop on Building Educational Applications Using NLP, Ann Arbor, Michigan, pp. 1\u20138. Stroudsburg, PA: Association for Computational Linguistics.","DOI":"10.3115\/1609829.1609830"},{"key":"S1351324916000024_ref006","unstructured":"Becker L. , Basu S. , and Vanderwende L. 2012. Mind the gap: learning to choose gaps for question generation. In Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 742\u201351. Montreal, Canada: Association for Computational Linguistics."},{"key":"S1351324916000024_ref055","unstructured":"Unspecified. 2006. Tiny invaders, National Geographic Explorer (Pioneer Edition) http:\/\/ngexplorer.cengage.com\/pioneer\/."},{"key":"S1351324916000024_ref010","doi-asserted-by":"publisher","DOI":"10.1021\/ed061p613"},{"key":"S1351324916000024_ref046","doi-asserted-by":"publisher","DOI":"10.5087\/dad.2012.201"},{"key":"S1351324916000024_ref019","doi-asserted-by":"publisher","DOI":"10.1207\/S15324818AME1503_5"},{"key":"S1351324916000024_ref028","unstructured":"Kolb P. 2008. Disco: a multilingual database of distributionally similar words. In Proceedings of KONVENS-2008 (Konferenz zur Verarbeitung nat\u00fcrlicher Sprache), pp. 5\u201312. Berlin."},{"key":"S1351324916000024_ref001","author":"Agarwal","year":"2011"},{"key":"S1351324916000024_ref016","unstructured":"Gates D. , Aist G. , Mostow J. , Mckeown M. , and Bey J. 2011. How to generate cloze questions from definitions: a syntactic approach. In Proceedings of the AAAI Symposium on Question Generation, pp. 19\u201322. Arlington, VA, AAAI Press."},{"key":"S1351324916000024_ref024","doi-asserted-by":"crossref","unstructured":"Huang Y.-T. , and Mostow J. 2015, June 22\u201326. Evaluating human and automated generation of distractors for diagnostic multiple-choice cloze questions to assess children\u2019s reading comprehension. In C. Conati , N. Heffernan , A. Mitrovic , and M. F. Verdejo (eds.), Proceedings of the 17th International Conference on Artificial Intelligence in Education, pp. 155\u201364. Madrid, Spain, Lecture Notes in Computer Science, vol. 9112. Switzerland: Springer International Publishing.","DOI":"10.1007\/978-3-319-19773-9_16"},{"key":"S1351324916000024_ref004","first-page":"7","volume-title":"The Workshop on NLP for Educational Resources. In conjunction with RANLP07","author":"Aldabe","year":"2007"},{"key":"S1351324916000024_ref042","unstructured":"Mostow J. , and Jang H. 2012, June 7. Generating diagnostic multiple choice comprehension cloze questions. In NAACL-HLT 2012 7th Workshop on Innovative Use of NLP for Building Educational Applications, pp. 136\u201346. Montr\u00e9al, Association for Computational Linguistics."},{"key":"S1351324916000024_ref044","first-page":"13","volume-title":"Children\u2019s Reading Comprehension and Assessment","author":"Pearson","year":"2005"},{"key":"S1351324916000024_ref054","doi-asserted-by":"crossref","unstructured":"Toutanova K. , Klein D. , Manning C. , and Singer Y. 2003. Feature-rich part-of-speech tagging with a cyclic dependency network. Proceedings of the Human Language Technology Conference and Annual Meeting of the North American Chapter of the Association for Computational Linguistics (HLT-NAACL), Edmonton, Canada, pp. 252\u20139.","DOI":"10.3115\/1073445.1073478"},{"key":"S1351324916000024_ref029","unstructured":"Kolb P. 2009. Experiments on the difference between semantic similarity and relatedness. In Proceedings of the 17th Nordic Conference on Computational Linguistics-NODALIDA\u201909, Odense, Denmark."},{"key":"S1351324916000024_ref034","author":"Lin","year":"2007"},{"key":"S1351324916000024_ref043","unstructured":"Niraula N. B. , Rus V. , Stefanescu D. , and Graesser A. C. 2014. Mining gap-fill questions from tutorial dialogues. In Proceedings of the 7th International Conference on Educational Data Mining, pp. 265\u20138. London, UK."},{"key":"S1351324916000024_ref041","unstructured":"Mostow J. , and Chen W. 2009, July 6\u201310. Generating instruction automatically for the reading strategy of self-questioning. In V. Dimitrova , R. Mizoguchi , B. D. Boulay , and A. Graesser (eds.), Proceedings of the 14th International Conference on Artificial Intelligence in Education, pp. 465\u201372. Brighton, UK: IOS Press."}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324916000024","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,2]],"date-time":"2025-06-02T17:53:21Z","timestamp":1748886801000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324916000024\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,4,14]]},"references-count":58,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2017,3]]}},"alternative-id":["S1351324916000024"],"URL":"https:\/\/doi.org\/10.1017\/s1351324916000024","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,4,14]]}}}