{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T00:56:26Z","timestamp":1784336186258,"version":"3.55.0"},"reference-count":332,"publisher":"Association for Computing Machinery (ACM)","issue":"10","license":[{"start":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T00:00:00Z","timestamp":1675296000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2023,10,31]]},"abstract":"<jats:p>Alongside huge volumes of research on deep learning models in NLP in the recent years, there has been much work on benchmark datasets needed to track modeling progress. Question answering and reading comprehension have been particularly prolific in this regard, with more than 80 new datasets appearing in the past 2 years. This study is the largest survey of the field to date. We provide an overview of the various formats and domains of the current resources, highlighting the current lacunae for future work. We further discuss the current classifications of \u201cskills\u201d that question answering\/reading comprehension systems are supposed to acquire and propose a new taxonomy. The supplementary materials survey the current multilingual resources and monolingual resources for languages other than English, and we discuss the implications of overfocusing on English. The study is aimed at both practitioners looking for pointers to the wealth of existing data and at researchers working on new resources.<\/jats:p>","DOI":"10.1145\/3560260","type":"journal-article","created":{"date-parts":[[2022,9,13]],"date-time":"2022-09-13T14:12:27Z","timestamp":1663078347000},"page":"1-45","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":105,"title":["QA Dataset Explosion: A Taxonomy of NLP Resources for Question Answering and Reading Comprehension"],"prefix":"10.1145","volume":"55","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4845-4023","authenticated-orcid":false,"given":"Anna","family":"Rogers","sequence":"first","affiliation":[{"name":"University of Copenhagen, Copenhagen K, Denmark and RIKEN"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8458-1727","authenticated-orcid":false,"given":"Matt","family":"Gardner","sequence":"additional","affiliation":[{"name":"Microsoft Semantic Machines, Washington, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1562-7909","authenticated-orcid":false,"given":"Isabelle","family":"Augenstein","sequence":"additional","affiliation":[{"name":"University of Copenhagen, Copenhagen K, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,2,2]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-6130"},{"key":"e_1_3_3_3_2","first-page":"307","volume-title":"Proceedings of the 17th Annual Conference of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201919)","author":"Abujabal Abdalghani","year":"2019","unstructured":"Abdalghani Abujabal, Rishiraj Saha Roy, Mohamed Yahya, and Gerhard Weikum. 2019. ComQA: A community-sourced dataset for complex factoid question answering with paraphrase clusters. In Proceedings of the 17th Annual Conference of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201919). 307\u2013317. https:\/\/aclweb.org\/anthology\/papers\/N\/N19\/N19-1027\/."},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1194"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018076"},{"key":"e_1_3_3_6_2","unstructured":"Douglas Adams. 2009. The Hitchhiker\u2019s Guide to the Galaxay . Ballantine Books New York NY. PR6051.D3352 H5 2009"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00471"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.164"},{"key":"e_1_3_3_9_2","volume-title":"Dialogue Acts in Verbmobil 2","author":"Alexandersson Jan","year":"1998","unstructured":"Jan Alexandersson, Bianka Buschbeck-Wolf, Tsutomu Fujinami, Michael Kipp, Stephan Koch, Elisabeth Maier, Norbert Reithinger, Birte Schmitz, and Melanie Siegel. 1998. Dialogue Acts in Verbmobil 2. Technical Report. Verbmobil."},{"key":"e_1_3_3_10_2","first-page":"520","volume-title":"Proceedings of the 19th Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201921)","author":"Anantha Raviteja","year":"2021","unstructured":"Raviteja Anantha, Svitlana Vakulenko, Zhucheng Tu, Shayne Longpre, Stephen Pulman, and Srinivas Chappidi. 2021. Open-domain question answering goes conversational via question rewriting. In Proceedings of the 19th Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201921). 520\u2013534. https:\/\/aclanthology.org\/2021.naacl-main.44."},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.421"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.118"},{"key":"e_1_3_3_14_2","article-title":"Multilingual extractive reading comprehension by runtime machine translation","author":"Asai Akari","year":"2018","unstructured":"Akari Asai, Akiko Eriguchi, Kazuma Hashimoto, and Yoshimasa Tsuruoka. 2018. Multilingual extractive reading comprehension by runtime machine translation. arXiv:1809.03275 [CS] (2018). http:\/\/arxiv.org\/abs\/1809.03275.","journal-title":"arXiv:1809.03275 [CS]"},{"key":"e_1_3_3_15_2","article-title":"XOR QA: Cross-lingual open-retrieval question answering","author":"Asai Akari","year":"2020","unstructured":"Akari Asai, Jungo Kasai, Jonathan H. Clark, Kenton Lee, Eunsol Choi, and Hannaneh Hajishirzi. 2020. XOR QA: Cross-lingual open-retrieval question answering. arXiv:2010.11856 [CS] (2020). http:\/\/arxiv.org\/abs\/2010.11856.","journal-title":"arXiv:2010.11856 [CS]"},{"key":"e_1_3_3_16_2","article-title":"Frames: A corpus for adding memory to goal-oriented dialogue systems","author":"Asri Layla El","year":"2017","unstructured":"Layla El Asri, Hannes Schulz, Shikhar Sharma, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, and Kaheer Suleman. 2017. Frames: A corpus for adding memory to goal-oriented dialogue systems. arXiv:1704.00057 [CS] (2017). http:\/\/arxiv.org\/abs\/1704.00057.","journal-title":"arXiv:1704.00057 [CS]"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.656"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1475"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1002\/9781118784235.eelt0369"},{"key":"e_1_3_3_20_2","article-title":"MS MARCO: A human generated machine reading comprehension dataset","author":"Bajaj Payal","year":"2016","unstructured":"Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, et\u00a0al. 2016. MS MARCO: A human generated machine reading comprehension dataset. arXiv:1611.09268 [CS] (2016). http:\/\/arxiv.org\/abs\/1611.09268.","journal-title":"arXiv:1611.09268 [CS]"},{"key":"e_1_3_3_21_2","volume-title":"Proceedings of the 5th International Conference on Learning Representations (ICLR\u201917)","author":"Bajgar Ondrej","year":"2017","unstructured":"Ondrej Bajgar, Rudolf Kadlec, and Jan Kleindienst. 2017. Embracing data abundance: BookTest dataset for reading comprehension. In Proceedings of the 5th International Conference on Learning Representations (ICLR\u201917). https:\/\/openreview.net\/pdf?id=H1U4mhVFe."},{"key":"e_1_3_3_22_2","first-page":"56","volume-title":"Proceedings of the Workshop on Modeling, Learning, and Mining for Cross\/Multilinguality (MultiLingMine\u201916) Co-located with the 2016 European Conference on Information Retrieval (ECIR\u201916)","author":"Banerjee Somnath","year":"2016","unstructured":"Somnath Banerjee, Sudip Kumar Naskar, and Paolo Rosso. 2016. The first cross-script code-mixed question answering corpus. In Proceedings of the Workshop on Modeling, Learning, and Mining for Cross\/Multilinguality (MultiLingMine\u201916) Co-located with the 2016 European Conference on Information Retrieval (ECIR\u201916). 1\u201310. 56\u201365. http:\/\/ceur-ws.org\/Vol-1589\/MultiLingMine6.pdf."},{"key":"e_1_3_3_23_2","volume-title":"The Stanford Encyclopedia of Philosophy (Spring 2019 ed.)","author":"Bartha Paul","year":"2019","unstructured":"Paul Bartha. 2019. Analogy and analogical reasoning. In The Stanford Encyclopedia of Philosophy (Spring 2019 ed.), Edward N. Zalta (Ed.). Metaphysics Research Lab, Stanford University, Stanford, CA. https:\/\/plato.stanford.edu\/archives\/spr2019\/entries\/reasoning-analogy\/."},{"key":"e_1_3_3_24_2","unstructured":"Emily M. Bender. 2019. The #BenderRule: On naming the languages we study and why it matters. The Gradient . Retrieved September 16 2022 from https:\/\/thegradient.pub\/the-benderrule-on-naming-the-languages-we-study-and-why-it-matters\/."},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00041"},{"key":"e_1_3_3_26_2","first-page":"1533","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201913)","author":"Berant Jonathan","year":"2013","unstructured":"Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on Freebase from question-answer pairs. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201913). 1533\u20131544. https:\/\/www.aclweb.org\/anthology\/D13-1160."},{"key":"e_1_3_3_27_2","first-page":"1499","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914)","author":"Berant Jonathan","year":"2014","unstructured":"Jonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, and Christopher D. Manning. 2014. Modeling biological processes for reading comprehension. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914). 1499\u20131510."},{"key":"e_1_3_3_28_2","doi-asserted-by":"crossref","unstructured":"Yevgeni Berzak Jonathan Malmaud and Roger Levy. 2020. STARC: Structured annotations for reading comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920) . 5726\u20135735. https:\/\/www.aclweb.org\/anthology\/2020.acl-main.507.","DOI":"10.18653\/v1\/2020.acl-main.507"},{"key":"e_1_3_3_29_2","unstructured":"Chandra Bhagavatula Ronan Le Bras Chaitanya Malaviya Keisuke Sakaguchi Ari Holtzman Hannah Rashkin Doug Downey Wen-Tau Yih and Yejin Choi. 2019. Abductive commonsense reasoning. In Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919) . https:\/\/openreview.net\/forum?id=Byg1v1HKDB."},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.703"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6239"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.442"},{"key":"e_1_3_3_33_2","unstructured":"Simon Blackburn. 2008. Inference. Retrieved September 16 2022 from https:\/\/www.oxfordreference.com\/view\/10.1093\/acref\/9780199541430.001.0001\/acref-9780199541430."},{"key":"e_1_3_3_34_2","unstructured":"Simon Blackburn. 2008. Reasoning. Retrieved September 16 2022 from https:\/\/www.oxfordreference.com\/view\/10.1093\/acref\/9780199541430.001.0001\/acref-9780199541430."},{"key":"e_1_3_3_35_2","unstructured":"Elisa Bone and Mike Prosser. 2020. Multiple Choice Questions: An Introductory Guide. Retrieved September 16 2022 from https:\/\/melbourne-cshe.unimelb.edu.au\/__data\/assets\/pdf_file\/0010\/3430648\/multiple-choice-questions_final.pdf."},{"key":"e_1_3_3_36_2","article-title":"Large-scale simple question answering with memory networks","author":"Bordes Antoine","year":"2015","unstructured":"Antoine Bordes, Nicolas Usunier, Sumit Chopra, and Jason Weston. 2015. Large-scale simple question answering with memory networks. arXiv:1506.02075 [CS] (2015). http:\/\/arxiv.org\/abs\/1506.02075.","journal-title":"arXiv:1506.02075 [CS]"},{"key":"e_1_3_3_37_2","article-title":"What question answering can learn from trivia nerds","author":"Boyd-Graber Jordan","year":"2019","unstructured":"Jordan Boyd-Graber. 2019. What question answering can learn from trivia nerds. arXiv:1910.14464 [CS] (2019). http:\/\/arxiv.org\/abs\/1910.14464.","journal-title":"arXiv:1910.14464 [CS]"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-94042-7_9"},{"key":"e_1_3_3_39_2","article-title":"Language models are few-shot learners","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, et\u00a0al. 2020. Language models are few-shot learners. arXiv:2005.14165 [CS] (2020). http:\/\/arxiv.org\/abs\/2005.14165.","journal-title":"arXiv:2005.14165 [CS]"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1547"},{"issue":"2","key":"e_1_3_3_41_2","first-page":"23","article-title":"A review of public datasets in question answering research","volume":"54","author":"Cambazoglu B. Barla","year":"2020","unstructured":"B. Barla Cambazoglu, Mark Sanderson, Falk Scholer, and Bruce Croft. 2020. A review of public datasets in question answering research. ACM SIGIR Forum 54, 2 (2020), 23. http:\/\/www.sigir.org\/wp-content\/uploads\/2020\/12\/p07.pdf.","journal-title":"ACM SIGIR Forum"},{"key":"e_1_3_3_42_2","doi-asserted-by":"crossref","first-page":"7302","DOI":"10.18653\/v1\/2020.acl-main.652","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920)","author":"Campos Jon Ander","year":"2020","unstructured":"Jon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu, Mark Cieliebak, and Eneko Agirre. 2020. DoQA-accessing domain-specific FAQs via conversational QA. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920). 7302\u20137314. https:\/\/aclanthology.org\/2020.acl-main.652\/."},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.422"},{"key":"e_1_3_3_44_2","article-title":"Automatic Spanish translation of the SQuAD dataset for multilingual question answering","author":"Carrino Casimiro Pio","year":"2019","unstructured":"Casimiro Pio Carrino, Marta R. Costa-Juss\u00e0, and Jos\u00e9 A. R. Fonollosa. 2019. Automatic Spanish translation of the SQuAD dataset for multilingual question answering. arXiv:1912.05200 [CS] (2019). http:\/\/arxiv.org\/abs\/1912.05200.","journal-title":"arXiv:1912.05200 [CS]"},{"key":"e_1_3_3_45_2","doi-asserted-by":"crossref","first-page":"1269","DOI":"10.18653\/v1\/2020.acl-main.117","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920)","author":"Castelli Vittorio","year":"2020","unstructured":"Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg, et\u00a0al. 2020. The TechQA dataset. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920). 1269\u20131278. https:\/\/www.aclweb.org\/anthology\/2020.acl-main.117."},{"key":"e_1_3_3_46_2","article-title":"Evaluation of text generation: A survey","author":"Celikyilmaz Asli","year":"2020","unstructured":"Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020. Evaluation of text generation: A survey. arXiv:2006.14799 [CS] (2020). http:\/\/arxiv.org\/abs\/2006.14799.","journal-title":"arXiv:2006.14799 [CS]"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-3204"},{"key":"e_1_3_3_48_2","unstructured":"Graham Chapman John Cleese Terry Gilliam Eric Idle Terry Jones Michael Palin John Goldstone Spike Milligan Monty Python (Comedy troupe) Handmade Films and Criterion Collection (Firm). 1999. Life of Brian ."},{"key":"e_1_3_3_49_2","first-page":"1135","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Chattopadhyay Prithvijit","year":"2017","unstructured":"Prithvijit Chattopadhyay, Ramakrishna Vedantam, Ramprasaath R. Selvaraju, Dhruv Batra, and Devi Parikh. 2017. Counting everyday objects in everyday scenes. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917). 1135\u20131144. https:\/\/openaccess.thecvf.com\/content_cvpr_2017\/html\/Chattopadhyay_Counting_Everyday_Objects_CVPR_2017_paper.html."},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.345"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-5817"},{"key":"e_1_3_3_52_2","first-page":"6521","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920)","author":"Chen Anthony","year":"2020","unstructured":"Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2020. MOCHA: A dataset for training and evaluating generative reading comprehension metrics. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920). 6521\u20136532. https:\/\/www.aclweb.org\/anthology\/2020.emnlp-main.528."},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1171"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-tutorials.8"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.91"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.343"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.300"},{"key":"e_1_3_3_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/371578.371581"},{"key":"e_1_3_3_59_2","first-page":"391","volume-title":"Proceedings of Machine Learning Research","author":"Cho Minseok","year":"2018","unstructured":"Minseok Cho, Reinald Kim Amplayo, Seung-Won Hwang, and Jonghyuck Park. 2018. Adversarial TableQA: Attention supervision for question answering on tables. In Proceedings of Machine Learning Research. 391\u2013406. http:\/\/proceedings.mlr.press\/v95\/cho18a\/cho18a.pdf."},{"key":"e_1_3_3_60_2","doi-asserted-by":"crossref","first-page":"2174","DOI":"10.18653\/v1\/D18-1241","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918)","author":"Choi Eunsol","year":"2018","unstructured":"Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-Tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. QuAC: Question answering in context. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918). 2174\u20132184. http:\/\/aclweb.org\/anthology\/D18-1241."},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00377"},{"key":"e_1_3_3_62_2","unstructured":"Sagnik Ray Choudhury Anna Rogers and Isabelle Augenstein. 2022. Machine reading fast and slow: When do models \u201cUnderstand\u201d language? In Proceedings of the 29th International Conference on Computational Linguistics . International Committee on Computational Linguistics 78\u201393. https:\/\/aclanthology.org\/2022.coling-1.8."},{"key":"e_1_3_3_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3358016"},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.493"},{"key":"e_1_3_3_65_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1078"},{"key":"e_1_3_3_66_2","first-page":"2924","volume-title":"Proceedings of the 17th Annual Conference of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201919)","author":"Clark Christopher","year":"2019","unstructured":"Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes\/no questions. In Proceedings of the 17th Annual Conference of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201919). 2924\u20132936. https:\/\/aclweb.org\/anthology\/papers\/N\/N19\/N19-1300\/."},{"key":"e_1_3_3_67_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00317"},{"key":"e_1_3_3_68_2","article-title":"Think you have solved question answering? Try ARC, the AI2 reasoning challenge","author":"Clark Peter","year":"2018","unstructured":"Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv:1803.05457 [CS] (2018). http:\/\/arxiv.org\/abs\/1803.05457.","journal-title":"arXiv:1803.05457 [CS]"},{"key":"e_1_3_3_69_2","first-page":"5450","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201920)","author":"Colas Anthony","year":"2020","unstructured":"Anthony Colas, Seokhwan Kim, Franck Dernoncourt, Siddhesh Gupte, Zhe Wang, and Doo Soon Kim. 2020. TutorialVQA: Question answering dataset for tutorial videos. In Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201920). 5450\u20135455. https:\/\/www.aclweb.org\/anthology\/2020.lrec-1.670."},{"key":"e_1_3_3_70_2","article-title":"Event-QA: A dataset for event-centric question answering over knowledge graphs","author":"Costa Tarc\u00edsio Souza","year":"2020","unstructured":"Tarc\u00edsio Souza Costa, Simon Gottschalk, and Elena Demidova. 2020. Event-QA: A dataset for event-centric question answering over knowledge graphs. arXiv:2004.11861 [CS] (2020). http:\/\/arxiv.org\/abs\/2004.11861.","journal-title":"arXiv:2004.11861 [CS]"},{"key":"e_1_3_3_71_2","doi-asserted-by":"publisher","DOI":"10.3233\/IA-190018"},{"key":"e_1_3_3_72_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1600"},{"key":"e_1_3_3_73_2","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201918)","author":"Cui Yiming","year":"2018","unstructured":"Yiming Cui, Ting Liu, Zhipeng Chen, Wentao Ma, Shijin Wang, and Guoping Hu. 2018. Dataset for the first evaluation on chinese machine reading comprehension. In Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201918). https:\/\/www.aclweb.org\/anthology\/L18-1431."},{"key":"e_1_3_3_74_2","first-page":"1777","volume-title":"Proceedings of the International Conference on Computational Linguistics (COLING\u201916)","author":"Cui Yiming","year":"2016","unstructured":"Yiming Cui, Ting Liu, Zhipeng Chen, Shijin Wang, and Guoping Hu. 2016. Consensus attention-based neural networks for chinese reading comprehension. In Proceedings of the International Conference on Computational Linguistics (COLING\u201916). 1777\u20131786. https:\/\/www.aclweb.org\/anthology\/C16-1167."},{"key":"e_1_3_3_75_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.589"},{"key":"e_1_3_3_76_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1606"},{"key":"e_1_3_3_77_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.365"},{"key":"e_1_3_3_78_2","volume-title":"The Stanford Encyclopedia of Philosophy (Summer 2019 ed.)","author":"Demey Lorenz","year":"2019","unstructured":"Lorenz Demey, Barteld Kooi, and Joshua Sack. 2019. Logic and probability. In The Stanford Encyclopedia of Philosophy (Summer 2019 ed.), Edward N. Zalta (Ed.). Metaphysics Research Lab, Stanford University, Stanford, CA. https:\/\/plato.stanford.edu\/archives\/sum2019\/entries\/logic-probability\/."},{"key":"e_1_3_3_79_2","article-title":"FQuAD: French question answering dataset","author":"d\u2019Hoffschmidt Martin","year":"2020","unstructured":"Martin d\u2019Hoffschmidt, Maxime Vidal, Wacim Belblidia, and Tom Brendl\u00e9. 2020. FQuAD: French question answering dataset. arXiv:2002.06071 [CS] (2020). http:\/\/arxiv.org\/abs\/2002.06071.","journal-title":"arXiv:2002.06071 [CS]"},{"key":"e_1_3_3_80_2","article-title":"Wizard of Wikipedia: Knowledge-powered conversational agents","author":"Dinan Emily","year":"2018","unstructured":"Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2018. Wizard of Wikipedia: Knowledge-powered conversational agents. arXiv:1811.01241 [CS] (2018). http:\/\/arxiv.org\/abs\/1811.01241.","journal-title":"arXiv:1811.01241 [CS]"},{"key":"e_1_3_3_81_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-2114"},{"key":"e_1_3_3_82_2","volume-title":"The Stanford Encyclopedia of Philosophy (Summer 2017 ed.)","author":"Douven Igor","year":"2017","unstructured":"Igor Douven. 2017. Abduction. In The Stanford Encyclopedia of Philosophy (Summer 2017 ed.), Edward N. Zalta (Ed.). Metaphysics Research Lab, Stanford University, Stanford, CA. https:\/\/plato.stanford.edu\/archives\/sum2017\/entries\/abduction\/."},{"key":"e_1_3_3_83_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-5820"},{"key":"e_1_3_3_84_2","first-page":"2368","volume-title":"Proceedings of the 17th Annual Conference of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201919)","author":"Dua Dheeru","year":"2019","unstructured":"Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 17th Annual Conference of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201919). 2368\u20132378. https:\/\/aclweb.org\/anthology\/papers\/N\/N19\/N19-1246\/."},{"key":"e_1_3_3_85_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-73618-1_86"},{"key":"e_1_3_3_86_2","doi-asserted-by":"crossref","first-page":"7839","DOI":"10.18653\/v1\/2020.acl-main.701","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920)","author":"Dunietz Jesse","year":"2020","unstructured":"Jesse Dunietz, Greg Burnham, Akash Bharadwaj, Owen Rambow, Jennifer Chu-Carroll, and Dave Ferrucci. 2020. To test machine comprehension, start by defining comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920). 7839\u20137859. https:\/\/www.aclweb.org\/anthology\/2020.acl-main.701."},{"key":"e_1_3_3_87_2","article-title":"SearchQA: A new Q&A dataset augmented with context from a search engine","author":"Dunn Matthew","year":"2017","unstructured":"Matthew Dunn, Levent Sagun, Mike Higgins, V. Ugur Guney, Volkan Cirik, and Kyunghyun Cho. 2017. SearchQA: A new Q&A dataset augmented with context from a search engine. arXiv:1704.05179 [CS] (2017). http:\/\/arxiv.org\/abs\/1704.05179.","journal-title":"arXiv:1704.05179 [CS]"},{"key":"e_1_3_3_88_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.693"},{"key":"e_1_3_3_89_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58219-7_1"},{"key":"e_1_3_3_90_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1605"},{"key":"e_1_3_3_91_2","article-title":"Key-value retrieval networks for task-oriented dialogue","author":"Eric Mihail","year":"2017","unstructured":"Mihail Eric and Christopher D. Manning. 2017. Key-value retrieval networks for task-oriented dialogue. arXiv:1705.05414 [CS] (2017). http:\/\/arxiv.org\/abs\/1705.05414.","journal-title":"arXiv:1705.05414 [CS]"},{"key":"e_1_3_3_92_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00298"},{"key":"e_1_3_3_93_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1346"},{"key":"e_1_3_3_94_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1091"},{"key":"e_1_3_3_95_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2020.3010650"},{"key":"e_1_3_3_96_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.570"},{"key":"e_1_3_3_97_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.86"},{"key":"e_1_3_3_98_2","article-title":"A dataset and baselines for visual question answering on art","author":"Garcia Noa","year":"2020","unstructured":"Noa Garcia, Chentao Ye, Zihua Liu, Qingtao Hu, Mayu Otani, Chenhui Chu, Yuta Nakashima, and Teruko Mitamura. 2020. A dataset and baselines for visual question answering on art. arXiv:2008.12520 [CS] (2020). http:\/\/arxiv.org\/abs\/2008.12520.","journal-title":"arXiv:2008.12520 [CS]"},{"key":"e_1_3_3_99_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.117"},{"key":"e_1_3_3_100_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1909.11291"},{"key":"e_1_3_3_101_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.135"},{"key":"e_1_3_3_102_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6282"},{"key":"e_1_3_3_103_2","article-title":"Datasheets for datasets","author":"Gebru Timnit","year":"2020","unstructured":"Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daum\u00e9 III, and Kate Crawford. 2020. Datasheets for datasets. arXiv:1803.09010 [CS] (2020). http:\/\/arxiv.org\/abs\/1803.09010.","journal-title":"arXiv:1803.09010 [CS]"},{"key":"e_1_3_3_104_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1456"},{"key":"e_1_3_3_105_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1107"},{"key":"e_1_3_3_106_2","article-title":"DaNetQA: A yes\/no question answering dataset for the russian language","author":"Glushkova Taisia","year":"2020","unstructured":"Taisia Glushkova, Alexey Machnev, Alena Fenogenova, Tatiana Shavrina, Ekaterina Artemova, and Dmitry I. Ignatov. 2020. DaNetQA: A yes\/no question answering dataset for the russian language. arXiv:2010.02605 [CS] (2020). http:\/\/arxiv.org\/abs\/2010.02605.","journal-title":"arXiv:2010.02605 [CS]"},{"key":"e_1_3_3_107_2","article-title":"Assessing BERT\u2019s syntactic abilities","author":"Goldberg Yoav","year":"2019","unstructured":"Yoav Goldberg. 2019. Assessing BERT\u2019s syntactic abilities. arXiv:1901.05287 [CS] (2019). http:\/\/arxiv.org\/abs\/1901.05287.","journal-title":"arXiv:1901.05287 [CS]"},{"key":"e_1_3_3_108_2","first-page":"2930","volume-title":"Findings of ACL-IJCNLP\u201921","author":"Gonz\u00e1lez Ana Valeria","year":"2021","unstructured":"Ana Valeria Gonz\u00e1lez, Anna Rogers, and Anders S\u00f8gaard. 2021. On the interaction of belief bias and explanations. In Findings of ACL-IJCNLP\u201921. 2930\u20132942. https:\/\/aclanthology.org\/2021.findings-acl.259."},{"key":"e_1_3_3_109_2","first-page":"394","volume-title":"Proceedings of the 1st Joint Conference on Lexical and Computational Semantics (*SEM\u201912)","author":"Gordon Andrew","year":"2012","unstructured":"Andrew Gordon, Zornitsa Kozareva, and Melissa Roemmele. 2012. SemEval-2012 Task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In Proceedings of the 1st Joint Conference on Lexical and Computational Semantics (*SEM\u201912). 394\u2013398. https:\/\/aclweb.org\/anthology\/papers\/S\/S12\/S12-1052\/."},{"key":"e_1_3_3_110_2","doi-asserted-by":"crossref","first-page":"4089","DOI":"10.1109\/CVPR.2018.00430","volume-title":"Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Gordon Daniel","year":"2018","unstructured":"Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi. 2018. IQA: Visual question answering in interactive environments. In Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918). 4089\u20134098."},{"key":"e_1_3_3_111_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3449992"},{"key":"e_1_3_3_112_2","article-title":"MultiReQA: A cross-domain evaluation for retrieval question answering models","author":"Guo Mandy","year":"2020","unstructured":"Mandy Guo, Yinfei Yang, Daniel Cer, Qinlan Shen, and Noah Constant. 2020. MultiReQA: A cross-domain evaluation for retrieval question answering models. arXiv:2005.02507 [CS] (2020). http:\/\/arxiv.org\/abs\/2005.02507.","journal-title":"arXiv:2005.02507 [CS]"},{"key":"e_1_3_3_113_2","first-page":"34","volume-title":"Proceedings of the IJCNLP\u201917, Shared Tasks","author":"Guo Shangmin","year":"2017","unstructured":"Shangmin Guo, Kang Liu, Shizhu He, Cao Liu, Jun Zhao, and Zhuoyu Wei. 2017. IJCNLP-2017 Task 5: Multi-choice question answering in examinations. In Proceedings of the IJCNLP\u201917, Shared Tasks. 34\u201340. https:\/\/www.aclweb.org\/anthology\/I17-4005."},{"key":"e_1_3_3_114_2","volume-title":"Findings of ACL 2021","author":"Gupta Aditya","year":"2021","unstructured":"Aditya Gupta, Jiacheng Xu, Shyam Upadhyay, Diyi Yang, and Manaal Faruqui. 2021. Disfl-QA: A benchmark dataset for understanding disfluencies in question answering. In Findings of ACL 2021. https:\/\/arxiv.org\/abs\/2106.04016."},{"key":"e_1_3_3_115_2","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201918)","author":"Gupta Deepak","year":"2018","unstructured":"Deepak Gupta, Surabhi Kumari, Asif Ekbal, and Pushpak Bhattacharyya. 2018. MMQA: A multi-domain multi-lingual question-answering framework for English and Hindi. In Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201918). https:\/\/www.aclweb.org\/anthology\/L18-1440."},{"key":"e_1_3_3_116_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/694"},{"key":"e_1_3_3_117_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-3205"},{"key":"e_1_3_3_118_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-2017"},{"key":"e_1_3_3_119_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.597"},{"key":"e_1_3_3_120_2","article-title":"ANTIQUE: A non-factoid question answering benchmark","author":"Hashemi Helia","year":"2019","unstructured":"Helia Hashemi, Mohammad Aliannejadi, Hamed Zamani, and W. Bruce Croft. 2019. ANTIQUE: A non-factoid question answering benchmark. arXiv:1905.08957 [CS] (2019). http:\/\/arxiv.org\/abs\/1905.08957.","journal-title":"arXiv:1905.08957 [CS]"},{"key":"e_1_3_3_121_2","first-page":"1803","volume-title":"Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD\u201917)","author":"Hassan Naeemul","year":"2017","unstructured":"Naeemul Hassan, Fatma Arslan, Chengkai Li, and Mark Tremayne. 2017. Toward automated fact-checking: Detecting check-worthy factual claims by ClaimBuster. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD\u201917). ACM, New York, NY, 1803\u20131812. http:\/\/dblp.uni-trier.de\/db\/conf\/kdd\/kdd2017.html#HassanALT17."},{"key":"e_1_3_3_122_2","volume-title":"The Stanford Encyclopedia of Philosophy (Spring 2022 ed.)","author":"Hawthorne James","year":"2021","unstructured":"James Hawthorne. 2021. Inductive logic. In The Stanford Encyclopedia of Philosophy (Spring 2022 ed.), Edward N. Zalta (Ed.). Metaphysics Research Lab, Stanford University, Stanford, CA. https:\/\/plato.stanford.edu\/archives\/spr2021\/entries\/logic-inductive\/."},{"key":"e_1_3_3_123_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-2605"},{"key":"e_1_3_3_124_2","first-page":"309","volume-title":"Proceedings of the International Conference on Speech Prosody","author":"Hedberg Nancy","year":"2004","unstructured":"Nancy Hedberg, Juan M. Sosa, and Lorna Fadden. 2004. Meanings and configurations of questions in English. In Proceedings of the International Conference on Speech Prosody. 309\u2013312. https:\/\/www.isca-speech.org\/archive\/sp2004\/papers\/sp04_309.pdf."},{"key":"e_1_3_3_125_2","volume-title":"Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley","author":"Hemphill Charles T.","year":"1990","unstructured":"Charles T. Hemphill, John J. Godfrey, and George R. Doddington. 1990. The ATIS spoken language systems pilot corpus. In Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24\u201327, 1990. https:\/\/www.aclweb.org\/anthology\/H90-1021."},{"key":"e_1_3_3_126_2","first-page":"1693","volume-title":"Proceedings of the Conference on Neural Information Processing Systems (NeurIPS\u201915)","author":"Hermann Karl Moritz","year":"2015","unstructured":"Karl Moritz Hermann, Tom\u00e1\u0161 Ko\u010disk\u00fd, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS\u201915). 1693\u20131701. http:\/\/dl.acm.org\/citation.cfm?id=2969239.2969428."},{"key":"e_1_3_3_127_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1145"},{"key":"e_1_3_3_128_2","article-title":"The goldilocks principle: Reading children\u2019s books with explicit memory representations","author":"Hill Felix","year":"2015","unstructured":"Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2015. The goldilocks principle: Reading children\u2019s books with explicit memory representations. arXiv:1511.02301 [CS] (2015). http:\/\/arxiv.org\/abs\/1511.02301.","journal-title":"arXiv:1511.02301 [CS]"},{"key":"e_1_3_3_129_2","first-page":"1753","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201920)","author":"Horbach Andrea","year":"2020","unstructured":"Andrea Horbach, Itziar Aldabe, Marie Bexte, Oier Lopez de Lacalle, and Montse Maritxalar. 2020. Linguistic appropriateness and pedagogic usefulness of reading comprehension questions. In Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201920). 1753\u20131762. https:\/\/www.aclweb.org\/anthology\/2020.lrec-1.217."},{"key":"e_1_3_3_130_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1243"},{"key":"e_1_3_3_131_2","first-page":"6700","volume-title":"Proceedings of the 2019 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201919)","author":"Hudson Drew A.","year":"2019","unstructured":"Drew A. Hudson and Christopher D. Manning. 2019. GQA: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the 2019 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201919). 6700\u20136709. https:\/\/openaccess.thecvf.com\/content_CVPR_2019\/html\/Hudson_GQA_A_New_Dataset_for_Real-World_Visual_Reasoning_and_Compositional_CVPR_2019_paper.html."},{"key":"e_1_3_3_132_2","unstructured":"SRI International. 2011. SRI\u2019s Amex Travel Agent Data. Retrieved September 16 2022 from http:\/\/www.ai.sri.com\/~communic\/amex\/amex.html."},{"key":"e_1_3_3_133_2","first-page":"2758","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Jang Yunseok","year":"2017","unstructured":"Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017. TGIF-QA: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917). 2758\u20132766. https:\/\/openaccess.thecvf.com\/content_cvpr_2017\/html\/Jang_TGIF-QA_Toward_Spatio-Temporal_CVPR_2017_paper.html."},{"key":"e_1_3_3_134_2","volume-title":"Proceedings of the Text REtrival Conference (TREC\u201919)","author":"Jeffrey Dalton","year":"2019","unstructured":"Dalton Jeffrey, Xiong Chenyan, and Callan Jamie. 2019. CAsT 2019: The conversational assistance track overview. In Proceedings of the Text REtrival Conference (TREC\u201919)."},{"key":"e_1_3_3_135_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1215"},{"key":"e_1_3_3_136_2","doi-asserted-by":"publisher","DOI":"10.1145\/3184558.3191536"},{"key":"e_1_3_3_137_2","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3269247"},{"key":"e_1_3_3_138_2","first-page":"318","volume-title":"Proceedings of the 17th Annual Conference of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201919)","author":"Jiang Kelvin","year":"2019","unstructured":"Kelvin Jiang, Dekun Wu, and Hui Jiang. 2019. FreebaseQA: A new factoid QA data set matching trivia-style question-answer pairs with Freebase. In Proceedings of the 17th Annual Conference of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201919). 318\u2013323. https:\/\/aclweb.org\/anthology\/papers\/N\/N19\/N19-1028\/."},{"key":"e_1_3_3_139_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.443"},{"key":"e_1_3_3_140_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1259"},{"key":"e_1_3_3_141_2","first-page":"2901","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Johnson Justin","year":"2017","unstructured":"Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917). 2901\u20132910."},{"key":"e_1_3_3_142_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1147"},{"key":"e_1_3_3_143_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-industry.34"},{"key":"e_1_3_3_144_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Kaushik Divyansh","year":"2019","unstructured":"Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2019. Learning the difference that makes a difference with counterfactually-augmented data. In Proceedings of the International Conference on Learning Representations(ICLR\u201919). https:\/\/openreview.net\/forum?id=Sklgs0NFvr."},{"key":"e_1_3_3_145_2","first-page":"5481","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201920)","author":"Keraron Rachel","year":"2020","unstructured":"Rachel Keraron, Guillaume Lancrenon, Mathilde Bras, Fr\u00e9d\u00e9ric Allary, Gilles Moyse, Thomas Scialom, Edmundo-Pavel Soriano-Morales, and Jacopo Staiano. 2020. Project PIAF: Building a native French question-answering dataset. In Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201920). 5481\u20135490. https:\/\/www.aclweb.org\/anthology\/2020.lrec-1.673."},{"key":"e_1_3_3_146_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1023"},{"key":"e_1_3_3_147_2","article-title":"UnifiedQA: Crossing format boundaries with a single QA system","author":"Khashabi Daniel","year":"2020","unstructured":"Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. UnifiedQA: Crossing format boundaries with a single QA system. arXiv:2005.00700 [CS] (2020). https:\/\/arxiv.org\/abs\/2005.00700.","journal-title":"arXiv:2005.00700 [CS]"},{"key":"e_1_3_3_148_2","volume-title":"Proceedings of the 2017 International Joint Conference on Artificial Intelligence (IJCAI\u201917)","author":"Kim Kyung-Min","year":"2017","unstructured":"Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang. 2017. DeepStory: Video story QA by deep embedded memory networks. In Proceedings of the 2017 International Joint Conference on Artificial Intelligence (IJCAI\u201917). https:\/\/openreview.net\/forum?id=ryZczSz_bS."},{"key":"e_1_3_3_149_2","unstructured":"Seokhwan Kim Luis Ferdinando D\u2019Haro Rafael E. Banchs Matthew Henderson Jason Willisams and Koichiro Yoshino. 2016. Dialog State Tracking Challenge 5 Handbook v.3.1. Retrieved September 16 2022 from http:\/\/workshop.colips.org\/dstc5\/."},{"key":"e_1_3_3_150_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00023"},{"key":"e_1_3_3_151_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.502"},{"key":"e_1_3_3_152_2","volume-title":"The Stanford Encyclopedia of Philosophy (Winter 2017 ed.)","author":"Koons Robert","year":"2017","unstructured":"Robert Koons. 2017. Defeasible reasoning. In The Stanford Encyclopedia of Philosophy (Winter 2017 ed.), Edward N. Zalta (Ed.). Metaphysics Research Lab, Stanford University, Stanford, CA. https:\/\/plato.stanford.edu\/archives\/win2017\/entries\/reasoning-defeasible\/."},{"key":"e_1_3_3_153_2","article-title":"RuBQ: A Russian dataset for question answering over wikidata","author":"Korablinov Vladislav","year":"2020","unstructured":"Vladislav Korablinov and Pavel Braslavski. 2020. RuBQ: A Russian dataset for question answering over wikidata. arXiv:2005.10659 [CS] (2020). http:\/\/arxiv.org\/abs\/2005.10659.","journal-title":"arXiv:2005.10659 [CS]"},{"key":"e_1_3_3_154_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.393"},{"key":"e_1_3_3_155_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1026"},{"key":"e_1_3_3_156_2","doi-asserted-by":"crossref","DOI":"10.1162\/tacl_a_00276","article-title":"Natural questions: A benchmark for question answering research","author":"Kwiatkowski Tom","year":"2019","unstructured":"Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, et\u00a0al. 2019. Natural questions: A benchmark for question answering research. Transactions of Association for Computational Linguistics 7 (2019), 452\u2013466. https:\/\/ai.google\/research\/pubs\/pub47761.","journal-title":"Transactions of Association for Computational Linguistics"},{"key":"e_1_3_3_157_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1082"},{"key":"e_1_3_3_158_2","doi-asserted-by":"publisher","DOI":"10.1109\/SLT.2018.8639505"},{"key":"e_1_3_3_159_2","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201918)","author":"Lee Kyungjae","year":"2018","unstructured":"Kyungjae Lee, Kyoungho Yoon, Sunghyun Park, and Seung-Won Hwang. 2018. Semi-supervised training data generation for multilingual question answering. In Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201918). https:\/\/www.aclweb.org\/anthology\/L18-1437."},{"key":"e_1_3_3_160_2","first-page":"1369","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918)","author":"Lei Jie","year":"2018","unstructured":"Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg. 2018. TVQA: Localized, compositional video question answering. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918). 1369\u20131379. https:\/\/www.aclweb.org\/anthology\/papers\/D\/D18\/D18-1167\/."},{"key":"e_1_3_3_161_2","first-page":"552","volume-title":"Proceedings of the 13th International Conference on Principles of Knowledge Representation and Reasoning","author":"Levesque Hector J.","year":"2012","unstructured":"Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012. The Winograd Schema Challenge. In Proceedings of the 13th International Conference on Principles of Knowledge Representation and Reasoning. 552\u2013561."},{"key":"e_1_3_3_162_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/K17-1034"},{"key":"e_1_3_3_163_2","doi-asserted-by":"crossref","first-page":"7315","DOI":"10.18653\/v1\/2020.acl-main.653","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920)","author":"Lewis Patrick","year":"2020","unstructured":"Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. MLQA: Evaluating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920). 7315\u20137330. https:\/\/www.aclweb.org\/anthology\/2020.acl-main.653\/."},{"key":"e_1_3_3_164_2","article-title":"Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension","author":"Li Chia-Hsuan","year":"2018","unstructured":"Chia-Hsuan Li, Szu-Lin Wu, Chi-Liang Liu, and Hung-Yi Lee. 2018. Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension. arXiv:1804.00320 [CS] (2018). http:\/\/arxiv.org\/abs\/1804.00320.","journal-title":"arXiv:1804.00320 [CS]"},{"key":"e_1_3_3_165_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.90"},{"key":"e_1_3_3_166_2","article-title":"Molweni: A challenge multiparty dialogues-based machine reading comprehension dataset with discourse structure","author":"Li Jiaqi","year":"2020","unstructured":"Jiaqi Li, Ming Liu, Min-Yen Kan, Zihao Zheng, Zekun Wang, Wenqiang Lei, Ting Liu, and Bing Qin. 2020. Molweni: A challenge multiparty dialogues-based machine reading comprehension dataset with discourse structure. arXiv:2004.05080 [CS] (2020). http:\/\/arxiv.org\/abs\/2004.05080.","journal-title":"arXiv:2004.05080 [CS]"},{"key":"e_1_3_3_167_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.698"},{"key":"e_1_3_3_168_2","article-title":"Dataset and neural recurrent sequence labeling model for open-domain factoid question answering","author":"Li Peng","year":"2016","unstructured":"Peng Li, Wei Li, Zhengyan He, Xuguang Wang, Ying Cao, Jie Zhou, and Wei Xu. 2016. Dataset and neural recurrent sequence labeling model for open-domain factoid question answering. arXiv:1607.06275 [CS] (2016). http:\/\/arxiv.org\/abs\/1607.06275.","journal-title":"arXiv:1607.06275 [CS]"},{"key":"e_1_3_3_169_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1129"},{"key":"e_1_3_3_170_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.290"},{"key":"e_1_3_3_171_2","first-page":"742","volume-title":"Proceedings of the 11th Asian Conference on Machine Learning.","author":"Liang Yichan","year":"2019","unstructured":"Yichan Liang, Jianheng Li, and Jian Yin. 2019. A new multi-choice reading comprehension dataset for curriculum learning. In Proceedings of the 11th Asian Conference on Machine Learning. 742\u2013757. http:\/\/proceedings.mlr.press\/v101\/liang19a.html."},{"key":"e_1_3_3_172_2","article-title":"KorQuAD1.0: Korean QA dataset for machine reading comprehension","author":"Lim Seungyoung","year":"2019","unstructured":"Seungyoung Lim, Myungji Kim, and Jooyoul Lee. 2019. KorQuAD1.0: Korean QA dataset for machine reading comprehension. arXiv:1909.07005 [CS] (2019). http:\/\/arxiv.org\/abs\/1909.07005.","journal-title":"arXiv:1909.07005 [CS]"},{"key":"e_1_3_3_173_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.557"},{"key":"e_1_3_3_174_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-5808"},{"key":"e_1_3_3_175_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.229"},{"key":"e_1_3_3_176_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_3_177_2","article-title":"How can we accelerate progress towards human-like linguistic generalization?","author":"Linzen Tal","year":"2020","unstructured":"Tal Linzen. 2020. How can we accelerate progress towards human-like linguistic generalization? arXiv:2005.00955 [CS] (2020). https:\/\/arxiv.org\/pdf\/2005.00955.pdf.","journal-title":"arXiv:2005.00955 [CS]"},{"key":"e_1_3_3_178_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/501"},{"key":"e_1_3_3_179_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1227"},{"key":"e_1_3_3_180_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-32233-5_43"},{"key":"e_1_3_3_181_2","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv:1907.11692 [CS] (2019). http:\/\/arxiv.org\/abs\/1907.11692.","journal-title":"arXiv:1907.11692 [CS]"},{"key":"e_1_3_3_182_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1086"},{"key":"e_1_3_3_183_2","article-title":"MKQA: A linguistically diverse benchmark for multilingual open domain question answering","author":"Longpre Shayne","year":"2020","unstructured":"Shayne Longpre, Yi Lu, and Joachim Daiber. 2020. MKQA: A linguistically diverse benchmark for multilingual open domain question answering. arXiv:2007.15207 [CS] (2020). http:\/\/arxiv.org\/abs\/2007.15207.","journal-title":"arXiv:2007.15207 [CS]"},{"key":"e_1_3_3_184_2","article-title":"The Ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems","author":"Lowe Ryan","year":"2015","unstructured":"Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau. 2015. The Ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems. arXiv:1506.08909 [CS] (2015). http:\/\/arxiv.org\/abs\/1506.08909.","journal-title":"arXiv:1506.08909 [CS]"},{"key":"e_1_3_3_185_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1185"},{"key":"e_1_3_3_186_2","first-page":"61","article-title":"Multiple-choice tests can support deep learning!","author":"MacFarlane Leigh-Ann","year":"2017","unstructured":"Leigh-Ann MacFarlane and Genvi\u00e8ve Boulet. 2017. Multiple-choice tests can support deep learning! Proceedings of the Atlantic Universities\u2019 Teaching Showcase 21 (2017), 61\u201366. https:\/\/ojs.library.dal.ca\/auts\/article\/view\/8430.","journal-title":"Proceedings of the Atlantic Universities\u2019 Teaching Showcase"},{"key":"e_1_3_3_187_2","article-title":"A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering","author":"Maharaj Tegan","year":"2017","unstructured":"Tegan Maharaj, Nicolas Ballas, Anna Rohrbach, Aaron Courville, and Christopher Pal. 2017. A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering. arXiv:1611.07810 [CS] (2017). http:\/\/arxiv.org\/abs\/1611.07810.","journal-title":"arXiv:1611.07810 [CS]"},{"issue":"5","key":"e_1_3_3_188_2","first-page":"44","article-title":"Developing certification exam questions: More deliberate than you may think","volume":"63","author":"Marcham Cheryl L.","year":"2018","unstructured":"Cheryl L. Marcham, Treasa M. Turnbeaugh, Susan Gould, and Joel T. Nadler. 2018. Developing certification exam questions: More deliberate than you may think. Professional Safety 63, 5 (May 2018), 44\u201349. https:\/\/onepetro.org\/PS\/article\/63\/05\/44\/33528\/Developing-Certification-Exam-Questions-More.","journal-title":"Professional Safety"},{"key":"e_1_3_3_189_2","volume-title":"Proceedings of the 6th Language and Technology Conference: Human Language Technologies as a Challenge for Computer Science and Linguistics","author":"Marcinczuk Micha\u0142","year":"2013","unstructured":"Micha\u0142 Marcinczuk, Marcin Ptak, Adam Radziszewski, and Maciej Piasecki. 2013. Open dataset for development of Polish question answering systems. In Proceedings of the 6th Language and Technology Conference: Human Language Technologies as a Challenge for Computer Science and Linguistics. https:\/\/www.researchgate.net\/profile\/Maciej-Piasecki\/publication\/272685856_Open_dataset_for_development_of_Polish_Question_Answering_systems."},{"key":"e_1_3_3_190_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-acl.177"},{"key":"e_1_3_3_191_2","doi-asserted-by":"publisher","DOI":"10.1145\/2872427.2883044"},{"key":"e_1_3_3_192_2","article-title":"The natural language decathlon: Multitask learning as question answering","author":"McCann Bryan","year":"2018","unstructured":"Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018. The natural language decathlon: Multitask learning as question answering. arXiv:1806.08730 [CS, STAT] (2018). http:\/\/arxiv.org\/abs\/1806.08730.","journal-title":"arXiv:1806.08730 [CS, STAT]"},{"key":"e_1_3_3_193_2","first-page":"463","volume-title":"Machine Intelligence 4","author":"McCarthy John","year":"1969","unstructured":"John McCarthy and Patrick Hayes. 1969. Some philosophical problems from the standpoint of artificial intelligence. In Machine Intelligence 4, B. Meltzer and Donald Michie (Eds.). Edinburgh University Press, 463\u2013502."},{"key":"e_1_3_3_194_2","article-title":"BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance","author":"McCoy Tom","year":"2019","unstructured":"Tom McCoy, Junghyun Min, and Tal Linzen. 2019. BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance. arXiv:1911.02969 [CS] (2019). http:\/\/arxiv.org\/abs\/1911.02969.","journal-title":"arXiv:1911.02969 [CS]"},{"key":"e_1_3_3_195_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1334"},{"key":"e_1_3_3_196_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0079-7421(09)51009-2"},{"key":"e_1_3_3_197_2","doi-asserted-by":"crossref","first-page":"975","DOI":"10.18653\/v1\/2020.acl-main.92","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920)","author":"Miao Shen-Yun","year":"2020","unstructured":"Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2020. A diverse corpus for evaluating and developing English math word problem solvers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920). 975\u2013984. https:\/\/www.aclweb.org\/anthology\/2020.acl-main.92."},{"key":"e_1_3_3_198_2","doi-asserted-by":"crossref","first-page":"2381","DOI":"10.18653\/v1\/D18-1260","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918)","author":"Mihaylov Todor","year":"2018","unstructured":"Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? A new dataset for open book question answering. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918). 2381\u20132391. http:\/\/aclweb.org\/anthology\/D18-1260."},{"key":"e_1_3_3_199_2","article-title":"AmbigQA: Answering ambiguous open-domain questions","author":"Min Sewon","year":"2020","unstructured":"Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. AmbigQA: Answering ambiguous open-domain questions. arXiv:2004.10645 [CS] (2020). http:\/\/arxiv.org\/abs\/2004.10645.","journal-title":"arXiv:2004.10645 [CS]"},{"key":"e_1_3_3_200_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.364"},{"key":"e_1_3_3_201_2","article-title":"Towards question format independent numerical reasoning: A set of prerequisite tasks","author":"Mishra Swaroop","year":"2020","unstructured":"Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Sachdeva, and Chitta Baral. 2020. Towards question format independent numerical reasoning: A set of prerequisite tasks. arXiv preprint arXiv:2005.08516 (2020). https:\/\/arxiv.org\/abs\/2005.08516.","journal-title":"arXiv preprint arXiv:2005.08516"},{"key":"e_1_3_3_202_2","doi-asserted-by":"publisher","DOI":"10.1145\/3287560.3287596"},{"key":"e_1_3_3_203_2","unstructured":"Ashutosh Modi Tatjana Anikina Simon Ostermann and Manfred Pinkal. 2016. InScript: Narrative texts annotated with script information. In Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201916) . 3485\u20133493. https:\/\/www.aclweb.org\/anthology\/L16-1555."},{"key":"e_1_3_3_204_2","volume-title":"Proceedings of the 1st Workshop on NLP for COVID-19 at ACL\u201920","author":"M\u00f6ller Timo","year":"2020","unstructured":"Timo M\u00f6ller, Anthony Reina, Raghavan Jayakumar, and Malte Pietsch. 2020. COVID-QA: A question answering dataset for COVID-19. In Proceedings of the 1st Workshop on NLP for COVID-19 at ACL\u201920. https:\/\/www.aclweb.org\/anthology\/2020.nlpcovid19-acl.18."},{"key":"e_1_3_3_205_2","doi-asserted-by":"crossref","first-page":"46","DOI":"10.18653\/v1\/W17-0906","volume-title":"Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential, and Discourse-Level Semantics","author":"Mostafazadeh Nasrin","year":"2017","unstructured":"Nasrin Mostafazadeh, Michael Roth, Nathanael Chambers, and Annie Louis. 2017. LSDSem 2017 shared task: The story Cloze test. In Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential, and Discourse-Level Semantics. 46\u201351. http:\/\/www.aclweb.org\/anthology\/W17-0900."},{"key":"e_1_3_3_206_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-4612"},{"key":"e_1_3_3_207_2","volume-title":"Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV\u201917)","author":"Mun Jonghwan","year":"2017","unstructured":"Jonghwan Mun, Paul Hongsuck Seo, Ilchae Jung, and Bohyung Han. 2017. MarioQA: Answering questions by watching gameplay videos. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV\u201917). http:\/\/arxiv.org\/abs\/1612.01669."},{"key":"e_1_3_3_208_2","doi-asserted-by":"crossref","first-page":"27","DOI":"10.18653\/v1\/S17-2003","volume-title":"Proceedings of the 11th International Workshop on Semantic Evaluations (SemEval\u201917)","author":"Nakov Preslav","year":"2017","unstructured":"Preslav Nakov, Doris Hoogeveen, Llu\u00eds M\u00e0rquez, Alessandro Moschitti, Hamdy Mubarak, Timothy Baldwin, and Karin Verspoor. 2017. SemEval-2017 Task 3: Community question answering. In Proceedings of the 11th International Workshop on Semantic Evaluations (SemEval\u201917). 27\u201348. http:\/\/www.aclweb.org\/anthology\/S17-2003."},{"key":"e_1_3_3_209_2","doi-asserted-by":"crossref","unstructured":"Preslav Nakov Llu\u00eds M\u00e0rquez Walid Magdy Alessandro Moschitti Jim Glass and Bilal Randeree. 2015. SemEval-2015 Task 3: Answer selection in community question answering. In Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval\u201915) . 269\u2013281.","DOI":"10.18653\/v1\/S15-2047"},{"key":"e_1_3_3_210_2","doi-asserted-by":"crossref","unstructured":"Preslav Nakov Llu\u00eds M\u00e0rquez Alessandro Moschitti Walid Magdy Hamdy Mubarak abed Alhakim Freihat Jim Glass and Bilal Randeree. 2016. SemEval-2016 Task 3: Community question answering. 525\u2013545.","DOI":"10.18653\/v1\/S16-1083"},{"key":"e_1_3_3_211_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.233"},{"key":"e_1_3_3_212_2","article-title":"TORQUE: A reading comprehension dataset of temporal ordering questions","author":"Ning Qiang","year":"2020","unstructured":"Qiang Ning, Hao Wu, Rujun Han, Nanyun Peng, Matt Gardner, and Dan Roth. 2020. TORQUE: A reading comprehension dataset of temporal ordering questions. arXiv:2005.00242 [CS] (2020). http:\/\/arxiv.org\/abs\/2005.00242.","journal-title":"arXiv:2005.00242 [CS]"},{"key":"e_1_3_3_213_2","first-page":"2450","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920)","author":"Omura Kazumasa","year":"2020","unstructured":"Kazumasa Omura, Daisuke Kawahara, and Sadao Kurohashi. 2020. A method for building a commonsense inference dataset based on basic events. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920). 2450\u20132460. https:\/\/www.aclweb.org\/anthology\/2020.emnlp-main.192."},{"key":"e_1_3_3_214_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1241"},{"key":"e_1_3_3_215_2","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201918)","author":"Ostermann Simon","year":"2018","unstructured":"Simon Ostermann, Ashutosh Modi, Michael Roth, Stefan Thater, and Manfred Pinkal. 2018. MCScript: A novel dataset for assessing machine comprehension using script knowledge. In Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201918). https:\/\/www.aclweb.org\/anthology\/L18-1564."},{"key":"e_1_3_3_216_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S18-1119"},{"key":"e_1_3_3_217_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1258"},{"key":"e_1_3_3_218_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.391"},{"key":"e_1_3_3_219_2","article-title":"The LAMBADA dataset: Word prediction requiring a broad discourse context","author":"Paperno Denis","year":"2016","unstructured":"Denis Paperno, Germ\u00e1n Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fern\u00e1ndez. 2016. The LAMBADA dataset: Word prediction requiring a broad discourse context. arXiv:1606.06031 [CS] (2016). http:\/\/arxiv.org\/abs\/1606.06031.","journal-title":"arXiv:1606.06031 [CS]"},{"key":"e_1_3_3_220_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1142"},{"key":"e_1_3_3_221_2","article-title":"Generating natural questions from images for multimodal assistants","author":"Patel Alkesh","year":"2020","unstructured":"Alkesh Patel, Akanksha Bindal, Hadas Kotek, Christopher Klein, and Jason Williams. 2020. Generating natural questions from images for multimodal assistants. arXiv:2012.03678 [CS] (2020). http:\/\/arxiv.org\/abs\/2012.03678.","journal-title":"arXiv:2012.03678 [CS]"},{"key":"e_1_3_3_222_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-11382-1_23"},{"key":"e_1_3_3_223_2","doi-asserted-by":"crossref","first-page":"539","DOI":"10.1007\/978-3-319-24027-5_50","volume-title":"Experimental IR Meets Multilinguality, Multimodality, and Interaction.","author":"Pe\u00f1as Anselmo","year":"2015","unstructured":"Anselmo Pe\u00f1as, Christina Unger, Georgios Paliouras, and Ioannis Kakadiaris. 2015. Overview of the CLEF question answering track 2015. In Experimental IR Meets Multilinguality, Multimodality, and Interaction.Lecture Notes in Computer Science, Vol. 9283. Springer, 539\u2013544."},{"key":"e_1_3_3_224_2","article-title":"Introducing MANtIS: A novel multi-domain information seeking dialogues dataset","author":"Penha Gustavo","year":"2019","unstructured":"Gustavo Penha, Alexandru Balan, and Claudia Hauff. 2019. Introducing MANtIS: A novel multi-domain information seeking dialogues dataset. arXiv:1912.04639 [CS] (2019). http:\/\/arxiv.org\/abs\/1912.04639.","journal-title":"arXiv:1912.04639 [CS]"},{"key":"e_1_3_3_225_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1460"},{"key":"e_1_3_3_226_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-acl.196"},{"key":"e_1_3_3_227_2","unstructured":"Eric Price. 2014. The NIPS Experiment. Retrieved September 16 2022 from http:\/\/blog.mrtz.org\/2014\/12\/15\/the-nips-experiment.html."},{"key":"e_1_3_3_228_2","article-title":"Learning to deceive with attention-based explanations","author":"Pruthi Danish","year":"2019","unstructured":"Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C. Lipton. 2019. Learning to deceive with attention-based explanations. arXiv:1909.07913 [CS] (2019). http:\/\/arxiv.org\/abs\/1909.07913.","journal-title":"arXiv:1909.07913 [CS]"},{"key":"e_1_3_3_229_2","unstructured":"Lianhui Qin Aditya Gupta Shyam Upadhyay Luheng He Yejin Choi and Manaal Faruqui. 2021. TIMEDIAL: Temporal commonsense reasoning in dialog. arXiv:2106.04571 [CS.CL] (2021)."},{"key":"e_1_3_3_230_2","article-title":"A survey on neural machine reading comprehension","author":"Qiu Boyu","year":"2019","unstructured":"Boyu Qiu, Xu Chen, Jungang Xu, and Yingfei Sun. 2019. A survey on neural machine reading comprehension. arXiv:1906.03824 [CS] (2019). http:\/\/arxiv.org\/abs\/1906.03824.","journal-title":"arXiv:1906.03824 [CS]"},{"key":"e_1_3_3_231_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210124"},{"key":"e_1_3_3_232_2","doi-asserted-by":"crossref","unstructured":"Filip Radlinski Krisztian Balog Bill Byrne and Karthik Krishnamoorthi. 2019. Coached conversational preference elicitation: A case study in understanding movie preferences. In Proceedings of the 20th Annual SIGdial Meeting on Discourse and Dialogue . https:\/\/research.google\/pubs\/pub48414\/.","DOI":"10.18653\/v1\/W19-5941"},{"key":"e_1_3_3_233_2","doi-asserted-by":"publisher","DOI":"10.1145\/2740908.2743006"},{"key":"e_1_3_3_234_2","article-title":"Explaining and improving model behavior with k nearest neighbor representations","author":"Rajani Nazneen Fatema","year":"2020","unstructured":"Nazneen Fatema Rajani, Ben Krause, Wengpeng Yin, Tong Niu, Richard Socher, and Caiming Xiong. 2020. Explaining and improving model behavior with k nearest neighbor representations. arXiv:2010.09030 [CS] (2020). http:\/\/arxiv.org\/abs\/2010.09030.","journal-title":"arXiv:2010.09030 [CS]"},{"key":"e_1_3_3_235_2","first-page":"784","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL\u201918)","author":"Rajpurkar Pranav","year":"2018","unstructured":"Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don\u2019t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL\u201918). 784\u2013789. http:\/\/aclweb.org\/anthology\/P18-2124."},{"key":"e_1_3_3_236_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_3_3_237_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.603"},{"key":"e_1_3_3_238_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1043"},{"key":"e_1_3_3_239_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00266"},{"key":"e_1_3_3_240_2","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939778"},{"key":"e_1_3_3_241_2","doi-asserted-by":"crossref","first-page":"4902","DOI":"10.18653\/v1\/2020.acl-main.442","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920)","author":"Ribeiro Marco Tulio","year":"2020","unstructured":"Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with checklist. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL\u201920). 4902\u20134912. https:\/\/www.aclweb.org\/anthology\/2020.acl-main.442."},{"key":"e_1_3_3_242_2","first-page":"193","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201913)","author":"Richardson Matthew","year":"2013","unstructured":"Matthew Richardson, Christopher J. C. Burges, and Erin Renshaw. 2013. MCTest: A challenge dataset for the open-domain machine comprehension of text. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201913). 193\u2013203."},{"key":"e_1_3_3_243_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.758"},{"key":"e_1_3_3_244_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.655"},{"key":"e_1_3_3_245_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1904.04792"},{"key":"e_1_3_3_246_2","first-page":"6","volume-title":"Proceedings of the AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning","author":"Roemmele Melissa","year":"2011","unstructured":"Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S. Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In Proceedings of the AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning. 6."},{"key":"e_1_3_3_247_2","unstructured":"Anna Rogers. 2019. How the Transformers Broke NLP Leaderboards. Retrieved September 16 2022 fromhttps:\/\/hackingsemantics.xyz\/2019\/leaderboards\/."},{"key":"e_1_3_3_248_2","first-page":"2182","volume-title":"Proceedings of the Conference of the Association for Computational Linguistics (ACL\u201921)","author":"Rogers Anna","year":"2021","unstructured":"Anna Rogers. 2021. Changing the world by changing the data. In Proceedings of the Conference of the Association for Computational Linguistics (ACL\u201921). 2182\u20132194. https:\/\/aclanthology.org\/2021.acl-long.170."},{"key":"e_1_3_3_249_2","first-page":"1256","volume-title":"Findings of EMNLP\u201920","author":"Rogers Anna","year":"2020","unstructured":"Anna Rogers and Isabelle Augenstein. 2020. What can we do to improve peer review in NLP? In Findings of EMNLP\u201920. 1256\u20131262. https:\/\/www.aclweb.org\/anthology\/2020.findings-emnlp.112\/."},{"key":"e_1_3_3_250_2","first-page":"8722","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI\u201920)","author":"Rogers Anna","year":"2020","unstructured":"Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky. 2020. Getting closer to AI complete question answering: A set of prerequisite real tasks. In Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI\u201920). 8722\u20138731. https:\/\/aaai.org\/ojs\/index.php\/AAAI\/article\/view\/6398."},{"key":"e_1_3_3_251_2","article-title":"LAReQA: Language-agnostic answer retrieval from a multilingual pool","author":"Roy Uma","year":"2020","unstructured":"Uma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua, Aaron Phillips, and Yinfei Yang. 2020. LAReQA: Language-agnostic answer retrieval from a multilingual pool. arXiv:2004.05484 [CS] (2020). http:\/\/arxiv.org\/abs\/2004.05484.","journal-title":"arXiv:2004.05484 [CS]"},{"key":"e_1_3_3_252_2","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201921)","author":"Ruder Sebastian","year":"2021","unstructured":"Sebastian Ruder and Si Avirup. 2021. Multi-domain multilingual question answering. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201921)."},{"key":"e_1_3_3_253_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.418"},{"key":"e_1_3_3_254_2","first-page":"322","volume-title":"Proceedings of the 2018 EMNLP Workshop BlackboxNLP","author":"Rychalska Barbara","year":"2018","unstructured":"Barbara Rychalska, Dominika Basaj, Anna Wr\u00f3blewska, and Przemyslaw Biecek. 2018. Does it care what you asked? Understanding importance of verbs in deep learning QA system. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP. 322\u2013324. http:\/\/aclweb.org\/anthology\/W18-5436."},{"key":"e_1_3_3_255_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1024"},{"key":"e_1_3_3_256_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1233"},{"key":"e_1_3_3_257_2","article-title":"WinoGrande: An adversarial Winograd Schema Challenge at scale","author":"Sakaguchi Keisuke","year":"2019","unstructured":"Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. WinoGrande: An adversarial Winograd Schema Challenge at scale. arXiv:1907.10641 [CS] (2019). http:\/\/arxiv.org\/abs\/1907.10641.","journal-title":"arXiv:1907.10641 [CS]"},{"key":"e_1_3_3_258_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1454"},{"key":"e_1_3_3_259_2","article-title":"Beyond leaderboards: A survey of methods for revealing weaknesses in natural language inference data and models","author":"Schlegel Viktor","year":"2020","unstructured":"Viktor Schlegel, Goran Nenadic, and Riza Batista-Navarro. 2020. Beyond leaderboards: A survey of methods for revealing weaknesses in natural language inference data and models. arXiv:2005.14709 [CS] (2020). http:\/\/arxiv.org\/abs\/2005.14709.","journal-title":"arXiv:2005.14709 [CS]"},{"key":"e_1_3_3_260_2","volume-title":"Proceedings of the Language Resources and Evaluation Conference","author":"Schlegel Viktor","year":"2020","unstructured":"Viktor Schlegel, Marco Valentino, Andr\u00e9 Freitas, Goran Nenadic, and Riza Batista-Navarro. 2020. A framework for evaluation of machine reading comprehension gold standards. In Proceedings of the Language Resources and Evaluation Conference. http:\/\/arxiv.org\/abs\/2003.04642."},{"key":"e_1_3_3_261_2","article-title":"A survey of available corpora for building data-driven dialogue systems","author":"Serban Iulian Vlad","year":"2015","unstructured":"Iulian Vlad Serban, Ryan Lowe, Peter Henderson, Laurent Charlin, and Joelle Pineau. 2015. A survey of available corpora for building data-driven dialogue systems. arXiv:1512.05742 [CS, STAT] (2015). http:\/\/arxiv.org\/abs\/1512.05742.","journal-title":"arXiv:1512.05742 [CS, STAT]"},{"key":"e_1_3_3_262_2","article-title":"DRCD: A Chinese machine reading comprehension dataset","author":"Shao Chih Chieh","year":"2019","unstructured":"Chih Chieh Shao, Trois Liu, Yuting Lai, Yiying Tseng, and Sam Tsai. 2019. DRCD: A Chinese machine reading comprehension dataset. arXiv:1806.00920 [CS] (2019). http:\/\/arxiv.org\/abs\/1806.00920.","journal-title":"arXiv:1806.00920 [CS]"},{"key":"e_1_3_3_263_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1135"},{"key":"e_1_3_3_264_2","first-page":"518","article-title":"Overview of the NTCIR-11 QA-lab task","author":"Shibuki Hideyuki","year":"2014","unstructured":"Hideyuki Shibuki, Kotaro Sakamoto, Yoshionobu Kano, Teruko Mitamura, Madoka Ishioroshi, Kelly Y. Itakura, Di Wang, Tatsunori Mori, and Noriko Kando. 2014. Overview of the NTCIR-11 QA-lab task. In Proceedings of the 11th NTCIR Conference. 518\u2013529. http:\/\/research.nii.ac.jp\/ntcir\/workshop\/OnlineProceedings11\/pdf\/NTCIR\/OVERVIEW\/01-NTCIR11-OV-QALAB-ShibukiH.pdf.","journal-title":"Proceedings of the 11th NTCIR Conference."},{"key":"e_1_3_3_265_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1458"},{"key":"e_1_3_3_266_2","article-title":"A survey of code-switched speech and language processing","author":"Sitaram Sunayana","year":"2020","unstructured":"Sunayana Sitaram, Khyathi Raghavi Chandu, Sai Krishna Rallabandi, and Alan W. Black. 2020. A survey of code-switched speech and language processing. arXiv:1904.00784 [CS, STAT] (2020). http:\/\/arxiv.org\/abs\/1904.00784.","journal-title":"arXiv:1904.00784 [CS, STAT]"},{"key":"e_1_3_3_267_2","first-page":"1245","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (EACL\u201921)","author":"Soleimani Amir","year":"2021","unstructured":"Amir Soleimani, Christof Monz, and Marcel Worring. 2021. NLQuAD: A non-factoid long question answering data set. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (EACL\u201921). 1245\u20131255. https:\/\/aclanthology.org\/2021.eacl-main.106."},{"key":"e_1_3_3_268_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-6001"},{"key":"e_1_3_3_269_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1075"},{"key":"e_1_3_3_270_2","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI\u201920)","author":"Sugawara Saku","year":"2020","unstructured":"Saku Sugawara, Pontus Stenetorp, Kentaro Inui, and Akiko Aizawa. 2020. Assessing the benchmarking capacity of machine reading comprehension datasets. In Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI\u201920). http:\/\/arxiv.org\/abs\/1911.09241."},{"key":"e_1_3_3_271_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-2034"},{"key":"e_1_3_3_272_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1644"},{"key":"e_1_3_3_273_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.253"},{"key":"e_1_3_3_274_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00264"},{"key":"e_1_3_3_275_2","article-title":"TableQA: A large-scale Chinese Text-to-SQL dataset for table-aware SQL generation","author":"Sun Ningyuan","year":"2020","unstructured":"Ningyuan Sun, Xuefeng Yang, and Yunfeng Liu. 2020. TableQA: A large-scale Chinese Text-to-SQL dataset for table-aware SQL generation. arXiv:2006.06434 [CS] (2020). http:\/\/arxiv.org\/abs\/2006.06434.","journal-title":"arXiv:2006.06434 [CS]"},{"key":"e_1_3_3_276_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1140"},{"key":"e_1_3_3_277_2","volume-title":"Proceedings of the 33rd AAAI Conference on Artificial Intelligence (AAAI\u201919)","author":"Tafjord Oyvind","year":"2019","unstructured":"Oyvind Tafjord, Peter Clark, Matt Gardner, Wen-Tau Yih, and Ashish Sabharwal. 2019. QuaRel: A dataset and models for answering questions about qualitative relationships. In Proceedings of the 33rd AAAI Conference on Artificial Intelligence (AAAI\u201919)."},{"key":"e_1_3_3_278_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1608"},{"key":"e_1_3_3_279_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1059"},{"key":"e_1_3_3_280_2","first-page":"4149","volume-title":"Proceedings of the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL\u201922)","author":"Talmor Alon","year":"2019","unstructured":"Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL\u201922). 4149\u20134158. https:\/\/www.aclweb.org\/anthology\/papers\/N\/N19\/N19-1421\/."},{"key":"e_1_3_3_281_2","first-page":"12","volume-title":"Proceedings of the 9th International Conference on Learning Representations (ICLR\u201921)","author":"Talmor Alon","year":"2021","unstructured":"Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant. 2021. MultimodalQA: Complex question answering over text, tables and images. In Proceedings of the 9th International Conference on Learning Representations (ICLR\u201921). 12. https:\/\/openreview.net\/pdf\/f3dad930cb55abce99a229e35cc131a2db791b66.pdf."},{"key":"e_1_3_3_282_2","volume-title":"Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"Tapaswi Makarand","year":"2016","unstructured":"Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016. MovieQA: Understanding stories in movies through question-answering. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)."},{"key":"e_1_3_3_283_2","volume-title":"Proceedings of the 35th Conference on Neural Information Processing Systems, Datasets, and Benchmarks Track","author":"Thakur Nandan","year":"2021","unstructured":"Nandan Thakur, Nils Reimers, Andreas R\u00fcckl\u00e9, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Proceedings of the 35th Conference on Neural Information Processing Systems, Datasets, and Benchmarks Track. https:\/\/openreview.net\/forum?id=wCu6T5xFjeJ."},{"key":"e_1_3_3_284_2","volume-title":"Proceedings of the 1st International Workshop on Conversational Approaches to Information Retrieval (CAIR\u201917)","author":"Thomas Paul","year":"2017","unstructured":"Paul Thomas, Daniel McDuff, Mary Czerwinski, and Nick Craswell. 2017. MISC: A data set of information-seeking conversations. In Proceedings of the 1st International Workshop on Conversational Approaches to Information Retrieval (CAIR\u201917). https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2017\/07\/Thomas-etal-CAIR17.pdf."},{"key":"e_1_3_3_285_2","first-page":"1977","volume-title":"Proceedings of the 2019 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL\u201919).","author":"Thomason Jesse","year":"2019","unstructured":"Jesse Thomason, Daniel Gordon, and Yonatan Bisk. 2019. Shifting the baseline: Single modality performance on visual navigation & QA. In Proceedings of the 2019 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL\u201919). 1977\u20131983. https:\/\/www.aclweb.org\/anthology\/papers\/N\/N19\/N19-1197\/."},{"key":"e_1_3_3_286_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1074"},{"key":"e_1_3_3_287_2","doi-asserted-by":"publisher","DOI":"10.1145\/3176349.3176387"},{"key":"e_1_3_3_288_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-2623"},{"key":"e_1_3_3_289_2","doi-asserted-by":"publisher","DOI":"10.1186\/s12859-015-0564-6"},{"key":"e_1_3_3_290_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2016-876"},{"key":"e_1_3_3_291_2","first-page":"494","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL\u201917)","author":"Upadhyay Shyam","year":"2017","unstructured":"Shyam Upadhyay and Ming-Wei Chang. 2017. Annotating derivations: A new evaluation strategy and dataset for algebra word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL\u201917). 494\u2013504. https:\/\/www.aclweb.org\/anthology\/E17-1047."},{"key":"e_1_3_3_292_2","article-title":"TableQA: Question answering on tabular data","author":"Vakulenko Svitlana","year":"2017","unstructured":"Svitlana Vakulenko and Vadim Savenkov. 2017. TableQA: Question answering on tabular data. arXiv:1705.06504 [CS] (2017). http:\/\/arxiv.org\/abs\/1705.06504.","journal-title":"arXiv:1705.06504 [CS]"},{"key":"e_1_3_3_293_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-8643"},{"issue":"4","key":"e_1_3_3_294_2","doi-asserted-by":"crossref","first-page":"770","DOI":"10.1037\/0278-7393.28.4.770","article-title":"Temporal order relations in language comprehension","volume":"28","author":"Meer Elke van der","year":"2002","unstructured":"Elke van der Meer, Reinhard Beyer, Bertram Heinze, and Isolde Badel. 2002. Temporal order relations in language comprehension. Journal of Experimental Psychology. Learning, Memory, and Cognition 28, 4 (July 2002), 770\u2013779.","journal-title":"Journal of Experimental Psychology. Learning, Memory, and Cognition"},{"key":"e_1_3_3_295_2","volume-title":"Strategies of Discourse Comprehension","author":"Dijk Teun A. van","year":"1983","unstructured":"Teun A. van Dijk and Walter Kintsch. 1983. Strategies of Discourse Comprehension. Academic Press, New York, NY. P302 .D472 1983"},{"key":"e_1_3_3_296_2","article-title":"HEAD-QA: A healthcare dataset for complex reasoning","author":"Vilares David","year":"2019","unstructured":"David Vilares and Carlos G\u00f3mez-Rodr\u00edguez. 2019. HEAD-QA: A healthcare dataset for complex reasoning. arXiv:1906.04701 [CS] (2019). http:\/\/arxiv.org\/abs\/1906.04701.","journal-title":"arXiv:1906.04701 [CS]"},{"key":"e_1_3_3_297_2","doi-asserted-by":"publisher","DOI":"10.1145\/345508.345577"},{"key":"e_1_3_3_298_2","first-page":"127","volume-title":"Proceedings of the Association for Computational Linguistics Student Research Workshop (ACL-SRW\u201918)","author":"Wallace Eric","year":"2018","unstructured":"Eric Wallace and Jordan Boyd-Graber. 2018. Trick me if you can: Adversarial writing of trivia challenge questions. In Proceedings of the Association for Computational Linguistics Student Research Workshop (ACL-SRW\u201918). 127\u2013133. http:\/\/aclweb.org\/anthology\/P18-3018."},{"key":"e_1_3_3_299_2","article-title":"Universal adversarial triggers for attacking and analyzing NLP","author":"Wallace Eric","year":"2019","unstructured":"Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019. Universal adversarial triggers for attacking and analyzing NLP. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201919)http:\/\/arxiv.org\/abs\/1908.07125.","journal-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201919)"},{"key":"e_1_3_3_300_2","first-page":"8","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI\u201920)","author":"Wang Bingning","year":"2020","unstructured":"Bingning Wang, Ting Yao, Qi Zhang, Jingfang Xu, and Xiaochuan Wang. 2020. ReCO: A large scale Chinese reading comprehension dataset on opinion. In Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI\u201920). 8. https:\/\/www.aaai.org\/Papers\/AAAI\/2020GB\/AAAI-WangB.2547.pdf."},{"key":"e_1_3_3_301_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-020-09387-9"},{"key":"e_1_3_3_302_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2109.03438"},{"key":"e_1_3_3_303_2","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380120"},{"key":"e_1_3_3_304_2","first-page":"6895","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201920)","author":"Watarai Takuto","year":"2020","unstructured":"Takuto Watarai and Masatoshi Tsuchiya. 2020. Developing dataset of Japanese slot filling quizzes designed for evaluation of machine reading comprehension. In Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201920). 6895\u20136901. https:\/\/www.aclweb.org\/anthology\/2020.lrec-1.852."},{"key":"e_1_3_3_305_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-4005"},{"key":"e_1_3_3_306_2","article-title":"Towards AI-complete question answering: A set of prerequisite toy tasks","author":"Weston Jason","year":"2015","unstructured":"Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M. Rush, Bart van Merri\u00ebnboer, Armand Joulin, and Tomas Mikolov. 2015. Towards AI-complete question answering: A set of prerequisite toy tasks. arXiv preprint arXiv:1502.05698 (2015).","journal-title":"arXiv preprint arXiv:1502.05698"},{"key":"e_1_3_3_307_2","unstructured":"Michael White Graham Chapman John Cleese Eric Idle Terry Gilliam Terry Jones Michael Palin et\u00a0al. 2001. Monty Python and the Holy Grail ."},{"key":"e_1_3_3_308_2","first-page":"439","volume-title":"Proceedings of the Human Language Technology Conference of the NAACL, Main Conference","author":"Wong Yuk Wah","year":"2006","unstructured":"Yuk Wah Wong and Raymond Mooney. 2006. Learning for semantic parsing with statistical machine translation. In Proceedings of the Human Language Technology Conference of the NAACL, Main Conference. 439\u2013446. https:\/\/www.aclweb.org\/anthology\/N06-1056."},{"key":"e_1_3_3_309_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.370"},{"key":"e_1_3_3_310_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1496"},{"key":"e_1_3_3_311_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.330"},{"key":"e_1_3_3_312_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.34"},{"key":"e_1_3_3_313_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1237"},{"key":"e_1_3_3_314_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-5923"},{"key":"e_1_3_3_315_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1259"},{"key":"e_1_3_3_316_2","first-page":"2318","volume-title":"Proceedings of the 17th Annual Conference of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201919)","author":"Yatskar Mark","year":"2019","unstructured":"Mark Yatskar. 2019. A qualitative comparison of CoQA, SQuAD 2.0, and QuAC. In Proceedings of the 17th Annual Conference of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201919). 2318\u20132323. https:\/\/www.aclweb.org\/anthology\/papers\/N\/N19\/N19-1241\/."},{"key":"e_1_3_3_317_2","article-title":"On the faithfulness measurements for model interpretations","author":"Yin Fan","year":"2021","unstructured":"Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, and Kai-Wei Chang. 2021. On the faithfulness measurements for model interpretations. arXiv:2104.08782 [CS] (2021). http:\/\/arxiv.org\/abs\/2104.08782.","journal-title":"arXiv:2104.08782 [CS]"},{"key":"e_1_3_3_318_2","article-title":"Towards data distillation for end-to-end spoken conversational question answering","author":"You Chenyu","year":"2020","unstructured":"Chenyu You, Nuo Chen, Fenglin Liu, Dongchao Yang, and Yuexian Zou. 2020. Towards data distillation for end-to-end spoken conversational question answering. arXiv:2010.08923 [CS, EESS] (2020). http:\/\/arxiv.org\/abs\/2010.08923.","journal-title":"arXiv:2010.08923 [CS, EESS]"},{"key":"e_1_3_3_319_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919)","author":"Yu Weihao","year":"2019","unstructured":"Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2019. ReClor: A reading comprehension dataset requiring logical reasoning. In Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919). https:\/\/openreview.net\/forum?id=HJgJtT4tvB."},{"key":"e_1_3_3_320_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1009"},{"key":"e_1_3_3_321_2","volume-title":"Proceedings of the Conference of the Association for Computational Linguistics (ACL\u201919)","author":"Zellers Rowan","year":"2019","unstructured":"Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a machine really finish your sentence? In Proceedings of the Conference of the Association for Computational Linguistics (ACL\u201919). http:\/\/arxiv.org\/abs\/1905.07830."},{"key":"e_1_3_3_322_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.586"},{"key":"e_1_3_3_323_2","article-title":"ReCoRD: Bridging the gap between human and machine commonsense reading comprehension","author":"Zhang Sheng","year":"2018","unstructured":"Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. 2018. ReCoRD: Bridging the gap between human and machine commonsense reading comprehension. arXiv:1810.12885 [cs] (oct2018). arXiv:1810.12885 [cs] http:\/\/arxiv.org\/abs\/1810.12885.","journal-title":"arXiv:1810.12885 [cs]"},{"key":"e_1_3_3_324_2","article-title":"When do you need billions of words of pretraining data?","author":"Zhang Yian","year":"2020","unstructured":"Yian Zhang, Alex Warstadt, Haau-Sing Li, and Samuel R. Bowman. 2020. When do you need billions of words of pretraining data? arXiv:2011.04946 [cs] (nov2020). arXiv:2011.04946 [cs] http:\/\/arxiv.org\/abs\/2011.04946.","journal-title":"arXiv:2011.04946 [cs]"},{"key":"e_1_3_3_325_2","first-page":"449","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL\u201918)","author":"Zhang Zhuosheng","year":"2018","unstructured":"Zhuosheng Zhang and Hai Zhao. 2018. One-shot learning for question-answering in gaokao history challenge. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL\u201918). 449\u2013461. https:\/\/www.aclweb.org\/anthology\/C18-1038."},{"key":"e_1_3_3_326_2","volume-title":"Proceedings of the 38th International Conference on Machine Learning (ICML\u201921)","author":"Zhao Tony Z.","year":"2021","unstructured":"Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning (ICML\u201921). http:\/\/arxiv.org\/abs\/2102.09690."},{"key":"e_1_3_3_327_2","article-title":"Seq2SQL: Generating structured queries from natural language using reinforcement learning","author":"Zhong Victor","year":"2017","unstructured":"Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2SQL: Generating structured queries from natural language using reinforcement learning. arXiv:1709.00103 [CS] (2017). http:\/\/arxiv.org\/abs\/1709.00103.","journal-title":"arXiv:1709.00103 [CS]"},{"key":"e_1_3_3_328_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1332"},{"key":"e_1_3_3_329_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.254"},{"key":"e_1_3_3_330_2","article-title":"Retrieving and reading: A comprehensive survey on open-domain question answering","author":"Zhu Fengbin","year":"2021","unstructured":"Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and Tat-Seng Chua. 2021. Retrieving and reading: A comprehensive survey on open-domain question answering. arXiv:2101.00774 [CS] (2021). http:\/\/arxiv.org\/abs\/2101.00774.","journal-title":"arXiv:2101.00774 [CS]"},{"key":"e_1_3_3_331_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-017-1033-7"},{"key":"e_1_3_3_332_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.342"},{"key":"e_1_3_3_333_2","doi-asserted-by":"publisher","DOI":"10.3758\/s13423-015-0864-x"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3560260","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3560260","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:33Z","timestamp":1750186833000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3560260"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,2]]},"references-count":332,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2023,10,31]]}},"alternative-id":["10.1145\/3560260"],"URL":"https:\/\/doi.org\/10.1145\/3560260","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,2]]},"assertion":[{"value":"2021-07-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-08-11","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-02","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}