{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T04:33:23Z","timestamp":1778214803395,"version":"3.51.4"},"reference-count":44,"publisher":"China Science Publishing & Media Ltd.","issue":"1","license":[{"start":{"date-parts":[[2024,3,11]],"date-time":"2024-03-11T00:00:00Z","timestamp":1710115200000},"content-version":"vor","delay-in-days":70,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,2,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n               <jats:p>Pir\u00e1 is a reading comprehension dataset focused on the ocean, the Brazilian coast, and climate change, built from a collection of scientific abstracts and reports on these topics. This dataset represents a versatile language resource, particularly useful for testing the ability of current machine learning models to acquire expert scientific knowledge. Despite its potential, a detailed set of baselines has not yet been developed for Pir\u00e1. By creating these baselines, researchers can more easily utilize Pir\u00e1 as a resource for testing machine learning models across a wide range of question answering tasks. In this paper, we define six benchmarks over the Pir\u00e1 dataset, covering closed generative question answering, machine reading comprehension, information retrieval, open question answering, answer triggering, and multiple choice question answering. As part of this effort, we have also produced a curated version of the original dataset, where we fixed a number of grammar issues, repetitions, and other shortcomings. Furthermore, the dataset has been extended in several new directions, so as to face the aforementioned benchmarks: translation of supporting texts from English into Portuguese, classification labels for answerability, automatic paraphrases of questions and answers, and multiple choice candidates. The results described in this paper provide several points of reference for researchers interested in exploring the challenges provided by the Pir\u00e1 dataset.<\/jats:p>","DOI":"10.1162\/dint_a_00245","type":"journal-article","created":{"date-parts":[[2024,3,11]],"date-time":"2024-03-11T20:09:25Z","timestamp":1710187765000},"page":"29-63","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":6,"title":["Benchmarks for Pir\u00e1 2.0, a Reading Comprehension Dataset about the Ocean, the Brazilian Coast, and Climate Change"],"prefix":"10.3724","volume":"6","author":[{"given":"Paulo","family":"Pirozelli","sequence":"first","affiliation":[{"name":"Universidade de S\u00e3o Paulo Ringgold Standard Institution-Instituto de Estudos Avan\u00e7ados Av. 370, S\u00e3o Paulo 05508-900, Brazil."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marcos M.","family":"Jos\u00e9","sequence":"additional","affiliation":[{"name":"Universidade de S\u00e3o Paulo Ringgold Standard Institution, Escola Polit\u00e9cnica Av. 370, S\u00e3o Paulo 05508-900, Brazil."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Igor","family":"Silveira","sequence":"additional","affiliation":[{"name":"Universidade de S\u00e3o Paulo Ringgold Standard Institution-Instituto de Matem\u00e1tica e Estat\u00edstica Av. 370, S\u00e3o Paulo 05508-900, Brazil."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fl\u00e1vio","family":"Nakasato","sequence":"additional","affiliation":[{"name":"Universidade de S\u00e3o Paulo Ringgold Standard Institution, Escola Polit\u00e9cnica Av. 370, S\u00e3o Paulo 05508-900, Brazil."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sarajane M.","family":"Peres","sequence":"additional","affiliation":[{"name":"Universidade de S\u00e3o Paulo, Rua Cayowa\u00e1, 876 apto 141, S\u00c3O PAULO 05018-001, Brazil."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anarosa A. F.","family":"Brand\u00e3o","sequence":"additional","affiliation":[{"name":"Universidade de S\u00e3o Paulo Ringgold Standard Institution, Escola Polit\u00e9cnica Av. 370, S\u00e3o Paulo 05508-900, Brazil."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anna H. R.","family":"Costa","sequence":"additional","affiliation":[{"name":"Universidade de S\u00e3o Paulo Ringgold Standard Institution, Escola Polit\u00e9cnica Av. 370, S\u00e3o Paulo 05508-900, Brazil."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fabio G.","family":"Cozman","sequence":"additional","affiliation":[{"name":"Universidade de S\u00e3o Paulo Ringgold Standard Institution, Escola Polit\u00e9cnica Av. 370, S\u00e3o Paulo 05508-900, Brazil."}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"2026","published-online":{"date-parts":[[2024,2,1]]},"reference":[{"key":"2024041719551079100_ref1","first-page":"4544","article-title":"Pir\u00e1: A bilingual Portuguese-English dataset for question-answering about the ocean","volume-title":"Proceedings of the 30th ACM International Conference on Information & Knowledge Management. CIKM \u201821","author":"Paschoal","year":"2021"},{"key":"2024041719551079100_ref2","volume-title":"World Ocean Assessment I","author":"Nations","year":"2017"},{"key":"2024041719551079100_ref3","volume-title":"World Ocean Assessment II","author":"Nations","year":"2021"},{"key":"2024041719551079100_ref4","first-page":"5485","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"2024041719551079100_ref5","volume-title":"PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data.","author":"Carmo","year":"2020"},{"key":"2024041719551079100_ref6","first-page":"483","article-title":"mT5: A massively multilingual pre-trained text-to-text transformer","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, June 6-11, 2021","author":"Xue","year":"2021"},{"key":"2024041719551079100_ref7","article-title":"The brWaC Corpus: A new open resource for Brazilian Portuguese","volume-title":"Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)","author":"Wagner Filho","year":"2018"},{"key":"2024041719551079100_ref8","article-title":"Language models are few-shot learners","volume":"abs\/2005.14165","author":"Brown","year":"2020","journal-title":"CoRR"},{"key":"2024041719551079100_ref9","volume-title":"Gpt-4 technical report","author":"OpenAI","year":"2023"},{"key":"2024041719551079100_ref10","volume-title":"Sparks of artificial general intelligence: Early experiments with GPT-4","author":"Bubeck","year":"2023"},{"key":"2024041719551079100_ref11","first-page":"2383","article-title":"SQuAD:100, 000+questions for machine comprehension of text","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, November 1-4, 2016","author":"Rajpurkar","year":"2016"},{"key":"2024041719551079100_ref12","first-page":"452","article-title":"Natural questions: A benchmark for question answering research","volume-title":"Trans. Assoc. Comput.","author":"Kwiatkowski","year":"2019"},{"key":"2024041719551079100_ref13","first-page":"784","article-title":"Know what you don't know: Unanswerable questions for SQuAD","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, July 15-20, 2018, Volume 2: Short Papers","author":"Rajpurkar","year":"2018"},{"key":"2024041719551079100_ref14","article-title":"CoQA: A conversational question answering challenge","volume":"abs\/1808.07042","author":"Reddy","year":"2018","journal-title":"CoRR"},{"key":"2024041719551079100_ref15","first-page":"2429","volume-title":"What do models learn from question answering datasets? In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Sen","year":"2020"},{"key":"2024041719551079100_ref16","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, June 2-7, 2019","author":"Devlin","year":"2019"},{"key":"2024041719551079100_ref17","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","volume":"abs\/1907.11692","author":"Liu","year":"2019","journal-title":"CoRR"},{"key":"2024041719551079100_ref18","article-title":"GLUE: A multi-task benchmark and analysis platform for natural language understanding","volume-title":"7th International Conference on Learning Representations, ICLR 2019, May 6-9, 2019","author":"Wang","year":"2019"},{"key":"2024041719551079100_ref19","first-page":"785","article-title":"RACE: Large-scale reading comprehension dataset from examinations","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, September 9-11, 2017","author":"Lai","year":"2017"},{"key":"2024041719551079100_ref20","first-page":"403","article-title":"BERTimbau: Pretrained BERT models for Brazilian Portuguese","volume-title":"Intelligent Systems-9th Brazilian Conference, BRACIS 2020, October 20-23, 2020, Proceedings, Part I. Lecture Notes in Computer Science","author":"Souza","year":"2020"},{"issue":"4","key":"2024041719551079100_ref21","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1561\/1500000019","article-title":"The probabilistic relevance framework: BM25 and beyond","volume":"3","author":"Robertson","year":"2009","journal-title":"Foundations and Trends in Information Retrieval"},{"key":"2024041719551079100_ref22","first-page":"6769","article-title":"Dense passage retrieval for open-domain question answering","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, November 16-20, 2020","author":"Karpukhin","year":"2020"},{"key":"2024041719551079100_ref23","first-page":"419","article-title":"DEEPAG\u00c9: Answering questions in Portuguese about the Brazilian environment","volume-title":"Intelligent Systems-10th Brazilian Conference, BRACIS 2021, November 29-December 3, 2021, Proceedings, Part II. Lecture Notes in Computer Science","author":"Ca\u00e7\u00e3o","year":"2021"},{"key":"2024041719551079100_ref24","article-title":"Ethical and social risks of harm from language models","volume":"abs\/2112.04359","author":"Weidinger","year":"2021","journal-title":"CoRR"},{"key":"2024041719551079100_ref25","article-title":"LaMDA: Language models for dialog applications","volume":"abs\/2201.08239","author":"Thoppilan","year":"2022","journal-title":"CoRR"},{"key":"2024041719551079100_ref26","first-page":"2013","article-title":"WikiQA: A challenge dataset for open-domain question answering","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015, September 17-21, 2015","author":"Yang","year":"2015"},{"key":"2024041719551079100_ref27","first-page":"820","article-title":"SelQA: A new benchmark for selection-based question answering","volume-title":"28th IEEE International Conference on Tools with Artificial Intelligence, ICTAI 2016, November 6-8, 2016","author":"Jurczyk","year":"2016"},{"key":"2024041719551079100_ref28","first-page":"9146","article-title":"ReCO: A large scale chinese reading comprehension dataset on opinion","volume-title":"The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, February 7-12, 2020","author":"Wang","year":"2020"},{"key":"2024041719551079100_ref29","first-page":"87228731","article-title":"Getting closer to AI complete question answering: A set of prerequisite real tasks","volume-title":"The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, February 7-12, 2020","author":"Rogers","year":"2020"},{"key":"2024041719551079100_ref30","first-page":"464","article-title":"To answer or not to answer? filtering questions for QA systems","volume-title":"Intelligent Systems - 11th Brazilian Conference, BRACIS 2022, Proceedings, Part II. Lecture Notes in Computer Science","author":"Pirozelli","year":"2022"},{"key":"2024041719551079100_ref31","first-page":"13","article-title":"PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Zhang","year":"2020"},{"key":"2024041719551079100_ref32","first-page":"299","article-title":"PTT5-Paraphraser: Diversity and meaning fidelity in automatic Portuguese paraphrasing","volume-title":"Computational Processing of the Portuguese Language-15th International Conference, PROPOR 2022, March 21-23, 2022, Proceedings. Lecture Notes in Computer Science","author":"Pellicer","year":"2022"},{"key":"2024041719551079100_ref33","first-page":"278","article-title":"Integrating question answering and Text-to-SQL in Portuguese","volume-title":"Computational Processing of the Portuguese Language-15th International Conference, PROPOR 2022, March 21-23, 2022, Proceedings. Lecture Notes in Computer Science","author":"Jos\u00e9","year":"2022"},{"key":"2024041719551079100_ref34","doi-asserted-by":"crossref","first-page":"2580","DOI":"10.1609\/aaai.v30i1.10325","article-title":"Combining retrieval, statistics, and inference to answer elementary science questions","volume":"30","author":"Clark","year":"2016","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2024041719551079100_ref35","article-title":"Think you have solved question answering? Try ARC, the AI2 Reasoning Challenge","volume":"abs\/1803.05457","author":"Clark","year":"2018","journal-title":"CoRR"},{"key":"2024041719551079100_ref36","first-page":"785","article-title":"RACE: Large-scale reading comprehension dataset from examinations","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, September 9-11, 2017","author":"Lai","year":"2017"},{"key":"2024041719551079100_ref37","first-page":"248","article-title":"MedMCQA: A largescale multi-subject multi-choice dataset for medical domain question answering","volume-title":"Conference on Health, Inference, and Learning, CHIL 2022, 7-8 April 2022. Proceedings of Machine Learning Research","author":"Pal","year":"2022"},{"key":"2024041719551079100_ref38","first-page":"18","article-title":"MCTest: A challenge dataset for the open-domain machine comprehension of text","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing","author":"Richardson","year":"2013"},{"key":"2024041719551079100_ref39","first-page":"747","article-title":"SemEval-2018 Task 11: Machine comprehension using commonsense knowledge","volume-title":"Proceedings of the 12th International Workshop on Semantic Evaluation, SemEval@ NAACL-HLT 2018, 5-6 June, 2018","author":"Ostermann","year":"2018"},{"key":"2024041719551079100_ref40","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1558\/cj.v14i2-4.15-33","article-title":"A preliminary inquiry into using corpus word frequency data in the automatic generation of English language cloze tests","volume":"14","author":"Coniam","year":"1997","journal-title":"CALICO Journal"},{"issue":"1","key":"2024041719551079100_ref41","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1109\/TLT.2018.2889100","article-title":"Automatic multiple choice question generation from text: A survey","volume":"13","author":"CH","year":"2020","journal-title":"IEEE Transactions on Learning Technologies"},{"key":"2024041719551079100_ref42","doi-asserted-by":"crossref","first-page":"1358","DOI":"10.18653\/v1\/D18-1166","article-title":"RecipeQA: A challenge dataset for multimodal comprehension of cooking recipes","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, October 31-November 4, 2018","author":"Yagcioglu","year":"2018"},{"key":"2024041719551079100_ref43","first-page":"1896","article-title":"UnifiedQA: Crossing format boundaries with a single QA system","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020, 16-20 November 2020. Findings of ACL, EMNLP 2020","author":"Khashabi","year":"2020"},{"key":"2024041719551079100_ref44","first-page":"426","article-title":"University Entrance Exam as a Guiding Test for Artificial Intelligence","volume-title":"2017 Brazilian Conference on Intelligent Systems, BRACIS 2017, October 2-5, 2017","author":"Silveira","year":"2017"}],"container-title":["Data Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/dint\/article-pdf\/6\/1\/29\/2364147\/dint_a_00245.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/dint\/article-pdf\/6\/1\/29\/2364147\/dint_a_00245.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,3,14]],"date-time":"2025-03-14T07:43:31Z","timestamp":1741938211000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.sciengine.com\/doi\/10.1162\/dint_a_00245"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":44,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,2,1]]}},"URL":"https:\/\/doi.org\/10.1162\/dint_a_00245","relation":{},"ISSN":["2641-435X"],"issn-type":[{"value":"2641-435X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}