{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T19:19:11Z","timestamp":1776885551531,"version":"3.51.2"},"reference-count":49,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2023,6,22]],"date-time":"2023-06-22T00:00:00Z","timestamp":1687392000000},"content-version":"vor","delay-in-days":172,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,6,20]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>We introduce ART, a new corpus-level autoencoding approach for training dense retrieval models that does not require any labeled training data. Dense retrieval is a central challenge for open-domain tasks, such as Open QA, where state-of-the-art methods typically require large supervised datasets with custom hard-negative mining and denoising of positive examples. ART, in contrast, only requires access to unpaired inputs and outputs (e.g., questions and potential answer passages). It uses a new passage-retrieval autoencoding scheme, where (1) an input question is used to retrieve a set of evidence passages, and (2) the passages are then used to compute the probability of reconstructing the original question. Training for retrieval based on question reconstruction enables effective unsupervised learning of both passage and question encoders, which can be later incorporated into complete Open QA systems without any further finetuning. Extensive experiments demonstrate that ART obtains state-of-the-art results on multiple QA retrieval benchmarks with only generic initialization from a pre-trained language model, removing the need for labeled data and task-specific losses.1<\/jats:p>\n               <jats:p>Our code and model checkpoints are available at: https:\/\/github.com\/DevSinghSachan\/art.<\/jats:p>","DOI":"10.1162\/tacl_a_00564","type":"journal-article","created":{"date-parts":[[2023,6,23]],"date-time":"2023-06-23T02:13:26Z","timestamp":1687486406000},"page":"600-616","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":23,"title":["Questions Are All You Need to Train a Dense Passage Retriever"],"prefix":"10.1162","volume":"11","author":[{"given":"Devendra Singh","family":"Sachan","sequence":"first","affiliation":[{"name":"McGill University, Canada. sachande@mila.quebec"},{"name":"Mila - Quebec AI Institute, Canada. sachande@mila.quebec"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mike","family":"Lewis","sequence":"additional","affiliation":[{"name":"Meta AI, USA. mikelewis@meta.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dani","family":"Yogatama","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA. dyogatama@google.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Luke","family":"Zettlemoyer","sequence":"additional","affiliation":[{"name":"Meta AI, USA. lsz@meta.com"},{"name":"University of Washington, USA. lsz@meta.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joelle","family":"Pineau","sequence":"additional","affiliation":[{"name":"McGill University, Canada. jpineau@meta.com"},{"name":"Mila - Quebec AI Institute, Canada. jpineau@meta.com"},{"name":"Meta AI, USA. jpineau@meta.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Manzil","family":"Zaheer","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA. manzilzaheer@google.com"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2023,6,20]]},"reference":[{"key":"2023062214083023000_bib1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.46","article-title":"XOR QA: Cross-lingual open-retrieval question answering","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Asai","year":"2021"},{"key":"2023062214083023000_bib2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1611.09268","article-title":"MS MARCO: A human generated machine reading comprehension dataset","author":"Bajaj","year":"2016","journal-title":"arXiv preprint arXiv:1611.09268"},{"key":"2023062214083023000_bib3","article-title":"Semantic parsing on Freebase from question-answer pairs","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing","author":"Berant","year":"2013"},{"key":"2023062214083023000_bib4","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531863","article-title":"Inpars: Unsupervised dataset generation for information retrieval","volume-title":"Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Bonifacio","year":"2022"},{"key":"2023062214083023000_bib5","doi-asserted-by":"publisher","DOI":"10.1142\/9789812797926_0003","article-title":"Signature verification using a \u201csiamese\u201d time delay neural network","volume-title":"Advances in Neural Information Processing Systems","author":"Bromley","year":"1994"},{"key":"2023062214083023000_bib6","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2204.02311","article-title":"PaLM: Scaling language modeling with pathways","volume":"abs\/2204.02311","author":"Chowdhery","year":"2022","journal-title":"ArXiv"},{"key":"2023062214083023000_bib7","doi-asserted-by":"publisher","first-page":"454","DOI":"10.1162\/tacl_a_00317","article-title":"TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages","volume":"8","author":"Clark","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2023062214083023000_bib8","article-title":"Dialog inpainting: Turning documents into dialogs","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"Dai","year":"2022"},{"key":"2023062214083023000_bib9","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2023062214083023000_bib10","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.203","article-title":"Unsupervised corpus aware language model pre-training for dense passage retrieval","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Gao","year":"2022"},{"key":"2023062214083023000_bib11","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/K19-1049","article-title":"Learning dense representations for entity retrieval","volume-title":"Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL)","author":"Gillick","year":"2019"},{"key":"2023062214083023000_bib12","article-title":"Retrieval augmented language model pre-training","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Guu","year":"2020"},{"key":"2023062214083023000_bib13","article-title":"Poly-encoders: Architectures and pre-training strategies for fast and accurate multi-sentence scoring","volume-title":"International Conference on Learning Representations","author":"Humeau","year":"2020"},{"key":"2023062214083023000_bib14","article-title":"Unsupervised dense information retrieval with contrastive learning","author":"Izacard","year":"2022","journal-title":"Transactions on Machine Learning Research"},{"key":"2023062214083023000_bib15","article-title":"Distilling knowledge from reader to retriever for question answering","volume-title":"International Conference on Learning Representations","author":"Izacard","year":"2021"},{"key":"2023062214083023000_bib16","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1147","article-title":"TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Joshi","year":"2017"},{"key":"2023062214083023000_bib17","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.550","article-title":"Dense passage retrieval for open-domain question answering","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Karpukhin","year":"2020"},{"key":"2023062214083023000_bib18","doi-asserted-by":"publisher","first-page":"929","DOI":"10.1162\/tacl_a_00405","article-title":"Relevance-guided supervision for openqa with colbert","volume":"9","author":"Khattab","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2023062214083023000_bib19","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401075","article-title":"Colbert: Efficient and effective passage search via contextualized late interaction over BERT","volume-title":"Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Khattab","year":"2020"},{"key":"2023062214083023000_bib20","doi-asserted-by":"publisher","first-page":"453","DOI":"10.1162\/tacl_a_00276","article-title":"Natural questions: A benchmark for question answering research","volume":"7","author":"Kwiatkowski","year":"2019","journal-title":"Transactions of the Association of Computational Linguistics"},{"key":"2023062214083023000_bib21","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1612","article-title":"Latent retrieval for weakly supervised open domain question answering","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Lee","year":"2019"},{"key":"2023062214083023000_bib22","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.243","article-title":"The power of scale for parameter-efficient prompt tuning","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Lester","year":"2021"},{"issue":"6624","key":"2023062214083023000_bib23","doi-asserted-by":"publisher","first-page":"1092","DOI":"10.1126\/science.abq1158","article-title":"Competition-level code generation with alphacode","volume":"378","author":"Li","year":"2022","journal-title":"Science"},{"issue":"4","key":"2023062214083023000_bib24","doi-asserted-by":"publisher","first-page":"1","DOI":"10.2200\/S01123ED1V01Y202108HLT053","article-title":"Pretrained transformers for text ranking: BERT and beyond","volume":"14","author":"Lin","year":"2021","journal-title":"Synthesis Lectures on Human Language Technologies"},{"key":"2023062214083023000_bib25","article-title":"Efficient training of retrieval models using negative cache","volume-title":"Advances in Neural Information Processing Systems","author":"Lindgren","year":"2021"},{"key":"2023062214083023000_bib26","doi-asserted-by":"publisher","first-page":"329","DOI":"10.1162\/tacl_a_00369","article-title":"Sparse, dense, and attentional representations for text retrieval","volume":"9","author":"Yi","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2023062214083023000_bib27","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.eacl-main.92","article-title":"Zero-shot neural passage retrieval via domain-targeted synthetic question generation","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Ji","year":"2021"},{"key":"2023062214083023000_bib28","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.316","article-title":"Generation-augmented retrieval for open-domain question answering","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Mao","year":"2021"},{"key":"2023062214083023000_bib29","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2201.10005","article-title":"Text and code embeddings by contrastive pre-training","author":"Neelakantan","year":"2022","journal-title":"arXiv preprint arXiv:2201.10005"},{"key":"2023062214083023000_bib30","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1901.04085","article-title":"Passage re-ranking with BERT","author":"Nogueira","year":"2019","journal-title":"arXiv preprint arXiv:1901.04085"},{"key":"2023062214083023000_bib31","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.63","article-title":"Document ranking with a pretrained sequence-to-sequence model","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Nogueira","year":"2020"},{"key":"2023062214083023000_bib32","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.134","article-title":"Beyond [CLS] through ranking by generation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Nogueira dos Santos","year":"2020"},{"key":"2023062214083023000_bib33","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1807.03748","article-title":"Representation learning with contrastive predictive coding","author":"van den Oord","year":"2018","journal-title":"arXiv preprint arXiv:1807.03748"},{"key":"2023062214083023000_bib34","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.466","article-title":"RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Yingqi","year":"2021"},{"issue":"140","key":"2023062214083023000_bib35","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"2023062214083023000_bib36","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1264","article-title":"SQuAD: 100,000+ questions for machine comprehension of text","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing","author":"Rajpurkar","year":"2016"},{"key":"2023062214083023000_bib37","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.193","article-title":"Learning to retrieve passages without supervision","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Ram","year":"2022"},{"key":"2023062214083023000_bib38","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.224","article-title":"RocketQAv2: A joint training method for dense passage retrieval and passage re-ranking","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Ren","year":"2021"},{"key":"2023062214083023000_bib39","doi-asserted-by":"publisher","DOI":"10.1561\/1500000019","article-title":"The probabilistic relevance framework: BM25 and beyond","author":"Robertson","year":"2009","journal-title":"Foundations and Trends in Information Retrieval"},{"key":"2023062214083023000_bib40","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2204.07496","article-title":"Improving passage retrieval with zero-shot question generation","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Sachan","year":"2022"},{"key":"2023062214083023000_bib41","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.519","article-title":"End-to-end training of neural retrievers for open-domain question answering","volume-title":"Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP)","author":"Sachan","year":"2021"},{"key":"2023062214083023000_bib42","article-title":"End-to-end training of multi-document reader and retriever for open-domain question answering","volume-title":"Advances in Neural Information Processing Systems","author":"Sachan","year":"2021"},{"key":"2023062214083023000_bib43","article-title":"Multitask prompted training enables zero-shot task generalization","volume-title":"International Conference on Learning Representations","author":"Sanh","year":"2022"},{"key":"2023062214083023000_bib44","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.496","article-title":"Simple entity-centric questions challenge dense retrievers","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Sciavolino","year":"2021"},{"key":"2023062214083023000_bib45","article-title":"BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models","volume-title":"Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)","author":"Thakur","year":"2021"},{"key":"2023062214083023000_bib46","article-title":"Attention is all you need","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani","year":"2017"},{"key":"2023062214083023000_bib47","article-title":"Approximate nearest neighbor negative contrastive learning for dense text retrieval","volume-title":"International Conference on Learning Representations","author":"Xiong","year":"2021"},{"key":"2023062214083023000_bib48","article-title":"Adversarial retriever-ranker for dense text retrieval","volume-title":"International Conference on Learning Representations","author":"Zhang","year":"2022"},{"key":"2023062214083023000_bib49","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.493","article-title":"Hyperlink-induced pre-training for passage retrieval in open-domain question answering","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhou","year":"2022"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00564\/2134472\/tacl_a_00564.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00564\/2134472\/tacl_a_00564.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,23]],"date-time":"2023-06-23T02:13:43Z","timestamp":1687486423000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00564\/116466\/Questions-Are-All-You-Need-to-Train-a-Dense"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023]]},"references-count":49,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00564","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023]]},"published":{"date-parts":[[2023]]}}}