{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T06:13:31Z","timestamp":1768284811361,"version":"3.49.0"},"reference-count":40,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,2]]},"abstract":"<jats:p>In recent years, neural networks have shown impressive performance gains on long-standing AI problems, such as answering queries from text and machine translation. These advances raise the question of whether neural nets can be used at the core of query processing to derive answers from facts, even when the facts are expressed in natural language. If so, it is conceivable that we could relax the fundamental assumption of database management, namely, that our data is represented as fields of a pre-defined schema. Furthermore, such technology would enable combining information from text, images, and structured data seamlessly.<\/jats:p>\n          <jats:p>\n            This paper introduces\n            <jats:italic>neural databases<\/jats:italic>\n            , a class of systems that use NLP transformers as localized answer derivation engines. We ground the vision in NeuralDB, a system for querying facts represented as short natural language sentences. We demonstrate that recent natural language processing models, specifically transformers, can answer select-project-join queries if they are given a set of relevant facts. However, they cannot scale to non-trivial databases nor answer set-based and aggregation queries. Based on these insights, we identify specific research challenges that are needed to build neural databases. Some of the challenges require drawing upon the rich literature in data management, and others pose new research opportunities to the NLP community. Finally, we show that with preliminary solutions, NeuralDB can already answer queries over thousands of sentences with very high accuracy.\n          <\/jats:p>","DOI":"10.14778\/3447689.3447706","type":"journal-article","created":{"date-parts":[[2021,4,12]],"date-time":"2021-04-12T16:20:06Z","timestamp":1618244406000},"page":"1033-1039","source":"Crossref","is-referenced-by-count":30,"title":["From natural language processing to neural databases"],"prefix":"10.14778","volume":"14","author":[{"given":"James","family":"Thorne","sequence":"first","affiliation":[{"name":"University of Cambridge and Facebook AI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Majid","family":"Yazdani","sequence":"additional","affiliation":[{"name":"Facebook AI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marzieh","family":"Saeidi","sequence":"additional","affiliation":[{"name":"Facebook AI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fabrizio","family":"Silvestri","sequence":"additional","affiliation":[{"name":"Facebook AI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sebastian","family":"Riedel","sequence":"additional","affiliation":[{"name":"Facebook AI and University College London"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alon","family":"Halevy","sequence":"additional","affiliation":[{"name":"Facebook AI"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,4,12]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Large scale knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training","author":"Agarwal Oshin","year":"2020","unstructured":"Oshin Agarwal , Heming Ge , Siamak Shakeri , and Rami Al-Rfou . Large scale knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training , 2020 . Oshin Agarwal, Heming Ge, Siamak Shakeri, and Rami Al-Rfou. Large scale knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training, 2020."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.12"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1017\/S135132490000005X"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_2_1_5_1","volume-title":"International Conference on Learning Representations","author":"Asai Akari","year":"2020","unstructured":"Akari Asai , Kazuma Hashimoto , Hannaneh Hajishirzi , Richard Socher , and Caiming Xiong . Learning to retrieve reasoning paths over wikipedia graph for question answering . In International Conference on Learning Representations , 2020 . Akari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher, and Caiming Xiong. Learning to retrieve reasoning paths over wikipedia graph for question answering. In International Conference on Learning Representations, 2020."},{"key":"e_1_2_1_6_1","volume-title":"Language Models are Few-Shot Learners","author":"Brown Tom B","year":"2020","unstructured":"Tom B Brown , Benjamin Mann , Nick Ryder , Melanie Subbiah , Jared Kaplan , Prafulla Dhariwal , Arvind Neelakantan , Pranav Shyam , Girish Sastry , Amanda Askell , Sandhini Agarwal , Ariel Herbert-Voss , Gretchen Krueger , Tom Henighan , Rewon Child , Aditya Ramesh , Daniel M Ziegler , Jeffrey Wu , Clemens Winter , Christopher Hesse , Mark Chen , Eric Sigler , Mateusz Litwin , Scott Gray , Benjamin Chess , Jack Clark , Christopher Berner , Sam McCandlish , Alec Radford , Ilya Sutskever , and Dario Amodei . Language Models are Few-Shot Learners . 2020 . Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language Models are Few-Shot Learners. 2020."},{"key":"e_1_2_1_7_1","volume-title":"International Conference on Learning Representations","author":"Cao Nicola De","year":"2021","unstructured":"Nicola De Cao , Gautier Izacard , Sebastian Riedel , and Fabio Petroni . Autoregressive entity retrieval . In International Conference on Learning Representations , 2021 . Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. Autoregressive entity retrieval. In International Conference on Learning Representations, 2021."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1246"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018","author":"Elsahar Hady","year":"2018","unstructured":"Hady Elsahar , Pavlos Vougiouklis , Arslen Remaci , Christophe Gravier , Jonathon Hare , Frederique Laforest , and Elena Simperl . T-REx : A large scale alignment of natural language with knowledge base triples . In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018 ), Miyazaki, Japan , May 2018 . European Language Resources Association (ELRA). Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frederique Laforest, and Elena Simperl. T-REx: A large scale alignment of natural language with knowledge base triples. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan, May 2018. European Language Resources Association (ELRA)."},{"key":"e_1_2_1_11_1","volume-title":"Neural turing machines. CoRR, abs\/1410.5401","author":"Graves Alex","year":"2014","unstructured":"Alex Graves , Greg Wayne , and Ivo Danihelka . Neural turing machines. CoRR, abs\/1410.5401 , 2014 . Alex Graves, Greg Wayne, and Ivo Danihelka. Neural turing machines. CoRR, abs\/1410.5401, 2014."},{"key":"e_1_2_1_12_1","volume-title":"Retrieval-Augmented Language Model Pre-Training","author":"Guu Kelvin","year":"2020","unstructured":"Kelvin Guu , Kenton Lee , Zora Tung , Panupong Pasupat , and Ming-wei Chang. REALM : Retrieval-Augmented Language Model Pre-Training , 2020 . Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-wei Chang. REALM : Retrieval-Augmented Language Model Pre-Training, 2020."},{"key":"e_1_2_1_13_1","volume-title":"CIDR 2003, First Biennial Conference on Innovative Data Systems Research, Asilomar, CA, USA, January 5-8, 2003, Online Proceedings. www.cidrdb.org","author":"Halevy Alon Y.","year":"2003","unstructured":"Alon Y. Halevy , Oren Etzioni , AnHai Doan , Zachary G. Ives , Jayant Madhavan , Luke K. McDowell , and Igor Tatarinov . Crossing the structure chasm . In CIDR 2003, First Biennial Conference on Innovative Data Systems Research, Asilomar, CA, USA, January 5-8, 2003, Online Proceedings. www.cidrdb.org , 2003 . Alon Y. Halevy, Oren Etzioni, AnHai Doan, Zachary G. Ives, Jayant Madhavan, Luke K. McDowell, and Igor Tatarinov. Crossing the structure chasm. In CIDR 2003, First Biennial Conference on Innovative Data Systems Research, Asilomar, CA, USA, January 5-8, 2003, Online Proceedings. www.cidrdb.org, 2003."},{"key":"e_1_2_1_14_1","first-page":"2790","volume-title":"Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby , Andrei Giurgiu , Stanislaw Jastrzebski , Bruna Morrone , Quentin De Laroussilhe , Andrea Gesmundo , Mona Attariyan , and Sylvain Gelly . Parameter-efficient transfer learning for NLP. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors , Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research , pages 2790 -- 2799 , Long Beach, California, USA, 09- -15 Jun 2019 . PMLR. Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for NLP. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 2790--2799, Long Beach, California, USA, 09--15 Jun 2019. PMLR."},{"key":"e_1_2_1_15_1","volume-title":"Dense Passage Retrieval for Open-Domain Question Answering","author":"Karpukhin Vladimir","year":"2020","unstructured":"Vladimir Karpukhin , Barlas O\u011fuz , Sewon Min , Patrick Lewis , Ledell Wu , Sergey Edunov , Danqi Chen , and Wen-tau Yih. Dense Passage Retrieval for Open-Domain Question Answering . 2020 . Vladimir Karpukhin, Barlas O\u011fuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense Passage Retrieval for Open-Domain Question Answering. 2020."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196909"},{"key":"e_1_2_1_17_1","volume-title":"6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net","author":"Lample Guillaume","year":"2018","unstructured":"Guillaume Lample , Alexis Conneau , Marc'Aurelio Ranzato , Ludovic Denoyer , and Herv\u00e9 J\u00e9gou . Word translation without parallel data . In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net , 2018 . Guillaume Lample, Alexis Conneau, Marc'Aurelio Ranzato, Ludovic Denoyer, and Herv\u00e9 J\u00e9gou. Word translation without parallel data. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018."},{"key":"e_1_2_1_18_1","volume-title":"NeurIPS","author":"Lewis Patrick","year":"2020","unstructured":"Patrick Lewis , Ethan Perez , Aleksandara Piktus , Fabio Petroni , Vladimir Karpukhin , Naman Goyal , Heinrich K\u00fcttler , Mike Lewis , Wen-tau Yih, Tim Rockt\u00e4schel , Sebastian Riedel , and Douwe Kiela . Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks . In NeurIPS , 2020 . Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\u00fcttler, Mike Lewis, Wen-tau Yih, Tim Rockt\u00e4schel, Sebastian Riedel, and Douwe Kiela. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In NeurIPS, 2020."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735461.2735468"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.14778\/3421424.3421431"},{"key":"e_1_2_1_21_1","volume-title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu , Myle Ott , Naman Goyal , Jingfei Du , Mandar Joshi , Danqi Chen , Omer Levy , Mike Lewis , Luke Zettlemoyer , and Veselin Stoyanov . RoBERTa: A Robustly Optimized BERT Pretraining Approach . 2019 . Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A Robustly Optimized BERT Pretraining Approach. 2019."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5962"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196926"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1179"},{"key":"e_1_2_1_25_1","volume-title":"James Thorne, Yacine Jernite, Vassilis Plachouras, Tim Rockt\u00e4schel, et al. Kilt: a benchmark for knowledge intensive language tasks. arXiv preprint arXiv:2009.02252","author":"Petroni Fabio","year":"2020","unstructured":"Fabio Petroni , Aleksandra Piktus , Angela Fan , Patrick Lewis , Majid Yazdani , Nicola De Cao , James Thorne, Yacine Jernite, Vassilis Plachouras, Tim Rockt\u00e4schel, et al. Kilt: a benchmark for knowledge intensive language tasks. arXiv preprint arXiv:2009.02252 , 2020 . Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vassilis Plachouras, Tim Rockt\u00e4schel, et al. Kilt: a benchmark for knowledge intensive language tasks. arXiv preprint arXiv:2009.02252, 2020."},{"key":"e_1_2_1_26_1","volume-title":"Sebastian Riedel. In Proceedings of EMNLP-IJCNLP","author":"Petroni Fabio","unstructured":"Fabio Petroni , Tim Rockt\u00e4schel , Patrick Lewis , Anton Bakhtin , Yuxiang Wu , Alexander H. Miller , and Sebastian Riedel. In Proceedings of EMNLP-IJCNLP , Hong Kong, China. Fabio Petroni, Tim Rockt\u00e4schel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. In Proceedings of EMNLP-IJCNLP, Hong Kong, China."},{"key":"e_1_2_1_27_1","first-page":"1","article-title":"Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel , Noam Shazeer , Adam Roberts , Katherine Lee , Sharan Narang , Michael Matena , Yanqi Zhou , Wei Li , and Peter J. Liu . Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer . Journal of Machine Learning Research , 21 : 1 -- 67 , 2020 . Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research, 21:1--67, 2020.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-2124"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294996.3295136"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1233"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969442.2969512"},{"key":"e_1_2_1_32_1","volume-title":"Relational pretrained transformers towards democratizing data preparation [vision]. CoRR, abs\/2012.02469","author":"Tang Nan","year":"2020","unstructured":"Nan Tang , Ju Fan , Fangyi Li , Jianhong Tu , Xiaoyong Du , Guoliang Li , Sam Madden , and Mourad Ouzzani . Relational pretrained transformers towards democratizing data preparation [vision]. CoRR, abs\/2012.02469 , 2020 . Nan Tang, Ju Fan, Fangyi Li, Jianhong Tu, Xiaoyong Du, Guoliang Li, Sam Madden, and Mourad Ouzzani. Relational pretrained transformers towards democratizing data preparation [vision]. CoRR, abs\/2012.02469, 2020."},{"key":"e_1_2_1_33_1","first-page":"1","volume-title":"ICLR","author":"Tenney Ian","year":"2019","unstructured":"Ian Tenney , Patrick Xia , Berlin Chen , Alex Wang , Adam Poliak , R. Thomas McCoy , Najoung Kim , Benjamin Van Durme , Samuel R Bowman , Dipanjan Das , and Ellie Pavlick . What do you learn from context? Probing for sentence structure in contextualized word representations . ICLR , pages 1 -- 17 , 2019 . Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R.Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, and Ellie Pavlick. What do you learn from context? Probing for sentence structure in contextualized word representations. ICLR, pages 1--17, 2019."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1074"},{"key":"e_1_2_1_35_1","volume-title":"Neural databases. CoRR, abs\/2010.06973","author":"Thorne James","year":"2020","unstructured":"James Thorne , Majid Yazdani , Marzieh Saeidi , Fabrizio Silvestri , Sebastian Riedel , and Alon Y. Halevy . Neural databases. CoRR, abs\/2010.06973 , 2020 . James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel, and Alon Y. Halevy. Neural databases. CoRR, abs\/2010.06973, 2020."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2629489"},{"key":"e_1_2_1_38_1","volume-title":"Towards ai-complete question answering: A set of prerequisite toy tasks. arXiv preprint arXiv:1502.05698","author":"Weston Jason","year":"2015","unstructured":"Jason Weston , Antoine Bordes , Sumit Chopra , Alexander M Rush , Bart van Merri\u00ebnboer , Armand Joulin , and Tomas Mikolov . Towards ai-complete question answering: A set of prerequisite toy tasks. arXiv preprint arXiv:1502.05698 , 2015 . Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart van Merri\u00ebnboer, Armand Joulin, and Tomas Mikolov. Towards ai-complete question answering: A set of prerequisite toy tasks. arXiv preprint arXiv:1502.05698, 2015."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00309"},{"key":"e_1_2_1_40_1","unstructured":"Jichuan Zeng Xi Victoria Lin Caiming Xiong Richard Socher Michael R. Lyu Irwin King and Steven C. H. Hoi. Photon: A robust cross-domain text-to-sql system.  Jichuan Zeng Xi Victoria Lin Caiming Xiong Richard Socher Michael R. Lyu Irwin King and Steven C. H. Hoi. Photon: A robust cross-domain text-to-sql system."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3447689.3447706","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:20:33Z","timestamp":1672226433000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3447689.3447706"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2]]},"references-count":40,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2021,2]]}},"alternative-id":["10.14778\/3447689.3447706"],"URL":"https:\/\/doi.org\/10.14778\/3447689.3447706","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,2]]}}}