{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T01:25:20Z","timestamp":1781054720864,"version":"3.54.1"},"reference-count":26,"publisher":"Association for Computing Machinery (ACM)","issue":"8","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,4]]},"abstract":"<jats:p>\n            Analysts frequently require data from multiple sources for their tasks, but finding these sources is challenging in exabyte-scale data lakes. In this paper, we address this problem for our enterprise's data lake by using machine-learning to identify related data sources. Leveraging queries made to the data lake over a month, we build a\n            <jats:italic>relevance model<\/jats:italic>\n            that determines whether two columns across two data streams are related or not. We then use the model to find relations at scale across tens of millions of column-pairs and thereafter construct a\n            <jats:italic>data relationship graph<\/jats:italic>\n            in a scalable fashion, processing a data lake that has 4.5 Petabytes of data in approximately 80 minutes. Using manually labeled datasets as ground-truth, we show that our techniques show improvements of at least 23% when compared to state-of-the-art methods.\n          <\/jats:p>","DOI":"10.14778\/3457390.3457403","type":"journal-article","created":{"date-parts":[[2021,10,21]],"date-time":"2021-10-21T22:48:38Z","timestamp":1634856518000},"page":"1392-1400","source":"Crossref","is-referenced-by-count":17,"title":["Discovering related data at scale"],"prefix":"10.14778","volume":"14","author":[{"given":"Sagar","family":"Bharadwaj","sequence":"first","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Praveen","family":"Gupta","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ranjita","family":"Bhagwan","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Saikat","family":"Guha","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,10,21]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1242572.1242591"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1010933404324"},{"key":"e_1_2_1_3_1","volume-title":"Aurum: A Data Discovery System. In 2018 IEEE 34th International Conference on Data Engineering (ICDE). IEEE","author":"Fernandez R. Castro","unstructured":"R. Castro Fernandez , Z. Abedjan , F. Koko , G. Yuan , S. Madden , and M. Stonebraker . 2018 . Aurum: A Data Discovery System. In 2018 IEEE 34th International Conference on Data Engineering (ICDE). IEEE , Paris, France, 1001--1012. R. Castro Fernandez, Z. Abedjan, F. Koko, G. Yuan, S. Madden, and M. Stonebraker. 2018. Aurum: A Data Discovery System. In 2018 IEEE 34th International Conference on Data Engineering (ICDE). IEEE, Paris, France, 1001--1012."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.14778\/1454159.1454166"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/872757.872796"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2006.9"},{"key":"e_1_2_1_7_1","volume-title":"2019 IEEE Symposium on Security and Privacy (SP). IEEE","author":"Chia P. H.","unstructured":"P. H. Chia , D. Desfontaines , I. M. Perera , D. Simmons-Marengo , C. Li , W. Day , Q. Wang , and M. Guevara . 2019. KHyperLogLog: Estimating Reidentifiability and Joinability of Large Data at Scale . In 2019 IEEE Symposium on Security and Privacy (SP). IEEE , San Francisco, CA, USA, 350--364. P. H. Chia, D. Desfontaines, I. M. Perera, D. Simmons-Marengo, C. Li, W. Day, Q. Wang, and M. Guevara. 2019. KHyperLogLog: Estimating Reidentifiability and Joinability of Large Data at Scale. In 2019 IEEE Symposium on Security and Privacy (SP). IEEE, San Francisco, CA, USA, 350--364."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/89086.89095"},{"key":"e_1_2_1_9_1","volume-title":"The Data Civilizer System. In CIDR 2017, 8th Biennial Conference on Innovative Data Systems Research, Chaminade, CA, USA, January 8--11, 2017, Online Proceedings. www.cidrdb.org","author":"Deng Dong","year":"2017","unstructured":"Dong Deng , Raul Castro Fernandez , Ziawasch Abedjan , Sibo Wang , Michael Stonebraker , Ahmed K. Elmagarmid , Ihab F. Ilyas , Samuel Madden , Mourad Ouzzani , and Nan Tang . 2017 . The Data Civilizer System. In CIDR 2017, 8th Biennial Conference on Innovative Data Systems Research, Chaminade, CA, USA, January 8--11, 2017, Online Proceedings. www.cidrdb.org , Chaminade, California. http:\/\/cidrdb.org\/cidr 2017\/papers\/p44-deng-cidr17.pdf Dong Deng, Raul Castro Fernandez, Ziawasch Abedjan, Sibo Wang, Michael Stonebraker, Ahmed K. Elmagarmid, Ihab F. Ilyas, Samuel Madden, Mourad Ouzzani, and Nan Tang. 2017. The Data Civilizer System. In CIDR 2017, 8th Biennial Conference on Innovative Data Systems Research, Chaminade, CA, USA, January 8--11, 2017, Online Proceedings. www.cidrdb.org, Chaminade, California. http:\/\/cidrdb.org\/cidr2017\/papers\/p44-deng-cidr17.pdf"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3196398.3196448"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.14778\/3231751.3231760"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcss.1997.1504"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1214\/aos\/1013203451"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/645927.672200"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/775152.775166"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352116"},{"key":"e_1_2_1_17_1","first-page":"219","article-title":"Natural language corpus data. In Beautiful data. O'Reilly Media, Boston, USA","volume":"14","author":"Norvig Peter","year":"2009","unstructured":"Peter Norvig . 2009 . Natural language corpus data. In Beautiful data. O'Reilly Media, Boston, USA , Chapter 14 , 219 -- 242 . Peter Norvig. 2009. Natural language corpus data. In Beautiful data. O'Reilly Media, Boston, USA, Chapter 14, 219--242.","journal-title":"Chapter"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_2_1_19_1","volume-title":"WebDB.","author":"Rostin Alexandra","unstructured":"Alexandra Rostin , Oliver Albrecht , Jana Bauckmann , Felix Naumann , and Ulf Leser . 2009. A machine learning approach to foreign key discovery . In WebDB. Providence, Rhode Island . Alexandra Rostin, Oliver Albrecht, Jana Bauckmann, Felix Naumann, and Ulf Leser. 2009. A machine learning approach to foreign key discovery. In WebDB. Providence, Rhode Island."},{"key":"e_1_2_1_20_1","first-page":"97","article-title":"On estimation of the size of the dictionary of a long text on the basis of a sample","volume":"19","author":"Shlosser A","year":"1981","unstructured":"A Shlosser . 1981 . On estimation of the size of the dictionary of a long text on the basis of a sample . Engineering Cybernetics 19 , 1 (1981), 97 -- 102 . A Shlosser. 1981. On estimation of the size of the dictionary of a long text on the basis of a sample. Engineering Cybernetics 19, 1 (1981), 97--102.","journal-title":"Engineering Cybernetics"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767865"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1186\/1758-2946-5-23"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352095"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-012-0280-z"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3300065"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994534"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3457390.3457403","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:46:30Z","timestamp":1672224390000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3457390.3457403"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4]]},"references-count":26,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2021,4]]}},"alternative-id":["10.14778\/3457390.3457403"],"URL":"https:\/\/doi.org\/10.14778\/3457390.3457403","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,4]]}}}