{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T09:56:58Z","timestamp":1773482218267,"version":"3.50.1"},"reference-count":9,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2012,8]]},"abstract":"<jats:p>In this demonstration, we present BlinkDB, a massively parallel, sampling-based approximate query processing framework for running interactive queries on large volumes of data. The key observation in BlinkDB is that one can make reasonable decisions in the absence of perfect answers. BlinkDB extends the Hive\/HDFS stack and can handle the same set of SPJA (selection, projection, join and aggregate) queries as supported by these systems. BlinkDB provides real-time answers along with statistical error guarantees, and can scale to petabytes of data and thousands of machines in a fault-tolerant manner. Our experiments using the TPC-H benchmark and on an anonymized real-world video content distribution workload from Conviva Inc. show that BlinkDB can execute a wide range of queries up to 150x faster than Hive on MapReduce and 10--150x faster than Shark (Hive on Spark) over tens of terabytes of data stored across 100 machines, all with an error of 2--10%.<\/jats:p>","DOI":"10.14778\/2367502.2367533","type":"journal-article","created":{"date-parts":[[2014,6,24]],"date-time":"2014-06-24T12:17:57Z","timestamp":1403612277000},"page":"1902-1905","source":"Crossref","is-referenced-by-count":63,"title":["Blink and it's done"],"prefix":"10.14778","volume":"5","author":[{"given":"Sameer","family":"Agarwal","sequence":"first","affiliation":[{"name":"UC Berkeley"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anand P.","family":"Iyer","sequence":"additional","affiliation":[{"name":"UC Berkeley"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aurojit","family":"Panda","sequence":"additional","affiliation":[{"name":"UC Berkeley"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Samuel","family":"Madden","sequence":"additional","affiliation":[{"name":"MIT CSAIL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Barzan","family":"Mozafari","sequence":"additional","affiliation":[{"name":"MIT CSAIL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ion","family":"Stoica","sequence":"additional","affiliation":[{"name":"UC Berkeley"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,8]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Apache Hive Project. http:\/\/hive.apache.org\/.  Apache Hive Project. http:\/\/hive.apache.org\/."},{"key":"e_1_2_1_2_1","unstructured":"Conviva Inc. http:\/\/www.conviva.com\/.  Conviva Inc. http:\/\/www.conviva.com\/."},{"key":"e_1_2_1_3_1","first-page":"281","volume-title":"NSDI","author":"Agarwal S.","year":"2012","unstructured":"S. Agarwal , S. Kandula , N. Bruno , M.-C. Wu , I. Stoica , and J. Zhou . Re-optimizing Data Parallel Computing . In NSDI , pages 281 -- 294 , 2012 . S. Agarwal, S. Kandula, N. Bruno, M.-C. Wu, I. Stoica, and J. Zhou. Re-optimizing Data Parallel Computing. In NSDI, pages 281--294, 2012."},{"key":"e_1_2_1_5_1","first-page":"805","volume-title":"SIGMOD","author":"Bruno N.","year":"2012","unstructured":"N. Bruno , S. Agarwal , S. Kandula , B. Shi , M.-C. Wu , and J. Zhou . Recurring Job Optimization in Scope . In SIGMOD , pages 805 -- 806 , 2012 . 10.1145\/2213836.2213959 N. Bruno, S. Agarwal, S. Kandula, B. Shi, M.-C. Wu, and J. Zhou. Recurring Job Optimization in Scope. In SIGMOD, pages 805--806, 2012. 10.1145\/2213836.2213959"},{"key":"e_1_2_1_6_1","first-page":"689","volume-title":"Shark: Fast Data Analysis Using Coarse-grained Distributed Memory. In SIGMOD Conference","author":"Engle C.","year":"2012","unstructured":"C. Engle Shark: Fast Data Analysis Using Coarse-grained Distributed Memory. In SIGMOD Conference , pages 689 -- 692 , 2012 . 10.1145\/2213836.2213934 C. Engle et al. Shark: Fast Data Analysis Using Coarse-grained Distributed Memory. In SIGMOD Conference, pages 689--692, 2012. 10.1145\/2213836.2213934"},{"key":"e_1_2_1_7_1","volume-title":"VLDB","author":"Garofalakis M.","year":"2001","unstructured":"M. Garofalakis and P. Gibbons . Approximate Query Processing: Taming the Terabytes . In VLDB , 2001 . Tutorial. M. Garofalakis and P. Gibbons. Approximate Query Processing: Taming the Terabytes. In VLDB, 2001. Tutorial."},{"issue":"12","key":"e_1_2_1_8_1","first-page":"1474","article-title":"The Researcher's Guide to the Data Deluge: Querying a Scientific Database in Just a Few Seconds","volume":"4","author":"Kersten M. L.","year":"2011","unstructured":"M. L. Kersten , S. Idreos , S. Manegold , and E. Liarou . The Researcher's Guide to the Data Deluge: Querying a Scientific Database in Just a Few Seconds . PVLDB , 4 ( 12 ): 1474 -- 1477 , 2011 . M. L. Kersten, S. Idreos, S. Manegold, and E. Liarou. The Researcher's Guide to the Data Deluge: Querying a Scientific Database in Just a Few Seconds. PVLDB, 4(12):1474--1477, 2011.","journal-title":"PVLDB"},{"key":"e_1_2_1_9_1","first-page":"296","volume-title":"CIDR","author":"Sidirourgos L.","year":"2011","unstructured":"L. Sidirourgos : Scientific Data Management With Bounds On Runtime and Quality . In CIDR , pages 296 -- 301 , 2011 . L. Sidirourgos et al. SciBORQ: Scientific Data Management With Bounds On Runtime and Quality. In CIDR, pages 296--301, 2011."},{"key":"e_1_2_1_10_1","first-page":"15","volume-title":"NSDI","author":"Zaharia M.","year":"2012","unstructured":"M. Zaharia Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing . In NSDI , pages 15 -- 28 , 2012 . M. Zaharia et al. Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing. In NSDI, pages 15--28, 2012."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/2367502.2367533","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:49:38Z","timestamp":1672224578000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/2367502.2367533"}},"subtitle":["interactive queries on very large data"],"short-title":[],"issued":{"date-parts":[[2012,8]]},"references-count":9,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2012,8]]}},"alternative-id":["10.14778\/2367502.2367533"],"URL":"https:\/\/doi.org\/10.14778\/2367502.2367533","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2012,8]]}}}