{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T22:15:59Z","timestamp":1765232159860},"reference-count":23,"publisher":"Association for Computing Machinery (ACM)","issue":"1-2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2010,9]]},"abstract":"<jats:p>\n            MapReduce is a computing paradigm that has gained a lot of attention in recent years from industry and research. Unlike parallel DBMSs, MapReduce allows non-expert users to run complex analytical tasks over very large data sets on very large clusters and clouds. However, this comes at a price: MapReduce processes tasks in a scan-oriented fashion. Hence, the performance of Hadoop --- an open-source implementation of MapReduce --- often does not match the one of a well-configured parallel DBMS. In this paper we propose a new type of system named Hadoop++: it boosts task performance without changing the Hadoop framework at all (Hadoop does not even 'notice it'). To reach this goal, rather than changing a working system (Hadoop), we\n            <jats:italic>inject<\/jats:italic>\n            our technology at the right places through UDFs only and affect Hadoop\n            <jats:italic>from inside<\/jats:italic>\n            . This has three important consequences: First, Hadoop++ significantly outperforms Hadoop. Second, any future changes of Hadoop may directly be used with Hadoop++ without rewriting any glue code. Third, Hadoop++ does not need to change the Hadoop interface. Our experiments show the superiority of Hadoop++ over both Hadoop and HadoopDB for tasks related to indexing and join processing.\n          <\/jats:p>","DOI":"10.14778\/1920841.1920908","type":"journal-article","created":{"date-parts":[[2014,6,24]],"date-time":"2014-06-24T12:17:57Z","timestamp":1403612277000},"page":"515-529","source":"Crossref","is-referenced-by-count":217,"title":["Hadoop++"],"prefix":"10.14778","volume":"3","author":[{"given":"Jens","family":"Dittrich","sequence":"first","affiliation":[{"name":"Saarland University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jorge-Arnulfo","family":"Quian\u00e9-Ruiz","sequence":"additional","affiliation":[{"name":"Saarland University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alekh","family":"Jindal","sequence":"additional","affiliation":[{"name":"Saarland University and International Max Planck Research School for Computer Science"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yagiz","family":"Kargin","sequence":"additional","affiliation":[{"name":"International Max Planck Research School for Computer Science"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vinay","family":"Setty","sequence":"additional","affiliation":[{"name":"International Max Planck Research School for Computer Science"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J\u00f6rg","family":"Schad","sequence":"additional","affiliation":[{"name":"Saarland University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2010,9]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Dbcolumn on MapReduce http:\/\/databasecolumn.vertica.com\/2008\/01\/mapreduce-a-major-step-back.html.  Dbcolumn on MapReduce http:\/\/databasecolumn.vertica.com\/2008\/01\/mapreduce-a-major-step-back.html."},{"key":"e_1_2_1_2_1","unstructured":"HDFS Bug http:\/\/issues.apache.org\/jira\/browse\/HDFS-96.  HDFS Bug http:\/\/issues.apache.org\/jira\/browse\/HDFS-96."},{"key":"e_1_2_1_3_1","volume-title":"HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads. PVLDB, 2(1)","author":"Abouzeid A.","year":"2009","unstructured":"A. Abouzeid , K. Bajda-Pawlikowski , D. Abadi , A. Silberschatz , and A. Rasin . HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads. PVLDB, 2(1) , 2009 . A. Abouzeid, K. Bajda-Pawlikowski, D. Abadi, A. Silberschatz, and A. Rasin. HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads. PVLDB, 2(1), 2009."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1739041.1739056"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/319983.319987"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1859127.1859141"},{"key":"e_1_2_1_7_1","volume-title":"Scope: Easy and Efficient Parallel Processing of Massive Data Sets. PVLDB, 1(2)","author":"Chaiken R.","year":"2008","unstructured":"R. Chaiken Scope: Easy and Efficient Parallel Processing of Massive Data Sets. PVLDB, 1(2) , 2008 . R. Chaiken et al. Scope: Easy and Efficient Parallel Processing of Massive Data Sets. PVLDB, 1(2), 2008."},{"key":"e_1_2_1_8_1","volume-title":"Mad Skills: New Analysis Practices for Big Data. PVLDB, 2(2)","author":"Cohen J.","year":"2009","unstructured":"J. Cohen , B. Dolan , M. Dunlap , J. Hellerstein , and C. Welton . Mad Skills: New Analysis Practices for Big Data. PVLDB, 2(2) , 2009 . J. Cohen, B. Dolan, M. Dunlap, J. Hellerstein, and C. Welton. Mad Skills: New Analysis Practices for Big Data. PVLDB, 2(2), 2009."},{"key":"e_1_2_1_9_1","volume-title":"NSDI","author":"Condie T.","year":"2010","unstructured":"T. Condie , N. Conway , P. Alvaro , J. M. Hellerstein , K. Elmeleegy , and R. Sears . MapReduce Online . In NSDI , 2010 . T. Condie, N. Conway, P. Alvaro, J. M. Hellerstein, K. Elmeleegy, and R. Sears. MapReduce Online. In NSDI, 2010."},{"key":"e_1_2_1_10_1","volume-title":"OSDI","author":"Dean J.","year":"2004","unstructured":"J. Dean and S. Ghemawat . Mapreduce: Simplified Data Processing on Large Clusters . In OSDI , 2004 . J. Dean and S. Ghemawat. Mapreduce: Simplified Data Processing on Large Clusters. In OSDI, 2004."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1629175.1629198"},{"key":"e_1_2_1_12_1","volume-title":"Building a HighLevel Dataflow System on Top of MapReduce: The Pig Experience. PVLDB, 2(2)","author":"Gates A.","year":"2009","unstructured":"A. Gates Building a HighLevel Dataflow System on Top of MapReduce: The Pig Experience. PVLDB, 2(2) , 2009 . A. Gates et al. Building a HighLevel Dataflow System on Top of MapReduce: The Pig Experience. PVLDB, 2(2), 2009."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1272996.1273005"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2010.5447919"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376726"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559865"},{"key":"e_1_2_1_17_1","volume-title":"VLDB","author":"Rao J.","year":"1999","unstructured":"J. Rao and K. A. Ross . Cache Conscious Indexing for Decision-Support in Main Memory . In VLDB , 1999 . J. Rao and K. A. Ross. Cache Conscious Indexing for Decision-Support in Main Memory. In VLDB, 1999."},{"key":"e_1_2_1_18_1","volume-title":"Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance. PVLDB, 3(1)","author":"Schad J.","year":"2010","unstructured":"J. Schad , J. Dittrich , and J.-A. Quiane-Ruiz . Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance. PVLDB, 3(1) , 2010 . J. Schad, J. Dittrich, and J.-A. Quiane-Ruiz. Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance. PVLDB, 3(1), 2010."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1629175.1629197"},{"key":"e_1_2_1_20_1","volume-title":"Hive - a warehousing solution over a map-reduce framework. PVLDB, 2(2)","author":"Thusoo A.","year":"2009","unstructured":"A. Thusoo Hive - a warehousing solution over a map-reduce framework. PVLDB, 2(2) , 2009 . A. Thusoo et al. Hive - a warehousing solution over a map-reduce framework. PVLDB, 2(2), 2009."},{"key":"e_1_2_1_21_1","volume-title":"CASCON","author":"Yan P.","year":"1994","unstructured":"P. Yan and P. Larson . Data Reduction Through Early Grouping . In CASCON , 1994 . P. Yan and P. Larson. Data Reduction Through Early Grouping. In CASCON, 1994."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2010.5447913"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1247480.1247602"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/1920841.1920908","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:45:33Z","timestamp":1672227933000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/1920841.1920908"}},"subtitle":["making a yellow elephant run like a cheetah (without it even noticing)"],"short-title":[],"issued":{"date-parts":[[2010,9]]},"references-count":23,"journal-issue":{"issue":"1-2","published-print":{"date-parts":[[2010,9]]}},"alternative-id":["10.14778\/1920841.1920908"],"URL":"https:\/\/doi.org\/10.14778\/1920841.1920908","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2010,9]]}}}