{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,4]],"date-time":"2026-03-04T09:41:48Z","timestamp":1772617308001,"version":"3.50.1"},"reference-count":16,"publisher":"Association for Computing Machinery (ACM)","issue":"1-2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2010,9]]},"abstract":"<jats:p>The growing demand for large-scale data mining and data analysis applications has led both industry and academia to design new types of highly scalable data-intensive computing platforms. MapReduce and Dryad are two popular platforms in which the dataflow takes the form of a directed acyclic graph of operators. These platforms lack built-in support for iterative programs, which arise naturally in many applications including data mining, web ranking, graph analysis, model fitting, and so on. This paper presents HaLoop, a modified version of the Hadoop MapReduce framework that is designed to serve these applications. HaLoop not only extends MapReduce with programming support for iterative applications, it also dramatically improves their efficiency by making the task scheduler loop-aware and by adding various caching mechanisms. We evaluated HaLoop on real queries and real datasets. Compared with Hadoop, on average, HaLoop reduces query runtimes by 1.85, and shuffles only 4% of the data between mappers and reducers.<\/jats:p>","DOI":"10.14778\/1920841.1920881","type":"journal-article","created":{"date-parts":[[2014,6,24]],"date-time":"2014-06-24T12:17:57Z","timestamp":1403612277000},"page":"285-296","source":"Crossref","is-referenced-by-count":486,"title":["HaLoop"],"prefix":"10.14778","volume":"3","author":[{"given":"Yingyi","family":"Bu","sequence":"first","affiliation":[{"name":"University of Washington, Seattle, WA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bill","family":"Howe","sequence":"additional","affiliation":[{"name":"University of Washington, Seattle, WA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Magdalena","family":"Balazinska","sequence":"additional","affiliation":[{"name":"University of Washington, Seattle, WA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael D.","family":"Ernst","sequence":"additional","affiliation":[{"name":"University of Washington, Seattle, WA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2010,9]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"http:\/\/www.nsf.gov\/pubs\/2008\/nsf08560\/nsf08560.htm. Accessed July 7 2010.  http:\/\/www.nsf.gov\/pubs\/2008\/nsf08560\/nsf08560.htm. Accessed July 7 2010."},{"issue":"1","key":"e_1_2_1_2_1","first-page":"922","article-title":"An architectural hybrid of MapReduce and DBMS technologies for analytical workloads","volume":"2","author":"Abouzeid Azza","year":"2009","unstructured":"Azza Abouzeid , Kamil Bajda-Pawlikowski , Daniel J. Abadi , Alexander Rasin , and Avi Silberschatz . HadoopDB : An architectural hybrid of MapReduce and DBMS technologies for analytical workloads . VLDB , 2 ( 1 ): 922 -- 933 , 2009 . Azza Abouzeid, Kamil Bajda-Pawlikowski, Daniel J. Abadi, Alexander Rasin, and Avi Silberschatz. HadoopDB: An architectural hybrid of MapReduce and DBMS technologies for analytical workloads. VLDB, 2(1):922--933, 2009.","journal-title":"VLDB"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/16894.16859"},{"key":"e_1_2_1_4_1","first-page":"137","volume-title":"OSDI","author":"Dean Jeffrey","year":"2004","unstructured":"Jeffrey Dean and Sanjay Ghemawat . MapReduce : Simplified data processing on large clusters . In OSDI , pages 137 -- 150 , 2004 . Jeffrey Dean and Sanjay Ghemawat. MapReduce: Simplified data processing on large clusters. In OSDI, pages 137--150, 2004."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/129888.129894"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/eScience.2008.59"},{"key":"e_1_2_1_7_1","volume-title":"http:\/\/hadoop.apache.org\/. Accessed","year":"2010","unstructured":"Hadoop. http:\/\/hadoop.apache.org\/. Accessed July 7, 2010 . Hadoop. http:\/\/hadoop.apache.org\/. Accessed July 7, 2010."},{"key":"e_1_2_1_8_1","volume-title":"http:\/\/hadoop.apache.org\/common\/docs\/current\/hdfs_design.html. Accessed","year":"2010","unstructured":"Hdfs. http:\/\/hadoop.apache.org\/common\/docs\/current\/hdfs_design.html. Accessed July 7, 2010 . Hdfs. http:\/\/hadoop.apache.org\/common\/docs\/current\/hdfs_design.html. Accessed July 7, 2010."},{"key":"e_1_2_1_9_1","volume-title":"http:\/\/hadoop.apache.org\/hive\/. Accessed","year":"2010","unstructured":"Hive. http:\/\/hadoop.apache.org\/hive\/. Accessed July 7, 2010 . Hive. http:\/\/hadoop.apache.org\/hive\/. Accessed July 7, 2010."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1272996.1273005"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/324133.324140"},{"key":"e_1_2_1_12_1","volume-title":"http:\/\/lucene.apache.org\/mahout\/. Accessed","year":"2010","unstructured":"Mahout. http:\/\/lucene.apache.org\/mahout\/. Accessed July 7, 2010 . Mahout. http:\/\/lucene.apache.org\/mahout\/. Accessed July 7, 2010."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807184"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376726"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559865"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.368511"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/1920841.1920881","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:53:59Z","timestamp":1672228439000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/1920841.1920881"}},"subtitle":["efficient iterative data processing on large clusters"],"short-title":[],"issued":{"date-parts":[[2010,9]]},"references-count":16,"journal-issue":{"issue":"1-2","published-print":{"date-parts":[[2010,9]]}},"alternative-id":["10.14778\/1920841.1920881"],"URL":"https:\/\/doi.org\/10.14778\/1920841.1920881","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2010,9]]}}}