{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,1]],"date-time":"2025-07-01T12:48:56Z","timestamp":1751374136283},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"5","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2019,1]]},"abstract":"<jats:p>Popular big data frameworks, ranging from Hadoop MapReduce to Spark, rely on garbage-collected languages, such as Java and Scala. Big data applications are especially sensitive to the effectiveness of garbage collection (i.e., GC), because they usually process a large volume of data objects that lead to heavy GC overhead. Lacking in-depth understanding of GC performance has impeded performance improvement in big data applications. In this paper, we conduct the first comprehensive evaluation on three popular garbage collectors, i.e., Parallel, CMS, and G1, using four representative Spark applications. By thoroughly investigating the correlation between these big data applications' memory usage patterns and the collectors' GC patterns, we obtain many findings about GC inefficiencies. We further propose empirical guidelines for application developers, and insightful optimization strategies for designing big-data-friendly garbage collectors.<\/jats:p>","DOI":"10.14778\/3303753.3303762","type":"journal-article","created":{"date-parts":[[2019,2,27]],"date-time":"2019-02-27T14:57:56Z","timestamp":1551279476000},"page":"570-583","source":"Crossref","is-referenced-by-count":15,"title":["An experimental evaluation of garbage collectors on big data applications"],"prefix":"10.14778","volume":"12","author":[{"given":"Lijie","family":"Xu","sequence":"first","affiliation":[{"name":"Institute of Software, Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tian","family":"Guo","sequence":"additional","affiliation":[{"name":"Worcester Polytechnic Institute"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wensheng","family":"Dou","sequence":"additional","affiliation":[{"name":"Institute of Software, Chinese Academy of Sciences and University of Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Software, Chinese Academy of Sciences and University of Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Wei","sequence":"additional","affiliation":[{"name":"Institute of Software, Chinese Academy of Sciences and University of Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,1]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Apache Hadoop. http:\/\/hadoop.apache.org\/.  Apache Hadoop. http:\/\/hadoop.apache.org\/."},{"key":"e_1_2_1_2_1","unstructured":"Apache Spark. http:\/\/spark.apache.org\/.  Apache Spark. http:\/\/spark.apache.org\/."},{"key":"e_1_2_1_3_1","unstructured":"Caching in Spark. https:\/\/spark.apache.org\/docs\/latest\/quick-start.html#caching.  Caching in Spark. https:\/\/spark.apache.org\/docs\/latest\/quick-start.html#caching."},{"key":"e_1_2_1_4_1","unstructured":"Early reclamation of large objects in G1. https:\/\/bugs.openjdk.java.net\/browse\/JDK-8027959.  Early reclamation of large objects in G1. https:\/\/bugs.openjdk.java.net\/browse\/JDK-8027959."},{"key":"e_1_2_1_5_1","unstructured":"Eclispe Memory Analyzer (MAT). https:\/\/www.eclipse.org\/mat\/.  Eclispe Memory Analyzer (MAT). https:\/\/www.eclipse.org\/mat\/."},{"key":"e_1_2_1_6_1","unstructured":"G1 native memory consumption. http:\/\/mail.openjdk.java.net\/pipermail\/hotspot-gc-use\/2017-February\/002609.html.  G1 native memory consumption. http:\/\/mail.openjdk.java.net\/pipermail\/hotspot-gc-use\/2017-February\/002609.html."},{"key":"e_1_2_1_7_1","unstructured":"Garbage-First Garbage Collector. https:\/\/docs.oracle.com\/javase\/8\/docs\/technotes\/guides\/vm\/gctuning\/g1_gc.html.  Garbage-First Garbage Collector. https:\/\/docs.oracle.com\/javase\/8\/docs\/technotes\/guides\/vm\/gctuning\/g1_gc.html."},{"key":"e_1_2_1_8_1","unstructured":"GC Ergonomics. https:\/\/docs.oracle.com\/javase\/8\/docs\/technotes\/guides\/vm\/gctuning\/ergonomics.html#ergonomics.  GC Ergonomics. https:\/\/docs.oracle.com\/javase\/8\/docs\/technotes\/guides\/vm\/gctuning\/ergonomics.html#ergonomics."},{"key":"e_1_2_1_9_1","unstructured":"HiBench: A big data benchmark suite. https:\/\/github.com\/intel-hadoop\/HiBench.  HiBench: A big data benchmark suite. https:\/\/github.com\/intel-hadoop\/HiBench."},{"key":"e_1_2_1_10_1","unstructured":"{JDK-8191565} Last-ditch Full GC should also move humongous objects. https:\/\/bugs.openjdk.java.net\/browse\/JDK-8191565.  {JDK-8191565} Last-ditch Full GC should also move humongous objects. https:\/\/bugs.openjdk.java.net\/browse\/JDK-8191565."},{"key":"e_1_2_1_11_1","unstructured":"JVM Generations. https:\/\/docs.oracle.com\/javase\/8\/docs\/technotes\/guides\/vm\/gctuning\/generations.html.  JVM Generations. https:\/\/docs.oracle.com\/javase\/8\/docs\/technotes\/guides\/vm\/gctuning\/generations.html."},{"key":"e_1_2_1_12_1","unstructured":"KDD Cup 2012 dataset. https:\/\/www.csie.ntu.edu.tw\/cjlin\/libsvmtools\/datasets\/binary.html.  KDD Cup 2012 dataset. https:\/\/www.csie.ntu.edu.tw\/cjlin\/libsvmtools\/datasets\/binary.html."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.14778\/3303753.3303762"},{"key":"e_1_2_1_14_1","unstructured":"Memory Management in the Java HotSpot Virtual Machine. http:\/\/www.oracle.com\/technetwork\/java\/javase\/memorymanagement-whitepaper-150215.pdf.  Memory Management in the Java HotSpot Virtual Machine. http:\/\/www.oracle.com\/technetwork\/java\/javase\/memorymanagement-whitepaper-150215.pdf."},{"key":"e_1_2_1_15_1","unstructured":"MLlib: Apache Spark's scalable machine learning library. https:\/\/spark.apache.org\/mllib\/.  MLlib: Apache Spark's scalable machine learning library. https:\/\/spark.apache.org\/mllib\/."},{"key":"e_1_2_1_16_1","unstructured":"OOM error caused by large array allocation in G1. http:\/\/mail.openjdk.java.net\/pipermail\/hotspot-gc-use\/2017-November\/002725.html.  OOM error caused by large array allocation in G1. http:\/\/mail.openjdk.java.net\/pipermail\/hotspot-gc-use\/2017-November\/002725.html."},{"key":"e_1_2_1_17_1","unstructured":"Project Tungsten: Bringing Apache Spark Closer to Bare Metal. https:\/\/databricks.com\/blog\/2015\/04\/28\/project-tungsten-bringing-spark-closer-to-bare-metal.html.  Project Tungsten: Bringing Apache Spark Closer to Bare Metal. https:\/\/databricks.com\/blog\/2015\/04\/28\/project-tungsten-bringing-spark-closer-to-bare-metal.html."},{"key":"e_1_2_1_18_1","unstructured":"Running Spark on YARN. https:\/\/spark.apache.org\/docs\/preview\/running-on-yarn.html.  Running Spark on YARN. https:\/\/spark.apache.org\/docs\/preview\/running-on-yarn.html."},{"key":"e_1_2_1_19_1","unstructured":"{SPARK-22713} OOM errors caused by the memory contention and memory leak in TaskMemoryManager. https:\/\/issues.apache.org\/jira\/browse\/SPARK-22713.  {SPARK-22713} OOM errors caused by the memory contention and memory leak in TaskMemoryManager. https:\/\/issues.apache.org\/jira\/browse\/SPARK-22713."},{"key":"e_1_2_1_20_1","unstructured":"Spark BigSQL Benchmark. https:\/\/amplab.cs.berkeley.edu\/benchmark\/.  Spark BigSQL Benchmark. https:\/\/amplab.cs.berkeley.edu\/benchmark\/."},{"key":"e_1_2_1_21_1","unstructured":"Spark executor GC taking long. http:\/\/stackoverflow.com\/questions\/38965787\/spark-executor-gc-taking-long.  Spark executor GC taking long. http:\/\/stackoverflow.com\/questions\/38965787\/spark-executor-gc-taking-long."},{"key":"e_1_2_1_22_1","unstructured":"SparkProfiler: Profiling Spark Applications for Performance Comparison and Diagnosis. https:\/\/github.com\/JerryLead\/SparkProfiler.  SparkProfiler: Profiling Spark Applications for Performance Comparison and Diagnosis. https:\/\/github.com\/JerryLead\/SparkProfiler."},{"key":"e_1_2_1_23_1","unstructured":"Twitter social graph. http:\/\/an.kaist.ac.kr\/traces\/WWW2010.html.  Twitter social graph. http:\/\/an.kaist.ac.kr\/traces\/WWW2010.html."},{"key":"e_1_2_1_24_1","unstructured":"Why does my JVM have access to less memory than -Xmx specifies? https:\/\/plumbr.io\/blog\/memory-leaks\/less-memory-than-xmx.  Why does my JVM have access to less memory than -Xmx specifies? https:\/\/plumbr.io\/blog\/memory-leaks\/less-memory-than-xmx."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2523616.2523625"},{"key":"e_1_2_1_26_1","first-page":"267","volume-title":"Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI)","author":"Ananthanarayanan G.","year":"2012"},{"key":"e_1_2_1_27_1","first-page":"265","volume-title":"Proceedings of the 9th USENIX Symposium on Operating Systems Design and Implementation, (OSDI)","author":"Ananthanarayanan G.","year":"2010"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824080"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2742797"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3135974.3135986"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3156818"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3092255.3092272"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491894.2466485"},{"key":"e_1_2_1_34_1","first-page":"29","volume-title":"Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI)","author":"Costa P.","year":"2012"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2815400.2815407"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2694344.2694361"},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the 15th Workshop on Hot Topics in Operating Systems (HotOS XV)","author":"Gog I.","year":"2015"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.14778\/3402707.3402746"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920903"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213840"},{"key":"e_1_2_1_41_1","first-page":"383","volume-title":"Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI)","author":"Lion D.","year":"2016"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994513"},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the 15th Workshop on Hot Topics in Operating Systems (HotOS XV)","author":"Maas M.","year":"2015"},{"key":"e_1_2_1_44_1","first-page":"349","volume-title":"Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI)","author":"Nguyen K.","year":"2016"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2694344.2694345"},{"key":"e_1_2_1_46_1","first-page":"293","volume-title":"Proceedings of the 12th USENIX Symposium on Networked Systems Design and Implementation (NSDI)","author":"Ousterhout K.","year":"2015"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559865"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2901318.2901319"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.14778\/2831360.2831365"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3190508.3190512"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE.2015.7381844"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2016.105"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/2892242.2892251"},{"key":"e_1_2_1_54_1","first-page":"15","volume-title":"Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI)","author":"Zaharia M.","year":"2012"},{"key":"e_1_2_1_55_1","first-page":"29","volume-title":"Proceedings of the 8th USENIX Symposium on Operating Systems Design and Implementation (OSDI)","author":"Zaharia M.","year":"2008"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3303753.3303762","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T09:28:41Z","timestamp":1672219721000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3303753.3303762"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,1]]},"references-count":55,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2019,1]]}},"alternative-id":["10.14778\/3303753.3303762"],"URL":"https:\/\/doi.org\/10.14778\/3303753.3303762","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2019,1]]}}}