{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T21:00:19Z","timestamp":1773090019712,"version":"3.50.1"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2018,2,28]],"date-time":"2018-02-28T00:00:00Z","timestamp":1519776000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Outstanding Youth Foundation of Hubei Province","award":["2016CFA032"],"award-info":[{"award-number":["2016CFA032"]}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["61772218, 61433019, U1435217"],"award-info":[{"award-number":["61772218, 61433019, U1435217"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2017YFC0803700"],"award-info":[{"award-number":["2017YFC0803700"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Comput. Syst."],"published-print":{"date-parts":[[2018,2,28]]},"abstract":"<jats:p>In-memory caching of intermediate data and active combining of data in shuffle buffers have been shown to be very effective in minimizing the recomputation and I\/O cost in big data processing systems such as Spark and Flink. However, it has also been widely reported that these techniques would create a large amount of long-living data objects in the heap. These generated objects may quickly saturate the garbage collector, especially when handling a large dataset, and hence, limit the scalability of the system. To eliminate this problem, we propose a lifetime-based memory management framework, which, by automatically analyzing the user-defined functions and data types, obtains the expected lifetime of the data objects and then allocates and releases memory space accordingly to minimize the garbage collection overhead. In particular, we present Deca,&lt;sup;&gt;1&lt;\/sup;&gt; a concrete implementation of our proposal on top of Spark, which transparently decomposes and groups objects with similar lifetimes into byte arrays and releases their space altogether when their lifetimes come to an end. When systems are processing very large data, Deca also provides field-oriented memory pages to ensure high compression efficiency. Extensive experimental studies using both synthetic and real datasets show that, in comparing to Spark, Deca is able to (1) reduce the garbage collection time by up to 99.9%, (2) reduce the memory consumption by up to 46.6% and the storage space by 23.4%, (3) achieve 1.2\u00d7 to 22.7\u00d7 speedup in terms of execution time in cases without data spilling and 16\u00d7 to 41.6\u00d7 speedup in cases with data spilling, and (4) provide similar performance compared to domain-specific systems.<\/jats:p>","DOI":"10.1145\/3310361","type":"journal-article","created":{"date-parts":[[2019,3,14]],"date-time":"2019-03-14T17:11:58Z","timestamp":1552583518000},"page":"1-47","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Deca"],"prefix":"10.1145","volume":"36","author":[{"given":"Xuanhua","family":"Shi","sequence":"first","affiliation":[{"name":"Huazhong University of Science and Technology, WuHan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhixiang","family":"Ke","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, WuHan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongluan","family":"Zhou","sequence":"additional","affiliation":[{"name":"University of Copenhagen, Copenhagen, Denmark"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hai","family":"Jin","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, WuHan China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lu","family":"Lu","sequence":"additional","affiliation":[{"name":"Alibaba Group, HangZhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ligang","family":"He","sequence":"additional","affiliation":[{"name":"University of Warwick, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhenyu","family":"Hu","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fei","family":"Wang","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,3,14]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/320384.320418"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/1924943.1924962"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1740390.1740400"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2742797"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1150402.1150412"},{"key":"e_1_2_1_6_1","unstructured":"Benchmark. 2014. Big Data Benchmark. Retrieved from http:\/\/tinyurl.com\/qg93r43.  Benchmark. 2014. Big Data Benchmark. Retrieved from http:\/\/tinyurl.com\/qg93r43."},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 26th International Conference on Software Engineering (ICSE\u201904)","author":"Blackburn Stephen M.","unstructured":"Stephen M. Blackburn , Perry Cheng , and Kathryn S . McKinley. 2004. Oil and water? High performance garbage collection in Java with MMTk . In Proceedings of the 26th International Conference on Software Engineering (ICSE\u201904) . IEEE Computer Society, Washington, DC, 137--146. Stephen M. Blackburn, Perry Cheng, and Kathryn S. McKinley. 2004. Oil and water? High performance garbage collection in Java with MMTk. In Proceedings of the 26th International Conference on Software Engineering (ICSE\u201904). IEEE Computer Society, Washington, DC, 137--146."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/988672.988752"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3092255.3092272"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491894.2466485"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/304065.304099"},{"key":"e_1_2_1_12_1","unstructured":"Cassandra. 2010. Cassandra Garbage Collection Tuning. Retrieved from http:\/\/tinyurl.com\/5u58mzc.  Cassandra. 2010. Cassandra Garbage Collection Tuning. Retrieved from http:\/\/tinyurl.com\/5u58mzc."},{"key":"e_1_2_1_13_1","unstructured":"Databricks. 2015. Tuning Java Garbage Collection for Spark Applications. Retrieved from http:\/\/tinyurl.com\/pd8kkau.  Databricks. 2015. Tuning Java Garbage Collection for Spark Applications. Retrieved from http:\/\/tinyurl.com\/pd8kkau."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1851476.1851593"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2815400.2815407"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2694344.2694361"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 15th USENIX Conference on Hot Topics in Operating Systems (HOTOS\u201915)","author":"Gog Ionel","year":"2015","unstructured":"Ionel Gog , Jana Giceva , Malte Schwarzkopf , Kapil Vaswani , Dimitrios Vytiniotis , Ganesan Ramalingan , Derek Murray , Steven Hand , and Michael Isard . 2015 . Broom: Sweeping out garbage collection from big data systems . In Proceedings of the 15th USENIX Conference on Hot Topics in Operating Systems (HOTOS\u201915) . USENIX Association, Berkeley, CA, 2--2. Ionel Gog, Jana Giceva, Malte Schwarzkopf, Kapil Vaswani, Dimitrios Vytiniotis, Ganesan Ramalingan, Derek Murray, Steven Hand, and Michael Isard. 2015. Broom: Sweeping out garbage collection from big data systems. In Proceedings of the 15th USENIX Conference on Hot Topics in Operating Systems (HOTOS\u201915). USENIX Association, Berkeley, CA, 2--2."},{"key":"e_1_2_1_18_1","volume-title":"Go GC: Prioritizing Low Latency and Simplicity.","author":"GC.","year":"2015","unstructured":"Go GC. 2015 . Go GC: Prioritizing Low Latency and Simplicity. Retrieved from https:\/\/blog.golang.org\/go15gc. GoGC. 2015. Go GC: Prioritizing Low Latency and Simplicity. Retrieved from https:\/\/blog.golang.org\/go15gc."},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201914)","author":"Gonzalez Joseph E.","year":"2014","unstructured":"Joseph E. Gonzalez , Reynold S. Xin , Ankur Dave , Daniel Crankshaw , Michael J. Franklin , and Ion Stoica . 2014 . GraphX: Graph processing in a distributed dataflow framework . In Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201914) . USENIX Association, Berkeley, CA, 599--613. Joseph E. Gonzalez, Reynold S. Xin, Ankur Dave, Daniel Crankshaw, Michael J. Franklin, and Ion Stoica. 2014. GraphX: Graph processing in a distributed dataflow framework. In Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201914). USENIX Association, Berkeley, CA, 599--613."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/2387880.2387893"},{"key":"e_1_2_1_21_1","unstructured":"HBaseGC. 2016. Tuning Java Garbage Collection for HBase. Retrieved from http:\/\/tinyurl.com\/j5hsd3x.  HBaseGC. 2016. Tuning Java Garbage Collection for HBase. Retrieved from http:\/\/tinyurl.com\/j5hsd3x."},{"key":"e_1_2_1_22_1","unstructured":"HiBench. 2016. HiBench Suite. Retrieved from http:\/\/tinyurl.com\/cns79vt.  HiBench. 2016. HiBench Suite. Retrieved from http:\/\/tinyurl.com\/cns79vt."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1629575.1629601"},{"key":"e_1_2_1_24_1","volume-title":"The Garbage Collection Handbook: The Art of Automatic Memory Management","author":"Jones Richard","unstructured":"Richard Jones , Antony Hosking , and Eliot Moss . 2011. The Garbage Collection Handbook: The Art of Automatic Memory Management . Chapman and Hall\/CRC. Richard Jones, Antony Hosking, and Eliot Moss. 2011. The Garbage Collection Handbook: The Art of Automatic Memory Management. Chapman and Hall\/CRC."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.5555\/1765931.1765948"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/1989323.1989426"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/358141.358147"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994513"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2872362.2872386"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766462.2767755"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 15th USENIX Conference on Hot Topics in Operating Systems (HOTOS\u201915)","author":"McSherry Frank","unstructured":"Frank McSherry , Michael Isard , and Derek G. Murray . 2015. Scalability&excl; but at what cost? In Proceedings of the 15th USENIX Conference on Hot Topics in Operating Systems (HOTOS\u201915) . USENIX Association, Berkeley, CA, 14--14. Frank McSherry, Michael Isard, and Derek G. Murray. 2015. Scalability&excl; but at what cost? In Proceedings of the 15th USENIX Conference on Hot Topics in Operating Systems (HOTOS\u201915). USENIX Association, Berkeley, CA, 14--14."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2509136.2509547"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1993498.1993513"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201916)","author":"Nguyen Khanh","year":"2016","unstructured":"Khanh Nguyen , Lu Fang , Guoqing Xu , Brian Demsky , Shan Lu , Sanazsadat Alamian , and Onur Mutlu . 2016 . Yak: A high-performance big-data-friendly garbage collector . In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201916) . USENIX Association, Berkeley, CA, 349--365. Khanh Nguyen, Lu Fang, Guoqing Xu, Brian Demsky, Shan Lu, Sanazsadat Alamian, and Onur Mutlu. 2016. Yak: A high-performance big-data-friendly garbage collector. In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201916). USENIX Association, Berkeley, CA, 349--365."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2694344.2694345"},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the 9th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201910)","author":"Power Russell","year":"2010","unstructured":"Russell Power and Jinyang Li . 2010 . Piccolo: Building fast, distributed programs with partitioned tables . In Proceedings of the 9th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201910) . USENIX Association, Berkeley, CA, 293--306. Russell Power and Jinyang Li. 2010. Piccolo: Building fast, distributed programs with partitioned tables. In Proceedings of the 9th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201910). USENIX Association, Berkeley, CA, 293--306."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2815400.2815418"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/604264.604282"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/2831360.2831365"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.14778\/2367502.2367513"},{"key":"e_1_2_1_41_1","unstructured":"SNAP. 2016. SNAP Dataset Collection. Retrieved from https:\/\/snap.stanford.edu\/data\/index.html.  SNAP. 2016. SNAP Dataset Collection. Retrieved from https:\/\/snap.stanford.edu\/data\/index.html."},{"key":"e_1_2_1_42_1","unstructured":"Soot. 2016. Soot Framework. Retrieved from http:\/\/sable.github.io\/soot\/.  Soot. 2016. Soot Framework. Retrieved from http:\/\/sable.github.io\/soot\/."},{"key":"e_1_2_1_43_1","unstructured":"SparkGC. 2016. Spark Garbage Collection Tuning. Retrieved from http:\/\/tinyurl.com\/hzf3gqm.  SparkGC. 2016. Spark Garbage Collection Tuning. Retrieved from http:\/\/tinyurl.com\/hzf3gqm."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1006\/inco.1996.2613"},{"key":"e_1_2_1_45_1","unstructured":"Tungsten. 2015. Project Tungsten of Spark. Retrieved from http:\/\/tinyurl.com\/mzw7hew.  Tungsten. 2015. Project Tungsten of Spark. Retrieved from http:\/\/tinyurl.com\/mzw7hew."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/2523616.2523633"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1002\/1096-9128(200005)12:7<519::AID-CPE497>3.0.CO;2-M"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/1755913.1755940"},{"key":"e_1_2_1_49_1","volume-title":"Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation (NSDI\u201912)","author":"Zaharia Matei","year":"2012","unstructured":"Matei Zaharia , Mosharaf Chowdhury , Tathagata Das , Ankur Dave , Justin Ma , Murphy McCauley , Michael J. Franklin , Scott Shenker , and Ion Stoica . 2012 . Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing . In Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation (NSDI\u201912) . USENIX Association, Berkeley, CA, 2--2. Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauley, Michael J. Franklin, Scott Shenker, and Ion Stoica. 2012. Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. In Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation (NSDI\u201912). USENIX Association, Berkeley, CA, 2--2."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.5555\/1855741.1855744"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2015.2427795"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/2038916.2038929"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-015-5371-1"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1977.1055714"}],"container-title":["ACM Transactions on Computer Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3310361","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3310361","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:53:37Z","timestamp":1750204417000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3310361"}},"subtitle":["A Garbage Collection Optimizer for In-Memory Data Processing"],"short-title":[],"issued":{"date-parts":[[2018,2,28]]},"references-count":54,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2018,2,28]]}},"alternative-id":["10.1145\/3310361"],"URL":"https:\/\/doi.org\/10.1145\/3310361","relation":{},"ISSN":["0734-2071","1557-7333"],"issn-type":[{"value":"0734-2071","type":"print"},{"value":"1557-7333","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,2,28]]},"assertion":[{"value":"2017-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-03-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}