{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:43:17Z","timestamp":1787017397986,"version":"build-2736575974"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,11,13]],"date-time":"2023-11-13T00:00:00Z","timestamp":1699833600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Database Syst."],"published-print":{"date-parts":[[2023,12,31]]},"abstract":"<jats:p>The use of storage tiering is becoming popular in data-intensive compute clusters due to the recent advancements in storage technologies. The Hadoop Distributed File System, for example, now supports storing data in memory, SSDs, and HDDs, while OctopusFS and hatS offer fine-grained storage tiering solutions. However, current big data platforms (such as Hadoop and Spark) are not exploiting the presence of storage tiers and the opportunities they present for performance optimizations. Specifically, schedulers and prefetchers will make decisions only based on data locality information and completely ignore the fact that local data are now stored on a variety of storage media with different performance characteristics. This article presents Trident, a scheduling and prefetching framework that is designed to make task assignment, resource scheduling, and prefetching decisions based on both locality and storage tier information. Trident formulates task scheduling as a minimum cost maximum matching problem in a bipartite graph and utilizes two novel pruning algorithms for bounding the size of the graph, while still guaranteeing optimality. In addition, Trident extends YARN\u2019s resource request model and proposes a new storage-tier-aware resource scheduling algorithm. Finally, Trident includes a cost-based data prefetching approach that coordinates with the schedulers for optimizing prefetching operations. Trident is implemented in both Spark and Hadoop and evaluated extensively using a realistic workload derived from Facebook traces as well as an industry-validated benchmark, demonstrating significant benefits in terms of application performance and cluster efficiency.<\/jats:p>","DOI":"10.1145\/3625389","type":"journal-article","created":{"date-parts":[[2023,9,24]],"date-time":"2023-09-24T04:24:16Z","timestamp":1695529456000},"page":"1-40","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Cost-based Data Prefetching and Scheduling in Big Data Platforms over Tiered Storage Systems"],"prefix":"10.1145","volume":"48","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8717-1691","authenticated-orcid":false,"given":"Herodotos","family":"Herodotou","sequence":"first","affiliation":[{"name":"Cyprus University of Technology, Cyprus"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1489-807X","authenticated-orcid":false,"given":"Elena","family":"Kakoulli","sequence":"additional","affiliation":[{"name":"Neapolis University Pafos, Cyprus and Cyprus University of Technology, Cyprus"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,11,13]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"crossref","first-page":"159","DOI":"10.1109\/CLUSTER.2011.26","volume-title":"Proceedings of the 2011 IEEE International Conference on Cluster Computing (CLUSTER)","author":"Abad Cristina L.","year":"2011","unstructured":"Cristina L. Abad, Yi Lu, and Roy H. Campbell. 2011. DARE: Adaptive data replication for efficient cluster scheduling. In Proceedings of the 2011 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, 159\u2013168."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS.2018.00034"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/2189750.2150984"},{"key":"e_1_3_1_5_2","unstructured":"Alluxio 2023. Alluxio: Data Orchestration for the Cloud . Retrieved September 18 2023 from http:\/\/www.alluxio.org\/"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1145\/1966445.1966472","volume-title":"Proceedings of the 6th European Conference on Computer Systems (EuroSys\u201911)","author":"Ananthanarayanan Ganesh","year":"2011","unstructured":"Ganesh Ananthanarayanan, Sameer Agarwal, Srikanth Kandula, Albert Greenberg, Ion Stoica, Duke Harlan, and Ed Harris. 2011. Scarlett: Coping with skewed popularity content in MapReduce clusters. In Proceedings of the 6th European Conference on Computer Systems (EuroSys\u201911). ACM, 287\u2013300."},{"key":"e_1_3_1_7_2","first-page":"12","volume-title":"Proceedings of the 13th Workshop on Hot Topics in Operating Systems (HotOS\u201911)","author":"Ananthanarayanan Ganesh","year":"2011","unstructured":"Ganesh Ananthanarayanan, Ali Ghodsi, Scott Shenker, and Ion Stoica. 2011. Disk-locality in datacenter computing considered irrelevant. In Proceedings of the 13th Workshop on Hot Topics in Operating Systems (HotOS\u201911). USENIX, 12\u201317."},{"key":"e_1_3_1_8_2","first-page":"267","volume-title":"Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201912)","author":"Ananthanarayanan Ganesh","year":"2012","unstructured":"Ganesh Ananthanarayanan, Ali Ghodsi, Andrew Warfield, Dhruba Borthakur, Srikanth Kandula, Scott Shenker, and Ion Stoica. 2012. PACMan: Coordinated memory caching for parallel jobs. In Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201912). USENIX, 267\u2013280."},{"key":"e_1_3_1_9_2","unstructured":"Apache Hadoop 2023. Apache Hadoop. Retrieved September 18 2023 from https:\/\/hadoop.apache.org"},{"key":"e_1_3_1_10_2","unstructured":"Apache Oozie 2023. Apache Oozie Workflow Scheduler for Hadoop. Retrieved September 18 2023 from https:\/\/oozie.apache.org\/"},{"key":"e_1_3_1_11_2","unstructured":"Apache Spark 2023. Apache Spark. Retrieved September 18 2023 from https:\/\/spark.apache.org"},{"key":"e_1_3_1_12_2","unstructured":"Azkaban 2023. Azkaban: Open-source Workflow Manager. Retrieved September 18 2023 from https:\/\/azkaban.github.io\/"},{"key":"e_1_3_1_13_2","first-page":"835","volume-title":"Proceedings of the 31st International Conference on Advanced Information Networking and Applications (AINA\u201917)","author":"Chen Chien-Hung","year":"2017","unstructured":"Chien-Hung Chen, Ting-Yuan Hsia, Yennun Huang, and Sy-Yen Kuo. 2017. Scheduling-aware data prefetching for data processing services in cloud. In Proceedings of the 31st International Conference on Advanced Information Networking and Applications (AINA\u201917). IEEE, 835\u2013842."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2892957"},{"key":"e_1_3_1_15_2","first-page":"2736","volume-title":"Proceedings of the 10th IEEE International Conference on Computer and Information Technology (ICCIT\u201910)","author":"Chen Quan","year":"2010","unstructured":"Quan Chen, D. Zhang, M. Guo, Q. Deng, and S. Guo. 2010. SAMR: A self-adaptive mapreduce scheduling algorithm in heterogeneous environment. In Proceedings of the 10th IEEE International Conference on Computer and Information Technology (ICCIT\u201910). IEEE, 2736\u20132743."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.14778\/2367502.2367519"},{"key":"e_1_3_1_17_2","first-page":"390","volume-title":"Proceedings of the IEEE International Symposium on Modeling, Analysis & Simulation of Computer and Telecommunication Systems (MASCOTS\u201911)","author":"Chen Yanpei","year":"2011","unstructured":"Yanpei Chen, Archana Ganapathi, Rean Griffith, and Randy Katz. 2011. The case for evaluating MapReduce performance using workload suites. In Proceedings of the IEEE International Symposium on Modeling, Analysis & Simulation of Computer and Telecommunication Systems (MASCOTS\u201911). IEEE, 390\u2013399."},{"key":"e_1_3_1_18_2","first-page":"97","volume-title":"Proceedings of the 15th IEEE International Conference on Cluster Computing (CLUSTER\u201914)","author":"Cheng Dazhao","year":"2014","unstructured":"Dazhao Cheng, Jia Rao, Yanfei Guo, and Xiaobo Zhou. 2014. Improving MapReduce performance in heterogeneous environments with adaptive task tuning. In Proceedings of the 15th IEEE International Conference on Cluster Computing (CLUSTER\u201914). ACM, 97\u2013108."},{"key":"e_1_3_1_19_2","volume-title":"Introduction to Algorithms","author":"Cormen Thomas H.","year":"2009","unstructured":"Thomas H. Cormen, Charles E. Leiserson, Ronald L. Rivest, and Clifford Stein. 2009. Introduction to Algorithms. MIT Press."},{"key":"e_1_3_1_20_2","first-page":"1","volume-title":"Proceedings of the 8th USENIX Workshop on Hot Topics in Storage and File Systems (HotStorage\u201916)","author":"Deslauriers Francis","year":"2016","unstructured":"Francis Deslauriers, Peter McCormick, George Amvrosiadis, Ashvin Goel, and Angela Demke Brown. 2016. Quartet: Harmonizing task scheduling and caching for cluster computing. In Proceedings of the 8th USENIX Workshop on Hot Topics in Storage and File Systems (HotStorage\u201916). USENIX, 1\u20135."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/2529989"},{"key":"e_1_3_1_22_2","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1145\/2987550.2987553","volume-title":"Proceedings of the 7th ACM Symposium on Cloud Computing (SoCC\u201916)","author":"Floratou Avrilia","year":"2016","unstructured":"Avrilia Floratou, Nimrod Megiddo, Navneet Potti, Fatma \u00d6zcan, Uday Kale, and Jan Schmitz-Hermes. 2016. Adaptive caching in big SQL using the HDFS cache. In Proceedings of the 7th ACM Symposium on Cloud Computing (SoCC\u201916). ACM, 321\u2013333."},{"key":"e_1_3_1_23_2","first-page":"61","volume-title":"Proceedings of the USENIX Annual Technical Conference (ATC\u201913)","author":"Gandhi Rohan","year":"2013","unstructured":"Rohan Gandhi, Di Xie, and Y. Charlie Hu. 2013. PIKACHU: How to rebalance load in optimizing MapReduce on heterogeneous clusters. In Proceedings of the USENIX Annual Technical Conference (ATC\u201913). USENIX, 61\u201366."},{"key":"e_1_3_1_24_2","first-page":"165","volume-title":"Proceedings of the 9th International Conference on Advanced Computing (ICoAC\u201917)","author":"Govindarajan Kannan","year":"2017","unstructured":"Kannan Govindarajan, Supun Kamburugamuve, Pulasthi Wickramasinghe, Vibhatha Abeykoon, and Geoffrey Fox. 2017. Task scheduling in big data-review, research challenges, and prospects. In Proceedings of the 9th International Conference on Advanced Computing (ICoAC\u201917). IEEE, 165\u2013173."},{"key":"e_1_3_1_25_2","unstructured":"GridGain 2023. GridGain In-Memory Data Platform for High-Performance Applications . Retrieved September 18 2023 from http:\/\/www.gridgain.com\/"},{"issue":"5","key":"e_1_3_1_26_2","doi-asserted-by":"crossref","first-page":"71","DOI":"10.14257\/ijgdc.2013.6.5.07","article-title":"Improving MapReduce performance by data prefetching in heterogeneous or shared environments","volume":"6","author":"Gu Tao","year":"2013","unstructured":"Tao Gu, Chuang Zuo, Qun Liao, Yulu Yang, and Tao Li. 2013. Improving MapReduce performance by data prefetching in heterogeneous or shared environments. Int. J. Grid Distrib. Comput. 6, 5 (2013), 71\u201382.","journal-title":"Int. J. Grid Distrib. Comput."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/2934872.2934908"},{"key":"e_1_3_1_28_2","unstructured":"Hadoop: Fair Scheduler 2023. Hadoop: Fair Scheduler. Retrieved September 18 2023 from https:\/\/hadoop.apache.org\/docs\/current\/hadoop-yarn\/hadoop-yarn-site\/FairScheduler.html"},{"key":"e_1_3_1_29_2","unstructured":"HDFS 2023. Centralized Cache Management in HDFS. Retrieved from https:\/\/hadoop.apache.org\/docs\/stable\/hadoop-project-dist\/hadoop-hdfs\/CentralizedCacheManagement.html"},{"key":"e_1_3_1_30_2","unstructured":"HDFS 2023. HDFS Archival Storage SSD & Memory. Retrieved September 18 2023 from https:\/\/hadoop.apache.org\/docs\/current\/hadoop-project-dist\/hadoop-hdfs\/ArchivalStorage.html"},{"key":"e_1_3_1_31_2","first-page":"133","volume-title":"Proceedings of the IEEE 35th International Conference on Data Engineering Workshops (ICDEW\u201919)","author":"Herodotou Herodotos","year":"2019","unstructured":"Herodotos Herodotou. 2019. AutoCache: Employing machine learning to automate caching in distributed file systems. In Proceedings of the IEEE 35th International Conference on Data Engineering Workshops (ICDEW\u201919). IEEE, 133\u2013139."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.14778\/3402707.3402746"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3381027"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.14778\/3357377.3357381"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.14778\/3461535.3461545"},{"key":"e_1_3_1_36_2","unstructured":"HiBench 2023. HiBench Suite. Retrieved September 18 2023 from https:\/\/github.com\/intel-hadoop\/HiBench"},{"key":"e_1_3_1_37_2","first-page":"295","volume-title":"Proceedings of the 8th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201911)","author":"Hindman Benjamin","year":"2011","unstructured":"Benjamin Hindman, Andy Konwinski, Matei Zaharia, Ali Ghodsi, Anthony D. Joseph, Randy Katz, Scott Shenker, and Ion Stoica. 2011. Mesos: A platform for fine-grained resource sharing in the data center. In Proceedings of the 8th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201911). USENIX, 295\u2013308."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-19294-4_9"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/1629575.1629601"},{"key":"e_1_3_1_40_2","first-page":"1","volume-title":"Proceedings of the 35th IEEE International Conference on Computer Communications (INFOCOM\u201916)","author":"Jiang Jingjie","year":"2016","unstructured":"Jingjie Jiang, Shiyao Ma, Bo Li, and Baochun Li. 2016. Symbiosis: Network-aware task scheduling in data-parallel frameworks. In Proceedings of the 35th IEEE International Conference on Computer Communications (INFOCOM\u201916). IEEE, 1\u20139."},{"key":"e_1_3_1_41_2","first-page":"65","volume-title":"Proceedings of the ACM International Conference on Management of Data (SIGMOD\u201917)","author":"Kakoulli Elena","year":"2017","unstructured":"Elena Kakoulli and Herodotos Herodotou. 2017. OctopusFS: A distributed file system with tiered storage management. In Proceedings of the ACM International Conference on Management of Data (SIGMOD\u201917). ACM, 65\u201378."},{"key":"e_1_3_1_42_2","first-page":"502","volume-title":"Proceedings of the 14th IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid\u201914)","author":"Krish K. R.","year":"2014","unstructured":"K. R. Krish, Ali Anwar, and Ali R. Butt. 2014. hatS: A heterogeneity-aware tiered storage for hadoop. In Proceedings of the 14th IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid\u201914). IEEE, 502\u2013511."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2019.02.007"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2678505"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2015.2442980"},{"issue":"9","key":"e_1_3_1_46_2","first-page":"308","article-title":"A comprehensive view of hadoop mapreduce scheduling algorithms","volume":"2","author":"Pakize Seyed Reza","year":"2014","unstructured":"Seyed Reza Pakize. 2014. A comprehensive view of hadoop mapreduce scheduling algorithms. Int. J. Comput. Netw. Commun. Secur. 2, 9 (2014), 308\u2013317.","journal-title":"Int. J. Comput. Netw. Commun. Secur."},{"key":"e_1_3_1_47_2","first-page":"1","volume-title":"Proceedings of the 24th IEEE International Conference on Parallel and Distributed Systems (ICPADS\u201918)","author":"Pan Fengfeng","year":"2018","unstructured":"Fengfeng Pan, Jin Xiong, Yijie Shen, Tianshi Wang, and Dejun Jiang. 2018. H-scheduler: Storage-aware task scheduling for heterogeneous-storage spark clusters. In Proceedings of the 24th IEEE International Conference on Parallel and Distributed Systems (ICPADS\u201918). IEEE, 1\u20139."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3465998.3466003"},{"issue":"1","key":"e_1_3_1_49_2","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1109\/TCC.2015.2396056","article-title":"HFSP: Bringing size-based scheduling to hadoop","volume":"5","author":"Pastorelli Mario","year":"2015","unstructured":"Mario Pastorelli, Damiano Carra, Matteo Dell\u2019Amico, and Pietro Michiardi. 2015. HFSP: Bringing size-based scheduling to hadoop. IEEE Trans. Cloud Comput. 5, 1 (2015), 43\u201356.","journal-title":"IEEE Trans. Cloud Comput."},{"key":"e_1_3_1_50_2","first-page":"50","volume-title":"Proceedings of the Third International Conference on Services in Emerging Markets","author":"Raj Aparna","year":"2012","unstructured":"Aparna Raj, Kamaldeep Kaur, Uddipan Dutta, V. Venkat Sandeep, and Shrisha Rao. 2012. Enhancement of hadoop clusters with virtualization using the capacity scheduler. In Proceedings of the Third International Conference on Services in Emerging Markets. IEEE, 50\u201357."},{"key":"e_1_3_1_51_2","first-page":"1","volume-title":"Proceedings of the 26th International Conference on Massive Storage Systems and Technology (MSST\u201910)","author":"Shvachko Konstantin","year":"2010","unstructured":"Konstantin Shvachko, Hairong Kuang, Sanjay Radia, and Robert Chansler. 2010. The hadoop distributed file system. In Proceedings of the 26th International Conference on Massive Storage Systems and Technology (MSST\u201910). IEEE, 1\u201310."},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2017.09.001"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2015.04.039"},{"key":"e_1_3_1_54_2","first-page":"82","volume-title":"International Conference on Algorithms and Architectures for Parallel Processing","author":"Sun Mingming","year":"2014","unstructured":"Mingming Sun, Hang Zhuang, Xuehai Zhou, Kun Lu, and Changlong Li. 2014. HPSO: Prefetching based scheduling to improve data locality for MapReduce clusters. In International Conference on Algorithms and Architectures for Parallel Processing. Springer, 82\u201395."},{"key":"e_1_3_1_55_2","first-page":"148","volume-title":"Proceedings of the 18th IEEE International Conference on Parallel and Distributed Systems (ICPADS\u201912)","author":"Sun Xiaoyu","year":"2012","unstructured":"Xiaoyu Sun, C. He, and Ying Lu. 2012. ESAMR: An enhanced self-adaptive MapReduce scheduling algorithm. In Proceedings of the 18th IEEE International Conference on Parallel and Distributed Systems (ICPADS\u201912). IEEE, 148\u2013155."},{"key":"e_1_3_1_56_2","first-page":"513","volume-title":"Proceedings of the 12th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201915)","author":"Suresh Lalith","year":"2015","unstructured":"Lalith Suresh, Marco Canini, Stefan Schmid, and Anja Feldmann. 2015. C3: Cutting tail latency in cloud data stores via adaptive replica selection. In Proceedings of the 12th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201915). USENIX Association, Oakland, CA, 513\u2013527."},{"key":"e_1_3_1_57_2","unstructured":"SWIM 2016. SWIM: Statistical Workload Injector for MapReduce . Retrieved September 18 2023 from https:\/\/github.com\/SWIMProjectUCB\/SWIM\/wiki"},{"key":"e_1_3_1_58_2","first-page":"1618","volume-title":"Proceedings of the 32nd IEEE International Conference on Computer Communications (INFOCOM\u201913)","author":"Tan Jian","year":"2013","unstructured":"Jian Tan, Xiaoqiao Meng, and Li Zhang. 2013. Coupling task progress for MapReduce resource-aware scheduling. In Proceedings of the 32nd IEEE International Conference on Computer Communications (INFOCOM\u201913). IEEE, 1618\u20131626."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-014-1335-2"},{"key":"e_1_3_1_60_2","first-page":"1","volume-title":"Proceedings of the 4th ACM Symposium on Cloud Computing (SoCC\u201913)","author":"Vavilapalli Vinod Kumar","year":"2013","unstructured":"Vinod Kumar Vavilapalli, Arun C. Murthy, Chris Douglas, Sharad Agarwal, Mahadev Konar, Robert Evans, et\u00a0al. 2013. Apache hadoop YARN: Yet another resource negotiator. In Proceedings of the 4th ACM Symposium on Cloud Computing (SoCC\u201913). ACM, 1\u201316."},{"key":"e_1_3_1_61_2","first-page":"761","volume-title":"Proceedings of the 7th IEEE International Conference on Cloud Computing (CLOUD\u201914)","author":"Wang Jiayin","year":"2014","unstructured":"Jiayin Wang, Yi Yao, Ying Mao, Bo Sheng, and Ningfang Mi. 2014. Fresh: Fair and efficient slot configuration and scheduling for hadoop clusters. In Proceedings of the 7th IEEE International Conference on Cloud Computing (CLOUD\u201914). IEEE, 761\u2013768."},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/NAS.2017.8026881"},{"issue":"1","key":"e_1_3_1_63_2","doi-asserted-by":"crossref","first-page":"190","DOI":"10.1109\/TNET.2014.2362745","article-title":"Map task scheduling in mapreduce with data locality: Throughput and heavy-traffic optimality","volume":"24","author":"Wang Weina","year":"2014","unstructured":"Weina Wang, Kai Zhu, Lei Ying, Jian Tan, and Li Zhang. 2014. Map task scheduling in mapreduce with data locality: Throughput and heavy-traffic optimality. IEEE\/ACM Trans. Netw. 24, 1 (2014), 190\u2013203.","journal-title":"IEEE\/ACM Trans. Netw."},{"key":"e_1_3_1_64_2","first-page":"245","volume-title":"Proceedings of the IEEE International Conference on Cluster Computing (CLUSTER\u201918)","author":"Xu Luna","year":"2018","unstructured":"Luna Xu, A. Butt, Seung-Hwan Lim, and R. Kannan. 2018. A heterogeneity-aware task scheduler for spark. In Proceedings of the IEEE International Conference on Cluster Computing (CLUSTER\u201918). IEEE, 245\u2013256."},{"key":"e_1_3_1_65_2","first-page":"52","volume-title":"Proceedings of the IEEE International Conference on Cluster Computing (CLUSTER\u201915)","author":"Yu Ze","year":"2015","unstructured":"Ze Yu, Min Li, Xin Yang, Han Zhao, and Xiaolin Li. 2015. Taming non-local stragglers using efficient prefetching in MapReduce. In Proceedings of the IEEE International Conference on Cluster Computing (CLUSTER\u201915). IEEE, 52\u201361."},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1145\/1755913.1755940"},{"key":"e_1_3_1_67_2","first-page":"29","volume-title":"Proceedings of the 8th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201908)","author":"Zaharia Matei","year":"2008","unstructured":"Matei Zaharia, Andy Konwinski, Anthony D. Joseph, Randy H. Katz, and Ion Stoica. 2008. Improving MapReduce performance in heterogeneous environments. In Proceedings of the 8th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201908). USENIX, 29\u201342."}],"container-title":["ACM Transactions on Database Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3625389","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3625389","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T12:36:33Z","timestamp":1750163793000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3625389"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,13]]},"references-count":66,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,12,31]]}},"alternative-id":["10.1145\/3625389"],"URL":"https:\/\/doi.org\/10.1145\/3625389","relation":{},"ISSN":["0362-5915","1557-4644"],"issn-type":[{"value":"0362-5915","type":"print"},{"value":"1557-4644","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,13]]},"assertion":[{"value":"2022-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-14","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}