{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T10:56:44Z","timestamp":1785927404034,"version":"3.56.0"},"reference-count":41,"publisher":"Springer Science and Business Media LLC","issue":"8","license":[{"start":{"date-parts":[[2024,1,29]],"date-time":"2024-01-29T00:00:00Z","timestamp":1706486400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,1,29]],"date-time":"2024-01-29T00:00:00Z","timestamp":1706486400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100031060","name":"EUROHPC","doi-asserted-by":"crossref","award":["956748"],"award-info":[{"award-number":["956748"]}],"id":[{"id":"10.13039\/100031060","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Universidad Carlos III"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2024,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The development of adaptive scheduling algorithms that take advantage of malleability has become a crucial area of research in many large-scale projects. Malleable workloads can improve the system\u2019s performance but, at the same time, provide an extra dimension to the scheduling problem. This paper proposes an adaptive, performance-based job scheduling method that emphasizes the backfilling concept with malleability. The proposed method performs the malleability operations only when the estimated execution time of the involved applications is better than or equal to the execution time on the allocated resources without reconfiguration. The reconfiguration feasibility is determined by performance models considering the application scalability and reconfiguration overheads. Different policies for implementing malleability are presented, each targeting a specific workload in terms of job size and scalability. The comprehensive evaluation shows an improvement in the slowdown up to 49% compared to the non-adaptive baseline scheduling algorithm.<\/jats:p>","DOI":"10.1007\/s11227-023-05882-0","type":"journal-article","created":{"date-parts":[[2024,1,29]],"date-time":"2024-01-29T03:02:53Z","timestamp":1706497373000},"page":"11556-11584","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Performance-driven scheduling for malleable workloads"],"prefix":"10.1007","volume":"80","author":[{"given":"Njoud O.","family":"Almaaitah","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David E.","family":"Singh","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taylan","family":"\u00d6zden","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jesus","family":"Carretero","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,1,29]]},"reference":[{"key":"5882_CR1","doi-asserted-by":"crossref","unstructured":"Utrera G, Tabik S, Corbalan J, Labarta J (2012) A job scheduling approach for multi-core clusters based on virtual malleability. In: Euro-Par 2012 Parallel Processing: 18th International Conference, Euro-Par 2012, Rhodes Island, Greece, August 27\u201331, 2012. Proceedings 18, pp 191\u2013203","DOI":"10.1007\/978-3-642-32820-6_20"},{"key":"5882_CR2","doi-asserted-by":"publisher","first-page":"295","DOI":"10.1007\/3-540-60153-8_35","volume-title":"Job Scheduling Strategies for Parallel Processing","author":"DA Lifka","year":"1995","unstructured":"Lifka DA (1995) The ANL\/IBM SP scheduling system. In: Feitelson DG, Rudolph L (eds) Job Scheduling Strategies for Parallel Processing. Springer, Berlin, pp 295\u2013303"},{"key":"5882_CR3","doi-asserted-by":"publisher","first-page":"69","DOI":"10.1016\/j.jpdc.2016.06.013","volume":"97","author":"C G\u00f3mez-Mart\u00edn","year":"2016","unstructured":"G\u00f3mez-Mart\u00edn C, Vega-Rodr\u00edguez MA, Gonz\u00e1lez-S\u00e1nchez J-L (2016) Fattened backfilling: an improved strategy for job scheduling in parallel systems. J Parallel Distrib Comput 97:69\u201377","journal-title":"J Parallel Distrib Comput"},{"key":"5882_CR4","doi-asserted-by":"crossref","unstructured":"Li B, Zhao D (2007) Performance impact of advance reservations from the grid on backfill algorithms. In: Sixth International Conference on Grid and Cooperative Computing (GCC 2007), pp 456\u2013461","DOI":"10.1109\/GCC.2007.96"},{"key":"5882_CR5","doi-asserted-by":"crossref","unstructured":"Srinivasan S, Kettimuthu R, Subramani V, Sadayappan P (2002) Selective reservation strategies for backfill job scheduling. In: Job Scheduling Strategies for Parallel Processing, vol 2537, pp 55\u201371","DOI":"10.1007\/3-540-36180-4_4"},{"key":"5882_CR6","doi-asserted-by":"publisher","unstructured":"Feitelson DG, Weil AM (1998) Utilization and predictability in scheduling the IBM SP2 with backfilling. In: Proceedings of the 1st Merged International Parallel Processing Symposium and Symposium on Parallel and Distributed Processing, IPPS\/SPDP 1998 1998-March, pp 542\u2013546. https:\/\/doi.org\/10.1109\/IPPS.1998.669970","DOI":"10.1109\/IPPS.1998.669970"},{"key":"5882_CR7","doi-asserted-by":"crossref","unstructured":"Tsafrir D, Feitelson DG (2006) The dynamics of backfilling: solving the mystery of why increased inaccuracy may help. In: 2006 IEEE International Symposium on Workload Characterization, pp 131\u2013141","DOI":"10.1109\/IISWC.2006.302737"},{"key":"5882_CR8","doi-asserted-by":"publisher","first-page":"122","DOI":"10.1007\/s11227-019-03004-3","volume":"76","author":"M Naghshnejad","year":"2020","unstructured":"Naghshnejad M, Singhal M (2020) A hybrid scheduling platform: a runtime prediction reliability aware scheduling platform to improve hpc scheduling performance. J Supercomput 76:122\u2013149","journal-title":"J Supercomput"},{"key":"5882_CR9","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1109\/71.932708","volume":"12","author":"A Mu\u2019alem","year":"2001","unstructured":"Mu\u2019alem A, Feitelson D (2001) Utilization, predictability, workloads, and user runtime estimates in scheduling the ibm sp2 with backfilling. IEEE Trans Parallel Distrib Syst 12:529\u2013543. https:\/\/doi.org\/10.1109\/71.932708","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"5882_CR10","unstructured":"EuroHPC JU (2023) The European high performance computing joint undertaking. https:\/\/eurohpc-ju.europa.eu\/research-innovation\/our-projects\/admire_en. Accessed 20 Aug 2023"},{"key":"5882_CR11","unstructured":"EuroHPC JU (2021) Programming environment for European exascale systems. https:\/\/www.deep-projects.eu\/. Accessed 20 Aug 2023"},{"key":"5882_CR12","unstructured":"EuroHPC (2021) Network interconnect for exascale systems. https:\/\/redsea-project.eu\/. Accessed 20 Aug 2023"},{"key":"5882_CR13","unstructured":"EuroHPC-RIA (2021) Towards EXtreme scale Technologies and Accelerators for euROhpc hw\/Sw supercomputing applications for exascale. https:\/\/textarossa.eu\/. Accessed 23 Aug 2023"},{"key":"5882_CR14","doi-asserted-by":"crossref","unstructured":"Sudarsan R, Ribbens CJ (2009) Scheduling resizable parallel applications. In: 2009 IEEE International Symposium on Parallel & Distributed Processing, pp 1\u201310","DOI":"10.1109\/IPDPS.2009.5161077"},{"key":"5882_CR15","doi-asserted-by":"crossref","unstructured":"Sanders P, Schreiber D (2022) Decentralized online scheduling of malleable np-hard jobs. In: European Conference on Parallel Processing, pp 119\u2013135","DOI":"10.1007\/978-3-031-12597-3_8"},{"key":"5882_CR16","unstructured":"SchedMD (2022) Scheduling configuration guide. https:\/\/slurm.schedmd.com\/sched_config.html. Accessed 25 Aug 2023"},{"key":"5882_CR17","doi-asserted-by":"publisher","first-page":"5960","DOI":"10.1007\/s11227-020-03506-5","volume":"77","author":"J Li","year":"2021","unstructured":"Li J, Zhang X, Han L, Ji Z, Dong X, Hu C (2021) Okcm: improving parallel task scheduling in high-performance computing systems using online learning. J Supercomput 77:5960\u20135983","journal-title":"J Supercomput"},{"issue":"10","key":"5882_CR18","doi-asserted-by":"publisher","first-page":"2899","DOI":"10.1016\/j.jpdc.2014.06.008","volume":"74","author":"H Casanova","year":"2014","unstructured":"Casanova H, Giersch A, Legrand A, Quinson M, Suter F (2014) Versatile, scalable, and accurate simulation of distributed applications and platforms. J Parallel Distrib Comput 74(10):2899\u20132917","journal-title":"J Parallel Distrib Comput"},{"key":"5882_CR19","doi-asserted-by":"publisher","first-page":"178","DOI":"10.1007\/978-3-319-61756-5_10","volume-title":"Job Scheduling Strategies for Parallel Processing","author":"P-F Dutot","year":"2017","unstructured":"Dutot P-F, Mercier M, Poquet M, Richard O (2017) Batsim: a realistic language-independent resources and jobs management systems simulator. In: Desai N, Cirne W (eds) Job Scheduling Strategies for Parallel Processing. Springer, Cham, pp 178\u2013197"},{"issue":"1","key":"5882_CR20","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1007\/s10586-019-02905-5","volume":"23","author":"C Galleguillos","year":"2020","unstructured":"Galleguillos C, Kiziltan Z, Netti A, Soto R (2020) Accasim: a customizable workload management simulator for job dispatching research in hpc systems. Clust Comput 23(1):107\u2013122. https:\/\/doi.org\/10.1007\/s10586-019-02905-5","journal-title":"Clust Comput"},{"key":"5882_CR21","unstructured":"Klus\u00e1\u010dek D, T\u00f3th v, Podoln\u00edkov\u00e1 G (2016) Complex job scheduling simulations with alea 4. SIMUTOOLS\u201916. ICST (Institute for Computer Sciences, Social-Informatics and Telecommunications Engineering), Brussels, BEL, pp 124\u2013129"},{"key":"5882_CR22","doi-asserted-by":"publisher","unstructured":"Jokanovic A, D\u2019Amico M, Corbalan J (2018) Evaluating slurm simulator with real-machine slurm and vice versa. In: 2018 IEEE\/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS), pp 72\u201382. https:\/\/doi.org\/10.1109\/PMBS.2018.8641556","DOI":"10.1109\/PMBS.2018.8641556"},{"key":"5882_CR23","doi-asserted-by":"publisher","first-page":"152","DOI":"10.1007\/978-3-319-77398-8_9","volume-title":"Job Scheduling Strategies for Parallel Processing","author":"GP Rodrigo","year":"2018","unstructured":"Rodrigo GP, Elmroth E, \u00d6stberg P-O, Ramakrishnan L (2018) Scsf: a scheduling simulation framework. In: Klus\u00e1\u010dek D, Cirne W, Desai N (eds) Job Scheduling Strategies for Parallel Processing. Springer, Cham, pp 152\u2013173"},{"key":"5882_CR24","doi-asserted-by":"publisher","unstructured":"\u00d6zden T, Beringer T, Mazaheri A, Fard HM, Wolf F (2022) Elastisim: a batch-system simulator for malleable workloads. In: Proceedings of the 51st International Conference on Parallel Processing. ICPP \u201922. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3545008.3545046","DOI":"10.1145\/3545008.3545046"},{"key":"5882_CR25","doi-asserted-by":"crossref","unstructured":"Calotoiu A, Hoefler T, Poke M, Wolf F (2013) Using automated performance modeling to find scalability bugs in complex codes. In: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, pp 1\u201312","DOI":"10.1145\/2503210.2503277"},{"key":"5882_CR26","doi-asserted-by":"crossref","unstructured":"Martin G, Marinescu M-C, Singh DE, Carretero J (2013) Flex-mpi: an mpi extension for supporting dynamic load balancing on heterogeneous non-dedicated systems. In: Euro-Par 2013 Parallel Processing: 19th International Conference, Aachen, Germany, August 26\u201330, 2013. Proceedings 19, pp 138\u2013149","DOI":"10.1007\/978-3-642-40047-6_16"},{"key":"5882_CR27","unstructured":"Ghafoor SK (2007) Modeling of an adaptive parallel system with malleable applications in a distributed computing environment"},{"key":"5882_CR28","doi-asserted-by":"crossref","unstructured":"Lina DH, Ghafoor S, Hines T (2023) Scheduling of elastic message passing applications on hpc systems. In: Job Scheduling Strategies for Parallel Processing: 25th International Workshop, JSSPP 2022, Virtual Event, June 3, 2022, Revised Selected Papers, pp 172\u2013191","DOI":"10.1007\/978-3-031-22698-4_9"},{"key":"5882_CR29","unstructured":"Feitelson D (2005) Parallel workloads archive cs.huji.ac.il. https:\/\/www.cs.huji.ac.il\/labs\/parallel\/workload\/index.html. Accessed 14 May 2023"},{"key":"5882_CR30","unstructured":"KIT (2023) Konfiguration des ForHLR II. https:\/\/www.scc.kit.edu\/dienste\/forhlr2.php. Accessed 25 Jul 2023"},{"key":"5882_CR31","unstructured":"Cruz GM, Singh DE, Marinescu M-C (2015) Optimization techniques for adaptability in mpi applications. Ph.D. thesis, Computer Science and Engineering Department-Universidad Carlos"},{"key":"5882_CR32","unstructured":"Silberschatz A, Galvin PB, Gagne G (2018) Operating system concepts, 10th edn. Wiley. http:\/\/os-book.com\/OS10\/index.html"},{"key":"5882_CR33","doi-asserted-by":"publisher","unstructured":"Feitelson DG, Weil AM (1998) Utilization and predictability in scheduling the IBM SP2 with backfilling. In: Proceedings of the 1st Merged International Parallel Processing Symposium and Symposium on Parallel and Distributed Processing, IPPS\/SPDP 1998 1998-March, pp 542\u2013546. https:\/\/doi.org\/10.1109\/IPPS.1998.669970","DOI":"10.1109\/IPPS.1998.669970"},{"key":"5882_CR34","doi-asserted-by":"publisher","first-page":"1487","DOI":"10.1007\/s11227-019-03004-3","volume":"68","author":"KH Khan","year":"2014","unstructured":"Khan KH, Qureshi K, Abd-El-Barr M (2014) An efficient grid scheduling strategy for data parallel applications. J Supercomput 68:1487\u20131502. https:\/\/doi.org\/10.1007\/s11227-019-03004-3","journal-title":"J Supercomput"},{"key":"5882_CR35","doi-asserted-by":"crossref","unstructured":"Feitelson DG, Rudolph L (1996) Toward convergence in job schedulers for parallel supercomputers. In: Job Scheduling Strategies for Parallel Processing: IPPS\u201996 Workshop Honolulu, Hawaii, April 16, 1996 Proceedings 2, pp 1\u201326","DOI":"10.1007\/BFb0022284"},{"key":"5882_CR36","unstructured":"Fan Y (2021) Job scheduling in high performance computing. arXiv preprint arXiv:2109.09269"},{"key":"5882_CR37","doi-asserted-by":"crossref","unstructured":"D\u2019Amico M, Jokanovic A, Corbalan J (2019) Holistic slowdown driven scheduling and resource management for malleable jobs. In: Proceedings of the 48th International Conference on Parallel Processing, pp 1\u201310","DOI":"10.1145\/3337821.3337909"},{"key":"5882_CR38","unstructured":"Sonmez O, Mohamed H, Lammers W, Epema D et al (2007) Scheduling malleable applications in multicluster systems. In: 2007 IEEE International Conference on Cluster Computing, pp 372\u2013381"},{"key":"5882_CR39","doi-asserted-by":"crossref","unstructured":"Kal\u00e9 LV, Kumar S, DeSouza J (2002) A malleable-job system for timeshared parallel machines. In: 2nd IEEE\/ACM International Symposium on Cluster Computing and the Grid (CCGRID\u201902), pp 230\u2013230","DOI":"10.1109\/CCGRID.2002.1017131"},{"key":"5882_CR40","doi-asserted-by":"crossref","unstructured":"Chadha M, John J, Gerndt M (2020) Extending slurm for dynamic resource-aware adaptive batch scheduling. In: 2020 IEEE 27th International Conference on High Performance Computing, Data, and Analytics (HiPC), pp 223\u2013232","DOI":"10.1109\/HiPC50609.2020.00036"},{"key":"5882_CR41","doi-asserted-by":"publisher","unstructured":"D\u2019Amico M, Jokanovic A, Corbalan J (2019) Holistic slowdown driven scheduling and resource management for malleable jobs. In: Proceedings of the 48th International Conference on Parallel Processing. ICPP \u201919. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3337821.3337909","DOI":"10.1145\/3337821.3337909"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-023-05882-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-023-05882-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-023-05882-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,6]],"date-time":"2024-05-06T07:05:00Z","timestamp":1714979100000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-023-05882-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,29]]},"references-count":41,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2024,5]]}},"alternative-id":["5882"],"URL":"https:\/\/doi.org\/10.1007\/s11227-023-05882-0","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-3361646\/v1","asserted-by":"object"}]},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,29]]},"assertion":[{"value":"23 December 2023","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 January 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}}]}}