{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T07:54:39Z","timestamp":1773215679921,"version":"3.50.1"},"reference-count":14,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T00:00:00Z","timestamp":1773100800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Deanship of Scientific Research at Northern Border University","award":["NBU-FFR-2026-2466-02"],"award-info":[{"award-number":["NBU-FFR-2026-2466-02"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>The widening performance gap between processor speed and memory access latency has made data locality a critical bottleneck in high-performance computing. In Non-Uniform Memory Access (NUMA) and distributed memory systems, remote accesses incur penalties far greater than local operations, degrading the efficiency of scientific and data-intensive workloads. This paper introduces CacheAware, a compiler\u2013runtime framework for data locality-aware scheduling. CacheAware leverages compiler analysis to annotate tasks with memory access footprints and combines this static information with runtime monitoring of cache miss patterns to guide scheduling and dynamic task migration. Unlike existing NUMA balancing or runtime tasking systems, CacheAware integrates both proactive and reactive strategies to minimize cache thrashing and remote memory fetches. Experimental evaluation on scientific benchmarks demonstrates reductions of up to 30% in cache misses and over 20% improvements in execution time compared to Linux AutoNUMA, NUMA-aware schedulers, and task-based runtimes. These results confirm that CacheAware provides a practical and scalable approach for enhancing data locality and accelerating workloads on modern distributed memory systems.<\/jats:p>","DOI":"10.3390\/computers15030181","type":"journal-article","created":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T09:55:38Z","timestamp":1773136538000},"page":"181","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["CacheAware: Data Locality-Aware Scheduling for Distributed Memory Systems"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-4577-6898","authenticated-orcid":false,"given":"Haifa A.","family":"Alanazi","sequence":"first","affiliation":[{"name":"Department of Information Systems and Computer Science, Faculty of Computing and Information Technology, Northern Border University, Rafha 91911, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Abdulaziz G.","family":"Alanazi","sequence":"additional","affiliation":[{"name":"Department of Information Systems and Computer Science, Faculty of Computing and Information Technology, Northern Border University, Rafha 91911, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5948-4260","authenticated-orcid":false,"given":"Nasser S.","family":"Albalawi","sequence":"additional","affiliation":[{"name":"Department of Information Systems and Computer Science, Faculty of Computing and Information Technology, Northern Border University, Rafha 91911, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,3,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Dongarra, J., Heroux, M.A., and Luszczek, P. (2015). HPCG Benchmark: A New Metric for Ranking High Performance Computing Systems, University of Tennessee.","DOI":"10.1177\/1094342015593158"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1145\/216585.216588","article-title":"Hitting the memory wall: Implications of the obvious","volume":"23","author":"Wulf","year":"1995","journal-title":"ACM SIGARCH Comput. Archit. News"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"455","DOI":"10.32996\/jcsts.2025.7.4.54","article-title":"Performance Optimization in NUMA and Multi-Socket Virtual Machine Environments: A Technical Analysis","volume":"7","author":"Kaprakattu","year":"2025","journal-title":"J. Comput. Sci. Technol. Stud."},{"key":"ref_4","unstructured":"Corbet, J. (2025, May 13). Autonuma: An Automatic Numa Balancing System for Linux. Available online: https:\/\/lwn.net\/Articles\/488709\/."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1002\/cpe.1631","article-title":"StarPU: A unified platform for task scheduling on heterogeneous multicore architectures","volume":"23","author":"Augonnet","year":"2011","journal-title":"Concurr. Comput. Pract. Exp."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Bosilca, G., Bouteiller, A., Danalis, A., H\u00e9rault, T., and Dongarra, J. (2013, January 10). PaRSEC: Exploiting heterogeneity to enhance scalability. Proceedings of the International ACM Conference on International Conference on Supercomputing, Eugene, OR, USA.","DOI":"10.1109\/MCSE.2013.98"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1145\/1498765.1498785","article-title":"Roofline: An insightful visual performance model for multicore architectures","volume":"52","author":"Williams","year":"2009","journal-title":"Commun. ACM"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Dashti, M., Fedorova, A., Funston, J., Gaudet, F., Qu\u00e9ma, V., and Roth, M. (2013, January 16\u201320). Traffic management: A holistic approach to memory placement on NUMA systems. Proceedings of the Eighteenth International Conference on Architectural Support for Programming Languages and Operating Systems, Houston, TX, USA.","DOI":"10.1145\/2451116.2451157"},{"key":"ref_9","first-page":"24","article-title":"Memory Hierarchy Optimization Strategies for HighPerformance Computing Architectures","volume":"6","author":"Vaithianathan","year":"2025","journal-title":"Int. J. Emerg. Trends Technol. Comput. Sci."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lu, Q., Alias, C., Bondhugula, U., Henretty, T., Krishnamoorthy, S., Ramanujam, J., Rountev, A., Sadayappan, P., Chen, Y., and Lin, H. (2009, January 12\u201316). Data layout transformation for enhancing data locality on nuca chip multiprocessors. Proceedings of the 2009 18th International Conference on Parallel Architectures and Compilation Techniques, Raleigh, NC, USA.","DOI":"10.1109\/PACT.2009.36"},{"key":"ref_11","unstructured":"Bondhugula, U., Hartono, A., Ramanujam, J., and Sadayappan, P. (2008, January 7\u201313). Pluto: A practical and fully automatic polyhedral program optimization system. Proceedings of the ACM SIGPLAN 2008 Conference on Programming Language Design and Implementation, Tucson, AZ, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Bauer, M., Treichler, S., and Aiken, A. (2012, January 10\u201316). Legion: Expressing locality and independence with logical regions. Proceedings of the SC\u201912: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, Salt Lake City, UT, USA.","DOI":"10.1109\/SC.2012.71"},{"key":"ref_13","unstructured":"Kale, L.V., and Krishnan, S. (October, January 26). CHARM++: A portable concurrent object oriented system based on C++. Proceedings of the Eighth Annual Conference on Object-Oriented Programming Systems, Languages, and Applications, Washington, DC, USA."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Hoefler, T., Belli, R., and Gerstenberger, R. (2010, January 13\u201319). Characterizing the influence of system noise on large-scale applications by simulation. 2010 ACM\/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis, New Orleans, LA, USA.","DOI":"10.1109\/SC.2010.12"}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/3\/181\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T10:03:42Z","timestamp":1773137022000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/3\/181"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,10]]},"references-count":14,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2026,3]]}},"alternative-id":["computers15030181"],"URL":"https:\/\/doi.org\/10.3390\/computers15030181","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,10]]}}}