{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T18:05:48Z","timestamp":1780509948869,"version":"3.54.1"},"reference-count":138,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2016,12,5]],"date-time":"2016-12-05T00:00:00Z","timestamp":1480896000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"MCTI\/RNP Brazil under the HPC4E project","award":["689772"],"award-info":[{"award-number":["689772"]}]},{"DOI":"10.13039\/501100002322","name":"CAPES","doi-asserted-by":"crossref","award":["PVE 117\/2013"],"award-info":[{"award-number":["PVE 117\/2013"]}],"id":[{"id":"10.13039\/501100002322","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2017,12,31]]},"abstract":"<jats:p>Shared memory architectures have recently experienced a large increase in thread-level parallelism, leading to complex memory hierarchies with multiple cache memory levels and memory controllers. These new designs created a Non-Uniform Memory Access (NUMA) behavior, where the performance and energy consumption of memory accesses depend on the place where the data is located in the memory hierarchy. Accesses to local caches or memory controllers are generally more efficient than accesses to remote ones. A common way to improve the locality and balance of memory accesses is to determine the mapping of threads to cores and data to memory controllers based on the affinity between threads and data. Such mapping techniques can operate at different hardware and software levels, which impacts their complexity, applicability, and the resulting performance and energy consumption gains. In this article, we introduce a taxonomy to classify different mapping mechanisms and provide a comprehensive overview of existing solutions.<\/jats:p>","DOI":"10.1145\/3006385","type":"journal-article","created":{"date-parts":[[2016,12,6]],"date-time":"2016-12-06T16:03:07Z","timestamp":1481040187000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":32,"title":["Affinity-Based Thread and Data Mapping in Shared Memory Systems"],"prefix":"10.1145","volume":"49","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9064-7806","authenticated-orcid":false,"given":"Matthias","family":"Diener","sequence":"first","affiliation":[{"name":"Informatics Institute, Federal University of Rio Grande do Sul, RS, Brazil"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eduardo H. M.","family":"Cruz","sequence":"additional","affiliation":[{"name":"Informatics Institute, Federal University of Rio Grande do Sul, RS, Brazil"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marco A. Z.","family":"Alves","sequence":"additional","affiliation":[{"name":"Department of Informatics, Federal University of Paran\u00e1, PR, Brazil"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Philippe O. A.","family":"Navaux","sequence":"additional","affiliation":[{"name":"Informatics Institute, Federal University of Rio Grande do Sul, RS, Brazil"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Israel","family":"Koren","sequence":"additional","affiliation":[{"name":"Department of Electrical 8 Computer Engineering, University of Massachusetts at Amherst, Amherst, MA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2016,12,5]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.v22:6"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897783"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/454\/1\/012010"},{"key":"e_1_2_1_4_1","unstructured":"Argonne National Laboratory. 2014. Using the Hydra Process Manager. Retrieved 2015-06-08 from https:\/\/wiki.mpich.org\/mpich\/index.php\/Using_the_Hydra_Process_Manager.  Argonne National Laboratory. 2014. Using the Hydra Process Manager. Retrieved 2015-06-08 from https:\/\/wiki.mpich.org\/mpich\/index.php\/Using_the_Hydra_Process_Manager."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854314"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1531793.1531803"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1177\/109434209100500306"},{"key":"e_1_2_1_9_1","first-page":"1","article-title":"Communication lower bounds and optimal algorithms for numerical linear algebra","author":"Ballard G.","year":"2014","unstructured":"G. Ballard , E. Carson , J. Demmel , M. Hoemmen , N. Knight , and O. Schwartz . 2014 . Communication lower bounds and optimal algorithms for numerical linear algebra . Acta Numerica 23 , May (2014), 1 -- 155 . G. Ballard, E. Carson, J. Demmel, M. Hoemmen, N. Knight, and O. Schwartz. 2014. Communication lower bounds and optimal algorithms for numerical linear algebra. Acta Numerica 23, May (2014), 1--155.","journal-title":"Acta Numerica 23"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306792"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485943"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2835238.2835239"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1454115.1454128"},{"key":"e_1_2_1_15_1","volume-title":"ACM\/IEEE Conference on Supercomputing (SC\u201900)","author":"Bircsak John","unstructured":"John Bircsak , Peter Craig , RaeLyn Crowell , Zarka Cvetanovic , Jonathan Harris , C. Alexander Nelson , and Carl D. Offner . 2000. Extending OpenMP for NUMA machines . In ACM\/IEEE Conference on Supercomputing (SC\u201900) . 163--181. John Bircsak, Peter Craig, RaeLyn Crowell, Zarka Cvetanovic, Jonathan Harris, C. Alexander Nelson, and Carl D. Offner. 2000. Extending OpenMP for NUMA machines. In ACM\/IEEE Conference on Supercomputing (SC\u201900). 163--181."},{"key":"e_1_2_1_16_1","volume-title":"USENIX Annual Technical Conference (ATC\u201911)","author":"Blagodurov Sergey","year":"2011","unstructured":"Sergey Blagodurov , Sergey Zhuravlev , Mohammad Dashti , and Alexandra Fedorova . 2011 . A case for NUMA-aware contention management on multicore systems . In USENIX Annual Technical Conference (ATC\u201911) . 557--571. Sergey Blagodurov, Sergey Zhuravlev, Mohammad Dashti, and Alexandra Fedorova. 2011. A case for NUMA-aware contention management on multicore systems. In USENIX Annual Technical Conference (ATC\u201911). 557--571."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/SFCS.1994.365680"},{"key":"e_1_2_1_18_1","volume-title":"Joint International Conference on Vector and Parallel Processing (CONPAR 90 -- VAPP IV). 405--416","author":"Jacques","unstructured":"Jacques E. Boillat and Peter G. Kropf. 1990. A fast distributed mapping algorithm . In Joint International Conference on Vector and Parallel Processing (CONPAR 90 -- VAPP IV). 405--416 . Jacques E. Boillat and Peter G. Kropf. 1990. A fast distributed mapping algorithm. In Joint International Conference on Vector and Parallel Processing (CONPAR 90 -- VAPP IV). 405--416."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/0743-7315(92)90051-N"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/106975.106994"},{"key":"e_1_2_1_21_1","volume-title":"July","author":"Brandfass Barbara","year":"2013","unstructured":"Barbara Brandfass , Thomas Alrutz , and Thomas Gerhold . 2013. Rank reordering for MPI communication optimization. Computers 8 Fluids 80 , July ( 2013 ), 372--380. Barbara Brandfass, Thomas Alrutz, and Thomas Gerhold. 2013. Rank reordering for MPI communication optimization. Computers 8 Fluids 80, July (2013), 372--380."},{"key":"e_1_2_1_22_1","volume-title":"Symposium on Experiences with Distributed and Multiprocessor Systems (SEDMS). 1--18","author":"Brecht Timothy","year":"1993","unstructured":"Timothy Brecht . 1993 . On the importance of parallel application placement in NUMA multiprocessors . In Symposium on Experiences with Distributed and Multiprocessor Systems (SEDMS). 1--18 . Timothy Brecht. 1993. On the importance of parallel application placement in NUMA multiprocessors. In Symposium on Experiences with Distributed and Multiprocessor Systems (SEDMS). 1--18."},{"key":"e_1_2_1_23_1","volume-title":"International Workshop on Power-Aware Computer Systems (PACS\u201900)","author":"Brooks David","year":"2000","unstructured":"David Brooks , Margaret Martonosi , John-David Wellman , and Pradip Bose . 2000 . Power-performance modeling and tradeoff analysis for a high end microprocessor . In International Workshop on Power-Aware Computer Systems (PACS\u201900) . 126--136. David Brooks, Margaret Martonosi, John-David Wellman, and Pradip Bose. 2000. Power-performance modeling and tradeoff analysis for a high end microprocessor. In International Workshop on Power-Aware Computer Systems (PACS\u201900). 126--136."},{"key":"e_1_2_1_24_1","volume-title":"Structuring the execution of OpenMP applications for multicore architectures","author":"Broquedis Fran\u00e7ois","unstructured":"Fran\u00e7ois Broquedis , Olivier Aumage , Brice Goglin , Samuel Thibault , Pierre-Andr\u00e9 Wacrenier , and Raymond Namyst . 2010a. Structuring the execution of OpenMP applications for multicore architectures . In IEEE International Parallel 8 Distributed Processing Symposium (IPDPS\u2019 10). 1--10. Fran\u00e7ois Broquedis, Olivier Aumage, Brice Goglin, Samuel Thibault, Pierre-Andr\u00e9 Wacrenier, and Raymond Namyst. 2010a. Structuring the execution of OpenMP applications for multicore architectures. In IEEE International Parallel 8 Distributed Processing Symposium (IPDPS\u201910). 1--10."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/PDP.2010.67"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-010-0136-3"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1177\/109434200001400404"},{"key":"e_1_2_1_29_1","volume-title":"European Workshop on OpenMP (EWOMP\u201902)","author":"Mark Bull J.","year":"2002","unstructured":"J. Mark Bull and Chris Johnson . 2002 . Data distribution, migration and replication on a cc-NUMA architecture . In European Workshop on OpenMP (EWOMP\u201902) . 1--5. J. Mark Bull and Chris Johnson. 2002. Data distribution, migration and replication on a cc-NUMA architecture. In European Workshop on OpenMP (EWOMP\u201902). 1--5."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/195473.195485"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/1183401.1183451"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1080695.1070001"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1105734.1105745"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2007.43"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/782814.782833"},{"key":"e_1_2_1_37_1","unstructured":"Jonathan Corbet. 2012a. AutoNUMA: The Other Approach to NUMA Scheduling. Retrieved 2015-06-08 from http:\/\/lwn.net\/Articles\/488709\/  Jonathan Corbet. 2012a. AutoNUMA: The Other Approach to NUMA Scheduling. Retrieved 2015-06-08 from http:\/\/lwn.net\/Articles\/488709\/"},{"key":"e_1_2_1_38_1","unstructured":"Jonathan Corbet. 2012b. Toward Better NUMA Scheduling. Retrieved 2015-06-08 from http:\/\/lwn.net\/Articles\/486858\/.  Jonathan Corbet. 2012b. Toward Better NUMA Scheduling. Retrieved 2015-06-08 from http:\/\/lwn.net\/Articles\/486858\/."},{"key":"e_1_2_1_39_1","unstructured":"Eduardo H. M. Cruz. 2012. Dynamic Detection of the Communication Pattern in Shared Memory Environments for Thread Mapping. Master\u2019s thesis.  Eduardo H. M. Cruz. 2012. Dynamic Detection of the Communication Pattern in Shared Memory Environments for Thread Mapping. Master\u2019s thesis."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.197"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2013.11.006"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/SBAC-PAD.2014.22"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2012.56"},{"key":"e_1_2_1_44_1","first-page":"685","article-title":"Communication-aware thread mapping using the translation lookaside buffer","volume":"22","author":"Cruz Eduardo H. M.","year":"2015","unstructured":"Eduardo H. M. Cruz , Matthias Diener , and Philippe O. A. Navaux . 2015 a. Communication-aware thread mapping using the translation lookaside buffer . Concurrency Computation: Practice and Experience 22 , 6 (2015), 685 -- 701 . Eduardo H. M. Cruz, Matthias Diener, and Philippe O. A. Navaux. 2015a. Communication-aware thread mapping using the translation lookaside buffer. Concurrency Computation: Practice and Experience 22, 6 (2015), 685--701.","journal-title":"Concurrency Computation: Practice and Experience"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/PDP.2015.25"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2013.6522311"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2451116.2451157"},{"key":"e_1_2_1_49_1","volume-title":"Catalyurek","author":"Devine Karen D.","year":"2006","unstructured":"Karen D. Devine , Erik G. Boman , Robert T. Heaphy , Rob H. Bisseling , and Umit V . Catalyurek . 2006 . Parallel hypergraph partitioning for scientific computing. In IEEE International Parallel 8 Distributed Processing Symposium (IPDPS\u2019 06). 124--133. Karen D. Devine, Erik G. Boman, Robert T. Heaphy, Rob H. Bisseling, and Umit V. Catalyurek. 2006. Parallel hypergraph partitioning for scientific computing. In IEEE International Parallel 8 Distributed Processing Symposium (IPDPS\u201906). 124--133."},{"key":"e_1_2_1_50_1","volume-title":"Evaluating Thread Placement Improvements in Multi-core Architectures. Master\u2019s thesis","author":"Diener Matthias","unstructured":"Matthias Diener . 2010. Evaluating Thread Placement Improvements in Multi-core Architectures. Master\u2019s thesis . Berlin Institute of Technology . Matthias Diener. 2010. Evaluating Thread Placement Improvements in Multi-core Architectures. Master\u2019s thesis. Berlin Institute of Technology."},{"key":"e_1_2_1_52_1","volume-title":"Navaux","author":"Diener Matthias","year":"2015","unstructured":"Matthias Diener , Eduardo H. M. Cruz , Marco A. Z. Alves , Mohammad S. Alhakeem , and Philippe O. A . Navaux . 2015 . Locality and balance for communication-aware thread mapping in multicore systems. In Euro-Par . 196--208. Matthias Diener, Eduardo H. M. Cruz, Marco A. Z. Alves, Mohammad S. Alhakeem, and Philippe O. A. Navaux. 2015. Locality and balance for communication-aware thread mapping in multicore systems. In Euro-Par. 196--208."},{"key":"e_1_2_1_53_1","volume-title":"Euromicro International Conference on Parallel, Distributed, and Network-based Processing (PDP\u201916)","author":"Diener Matthias","unstructured":"Matthias Diener , Eduardo H. M. Cruz , Marco A. Z. Alves , and Philippe O. A. Navaux . 2016. Communication in shared memory: Concepts, definitions, and efficient detection . In Euromicro International Conference on Parallel, Distributed, and Network-based Processing (PDP\u201916) . 151--158. Matthias Diener, Eduardo H. M. Cruz, Marco A. Z. Alves, and Philippe O. A. Navaux. 2016. Communication in shared memory: Concepts, definitions, and efficient detection. In Euromicro International Conference on Parallel, Distributed, and Network-based Processing (PDP\u201916). 151--158."},{"key":"e_1_2_1_54_1","first-page":"1","article-title":"Kernel-based thread and data mapping for improved memory affinity","volume":"26","author":"Diener Matthias","year":"2015","unstructured":"Matthias Diener , Eduardo H. M. Cruz , Marco A. Z. Alves , Philippe O. A. Navaux , Anselm Busse , and Hans-Ulrich Heiss . 2015 . Kernel-based thread and data mapping for improved memory affinity . IEEE Transactions on Parallel and Distributed Systems (TPDS) 26 , X (2015), 1 -- 14 . Matthias Diener, Eduardo H. M. Cruz, Marco A. Z. Alves, Philippe O. A. Navaux, Anselm Busse, and Hans-Ulrich Heiss. 2015. Kernel-based thread and data mapping for improved memory affinity. IEEE Transactions on Parallel and Distributed Systems (TPDS) 26, X (2015), 1--14.","journal-title":"IEEE Transactions on Parallel and Distributed Systems (TPDS)"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2013.57"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/PDP.2015.11"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628085"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2015.01.005"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.peva.2015.03.001"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCC.2010.114"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2013.6495009"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/CSE.2008.51"},{"key":"e_1_2_1_65_1","unstructured":"Fabrice Dupros Christiane Pousa Alexandre Carissimi and Jean-Fran\u00e7ois M\u00e9haut. 2010. Parallel simulations of seismic wave propagation on NUMA architectures. In Parallel Computing: From Multicores and GPU\u2019s to Petascale. 67--74.  Fabrice Dupros Christiane Pousa Alexandre Carissimi and Jean-Fran\u00e7ois M\u00e9haut. 2010. Parallel simulations of seismic wave propagation on NUMA architectures. In Parallel Computing: From Multicores and GPU\u2019s to Petascale. 67--74."},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-30961-8_2"},{"key":"e_1_2_1_67_1","volume-title":"Henry Burkhardt III, and James Rothnie","author":"Frank Steven","year":"1993","unstructured":"Steven Frank , Henry Burkhardt III, and James Rothnie . 1993 . The KSR1: Bridging the gap between shared memory and MPPs. In IEEE Compcon . 285--294. Steven Frank, Henry Burkhardt III, and James Rothnie. 1993. The KSR1: Bridging the gap between shared memory and MPPs. In IEEE Compcon. 285--294."},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.5194\/acp-9-2843-2009"},{"key":"e_1_2_1_69_1","volume-title":"Woodall","author":"Gabriel Edgar","year":"2004","unstructured":"Edgar Gabriel , Graham E. Fagg , George Bosilca , Thara Angskun , Jack J. Dongarra , Jeffrey M. Squyres , Vishal Sahay , Prabhanjan Kambadur , Brian Barrett , Andrew Lumsdaine , Ralph H. Castain , David J. Daniel , Richard L. Graham , and Timothy S . Woodall . 2004 . Open MPI: Goals , concept, and design of a next generation MPI implementation. In Recent Advances in Parallel Virtual Machine and Message Passing Interface (PVMMPI\u2019 04). 97--104. Edgar Gabriel, Graham E. Fagg, George Bosilca, Thara Angskun, Jack J. Dongarra, Jeffrey M. Squyres, Vishal Sahay, Prabhanjan Kambadur, Brian Barrett, Andrew Lumsdaine, Ralph H. Castain, David J. Daniel, Richard L. Graham, and Timothy S. Woodall. 2004. Open MPI: Goals, concept, and design of a next generation MPI implementation. In Recent Advances in Parallel Virtual Machine and Message Passing Interface (PVMMPI\u201904). 97--104."},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/2814328"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid.2016.91"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.19"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/PDP.2015.21"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5161101"},{"key":"e_1_2_1_76_1","doi-asserted-by":"crossref","unstructured":"William Gropp. 2002. MPICH2: A new start for MPI implementations. In Recent Advances in Parallel Virtual Machine and Message Passing Interface.   William Gropp. 2002. MPICH2: A new start for MPI implementations. In Recent Advances in Parallel Virtual Machine and Message Passing Interface.","DOI":"10.1007\/3-540-45825-5_5"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.5555\/520549.822785"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2011.59"},{"key":"e_1_2_1_79_1","unstructured":"Intel. 2012. Using KMP_AFFINITY to create OpenMP thread mapping to OS proc IDs. Retrieved 2015-06-08 from https:\/\/software.intel.com\/en-us\/articles\/using-kmp-affinity-to-create-openmp-thread-mapping-to-os-proc-ids.  Intel. 2012. Using KMP_AFFINITY to create OpenMP thread mapping to OS proc IDs. Retrieved 2015-06-08 from https:\/\/software.intel.com\/en-us\/articles\/using-kmp-affinity-to-create-openmp-thread-mapping-to-os-proc-ids."},{"key":"e_1_2_1_80_1","unstructured":"Intel. 2013. Intel Trace Analyzer and Collector. Retrieved from http:\/\/software.intel.com\/en-us\/intel-trace-analyzer.  Intel. 2013. Intel Trace Analyzer and Collector. Retrieved from http:\/\/software.intel.com\/en-us\/intel-trace-analyzer."},{"key":"e_1_2_1_81_1","volume-title":"Automatically optimized core mapping to subdomains of domain decomposition method on multicore parallel environments. Computers 8 Fluids 80 (jul","author":"Ito Satoshi","year":"2013","unstructured":"Satoshi Ito , Kazuya Goto , and Kenji Ono . 2013. Automatically optimized core mapping to subdomains of domain decomposition method on multicore parallel environments. Computers 8 Fluids 80 (jul 2013 ), 88--93. Satoshi Ito, Kazuya Goto, and Kenji Ono. 2013. Automatically optimized core mapping to subdomains of domain decomposition method on multicore parallel environments. Computers 8 Fluids 80 (jul 2013), 88--93."},{"key":"e_1_2_1_82_1","doi-asserted-by":"crossref","unstructured":"Emmanuel Jeannot and Guillaume Mercier. 2010. Near-optimal placement of MPI processes on hierarchical NUMA architectures. In Euro-Par Parallel Processing. 199--210.   Emmanuel Jeannot and Guillaume Mercier. 2010. Near-optimal placement of MPI processes on hierarchical NUMA architectures. In Euro-Par Parallel Processing. 199--210.","DOI":"10.1007\/978-3-642-15291-7_20"},{"key":"e_1_2_1_83_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2013.104"},{"key":"e_1_2_1_84_1","unstructured":"H. Jin M. Frumkin and J. Yan. 1999. The OpenMP Implementation of NAS Parallel Benchmarks and Its Performance. Technical Report October. NASA.  H. Jin M. Frumkin and J. Yan. 1999. The OpenMP Implementation of NAS Parallel Benchmarks and Its Performance. Technical Report October. NASA."},{"key":"e_1_2_1_85_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2012.47"},{"key":"e_1_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.5555\/645606.661329"},{"key":"e_1_2_1_87_1","doi-asserted-by":"publisher","DOI":"10.1137\/S1064827595287997"},{"key":"e_1_2_1_89_1","doi-asserted-by":"crossref","unstructured":"Tobias Klug Michael Ott Josef Weidendorfer and Carsten Trinitis. 2008. autopin -- automated optimization of thread-to-core pinning on multicore systems. In Transactions on High-Performance Embedded Architectures and Compilers (HiPEAC). 219--235.   Tobias Klug Michael Ott Josef Weidendorfer and Carsten Trinitis. 2008. autopin -- automated optimization of thread-to-core pinning on multicore systems. In Transactions on High-Performance Embedded Architectures and Compilers (HiPEAC). 219--235.","DOI":"10.1007\/978-3-642-19448-1_12"},{"key":"e_1_2_1_90_1","volume-title":"USENIX Annual Technical Conference (ATC\u201912)","author":"Lachaize Renaud","year":"2012","unstructured":"Renaud Lachaize , Baptiste Lepers , and Vivien Qu\u00e9ma . 2012 . MemProf: A memory profiler for NUMA multicore systems . In USENIX Annual Technical Conference (ATC\u201912) . 53--64. Renaud Lachaize, Baptiste Lepers, and Vivien Qu\u00e9ma. 2012. MemProf: A memory profiler for NUMA multicore systems. In USENIX Annual Technical Conference (ATC\u201912). 53--64."},{"key":"e_1_2_1_91_1","first-page":"576","article-title":"Affinity-on-next-touch: An extension to the linux kernel for NUMA architectures. Lecture Notes in Computer Science 6067 LNCS","volume":"1","author":"Lankes Stefan","year":"2010","unstructured":"Stefan Lankes , Boris Bierbaum , and Thomas Bemmerl . 2010 . Affinity-on-next-touch: An extension to the linux kernel for NUMA architectures. Lecture Notes in Computer Science 6067 LNCS , PART 1 (2010), 576 -- 585 . Stefan Lankes, Boris Bierbaum, and Thomas Bemmerl. 2010. Affinity-on-next-touch: An extension to the linux kernel for NUMA architectures. Lecture Notes in Computer Science 6067 LNCS, PART 1 (2010), 576--585.","journal-title":"PART"},{"key":"e_1_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.1145\/149439.133082"},{"key":"e_1_2_1_94_1","volume-title":"Free and Open Source Developers European Meeting (FOSDEM\u201911)","author":"Lattner Chris","year":"2011","unstructured":"Chris Lattner . 2011 . LLVM and clang: Advancing compiler technology . In Free and Open Source Developers European Meeting (FOSDEM\u201911) . Chris Lattner. 2011. LLVM and clang: Advancing compiler technology. In Free and Open Source Developers European Meeting (FOSDEM\u201911)."},{"key":"e_1_2_1_95_1","doi-asserted-by":"publisher","DOI":"10.1145\/139669.139706"},{"key":"e_1_2_1_97_1","doi-asserted-by":"publisher","DOI":"10.1145\/1504176.1504188"},{"key":"e_1_2_1_98_1","doi-asserted-by":"publisher","DOI":"10.1145\/2555243.2555271"},{"key":"e_1_2_1_99_1","doi-asserted-by":"publisher","DOI":"10.1145\/1088149.1088201"},{"key":"e_1_2_1_100_1","doi-asserted-by":"publisher","DOI":"10.1145\/1065010.1065034"},{"key":"e_1_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.982916"},{"key":"e_1_2_1_102_1","doi-asserted-by":"publisher","DOI":"10.1145\/2259016.2259046"},{"key":"e_1_2_1_103_1","doi-asserted-by":"publisher","DOI":"10.1145\/2688500.2688509"},{"key":"e_1_2_1_104_1","doi-asserted-by":"publisher","DOI":"10.1145\/1122971.1122987"},{"key":"e_1_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2010.08.015"},{"key":"e_1_2_1_106_1","volume-title":"International Parallel Processing Symposium (IPPS\u201995)","author":"Marchetti Michael","unstructured":"Michael Marchetti , Leonidas Kontothanassis , Ricardo Bianchini , and Michael L. Scott . 1995. Using simple page placement policies to reduce the cost of cache fills in coherent shared-memory systems . In International Parallel Processing Symposium (IPPS\u201995) . 480--485. Michael Marchetti, Leonidas Kontothanassis, Ricardo Bianchini, and Michael L. Scott. 1995. Using simple page placement policies to reduce the cost of cache fills in coherent shared-memory systems. In International Parallel Processing Symposium (IPPS\u201995). 480--485."},{"key":"e_1_2_1_107_1","volume-title":"Euromicro International Conference on Parallel, Distributed, and Network-based Processing (PDP\u201916)","author":"Mariano Artur","unstructured":"Artur Mariano , Matthias Diener , Christian Bischof , and Philippe O. A. Navaux . 2016. Analyzing and improving memory access patterns of large irregular applications on NUMA machines . In Euromicro International Conference on Parallel, Distributed, and Network-based Processing (PDP\u201916) . 382--387. Artur Mariano, Matthias Diener, Christian Bischof, and Philippe O. A. Navaux. 2016. Analyzing and improving memory access patterns of large irregular applications on NUMA machines. In Euromicro International Conference on Parallel, Distributed, and Network-based Processing (PDP\u201916). 382--387."},{"key":"e_1_2_1_108_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2015.68"},{"key":"e_1_2_1_109_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2010.5452060"},{"key":"e_1_2_1_110_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-03770-2_17"},{"key":"e_1_2_1_111_1","doi-asserted-by":"publisher","DOI":"10.5555\/2042476.2042483"},{"key":"e_1_2_1_112_1","doi-asserted-by":"publisher","DOI":"10.5555\/850941.852887"},{"key":"e_1_2_1_113_1","doi-asserted-by":"crossref","unstructured":"Dimitrios S. Nikolopoulos Theodore S. Papatheodorou Constantine D. Polychronopoulos Jes\u00fas Labarta and Eduard Ayguad\u00e9. 2000b. UPMLIB: A runtime system for tuning the memory performance of OpenMP programs on scalable shared-memory multiprocessors. In Languages Compilers and Run-Time Systems for Scalable Computers (LCR\u201900). 85--99.   Dimitrios S. Nikolopoulos Theodore S. Papatheodorou Constantine D. Polychronopoulos Jes\u00fas Labarta and Eduard Ayguad\u00e9. 2000b. UPMLIB: A runtime system for tuning the memory performance of OpenMP programs on scalable shared-memory multiprocessors. In Languages Compilers and Run-Time Systems for Scalable Computers (LCR\u201900). 85--99.","DOI":"10.1007\/3-540-40889-4_7"},{"key":"e_1_2_1_114_1","doi-asserted-by":"publisher","DOI":"10.5555\/370049.370454"},{"key":"e_1_2_1_115_1","doi-asserted-by":"publisher","DOI":"10.1145\/331532.331570"},{"key":"e_1_2_1_116_1","doi-asserted-by":"publisher","DOI":"10.1145\/1640089.1640117"},{"key":"e_1_2_1_117_1","volume-title":"Version 4.0.","author":"Architecture Review Board MP","year":"2013","unstructured":"Open MP Architecture Review Board . 2013. OpenMP Application Program Interface , Version 4.0. ( 2013 ). OpenMP Architecture Review Board. 2013. OpenMP Application Program Interface, Version 4.0. (2013)."},{"key":"e_1_2_1_118_1","unstructured":"Oracle. 2010. Solaris OS Tuning Features. Retrieved 2015-06-16 from http:\/\/docs.oracle.com\/cd\/E18659_01\/html\/821-1381\/aewda.html  Oracle. 2010. Solaris OS Tuning Features. Retrieved 2015-06-16 from http:\/\/docs.oracle.com\/cd\/E18659_01\/html\/821-1381\/aewda.html"},{"key":"e_1_2_1_119_1","volume-title":"Workshop on Programmability Issues for Multi-Core Computers (MULTIPROG\u201908)","author":"Ott Michael","year":"2008","unstructured":"Michael Ott , Tobias Klug , Josef Weidendorfer , and Carsten Trinitis . 2008 . autopin -- automated optimization of thread-to-core pinning on multicore systems . In Workshop on Programmability Issues for Multi-Core Computers (MULTIPROG\u201908) . Michael Ott, Tobias Klug, Josef Weidendorfer, and Carsten Trinitis. 2008. autopin -- automated optimization of thread-to-core pinning on multicore systems. In Workshop on Programmability Issues for Multi-Core Computers (MULTIPROG\u201908)."},{"key":"e_1_2_1_120_1","doi-asserted-by":"publisher","DOI":"10.1109\/SHPCC.1994.296682"},{"key":"e_1_2_1_121_1","unstructured":"Fran\u00e7ois Pellegrini. 2010. Scotch and Libscotch 5.1 User\u2019s Guide. Technical Report.  Fran\u00e7ois Pellegrini. 2010. Scotch and Libscotch 5.1 User\u2019s Guide. Technical Report."},{"key":"e_1_2_1_122_1","volume-title":"International Conference on Parallel Processing (ICPP\u201985)","author":"Pfister Gregory F.","year":"1985","unstructured":"Gregory F. Pfister , William C. Brantley , David A. George , Steve L. Harvey , Wally J. Kleinfelder , Kevin P. McAuliffe , Evelin S. Melton , V. Alan Norton , and Jodi Weiss . 1985 . The IBM research parallel processor prototype (RP3): Introduction and architecture . In International Conference on Parallel Processing (ICPP\u201985) . 764--771. Gregory F. Pfister, William C. Brantley, David A. George, Steve L. Harvey, Wally J. Kleinfelder, Kevin P. McAuliffe, Evelin S. Melton, V. Alan Norton, and Jodi Weiss. 1985. The IBM research parallel processor prototype (RP3): Introduction and architecture. In International Conference on Parallel Processing (ICPP\u201985). 764--771."},{"key":"e_1_2_1_123_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628077"},{"key":"e_1_2_1_124_1","doi-asserted-by":"publisher","DOI":"10.1145\/2248487.2151002"},{"key":"e_1_2_1_125_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2012.311"},{"key":"e_1_2_1_127_1","volume-title":"International Conference on High Performance Computing for Computational Science (VECPAR\u201910)","author":"Ribeiro Christiane Pousa","year":"2010","unstructured":"Christiane Pousa Ribeiro , Marcio Castro , Jean-Fran\u00e7ois M\u00e9haut , and Alexandre Carissimi . 2010 . Improving memory affinity of geophysics applications on NUMA platforms using Minas . In International Conference on High Performance Computing for Computational Science (VECPAR\u201910) . 279--292. Christiane Pousa Ribeiro, Marcio Castro, Jean-Fran\u00e7ois M\u00e9haut, and Alexandre Carissimi. 2010. Improving memory affinity of geophysics applications on NUMA platforms using Minas. In International Conference on High Performance Computing for Computational Science (VECPAR\u201910). 279--292."},{"key":"e_1_2_1_128_1","doi-asserted-by":"publisher","DOI":"10.1109\/SBAC-PAD.2009.16"},{"key":"e_1_2_1_129_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCC.2009.5202271"},{"key":"e_1_2_1_130_1","doi-asserted-by":"crossref","unstructured":"John Shalf Sudip Dosanjh and John Morrison. 2010. Exascale computing technology challenges. In High Performance Computing for Computational Science (VECPAR\u201910). 1--25.   John Shalf Sudip Dosanjh and John Morrison. 2010. Exascale computing technology challenges. In High Performance Computing for Computational Science (VECPAR\u201910). 1--25.","DOI":"10.1007\/978-3-642-19328-6_1"},{"key":"e_1_2_1_131_1","doi-asserted-by":"publisher","DOI":"10.1145\/169627.169699"},{"key":"e_1_2_1_132_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-013-0918-7"},{"key":"e_1_2_1_134_1","doi-asserted-by":"publisher","DOI":"10.1145\/1272996.1273004"},{"key":"e_1_2_1_135_1","doi-asserted-by":"publisher","DOI":"10.1145\/1366219.1366222"},{"key":"e_1_2_1_136_1","unstructured":"The Open MPI project. 2009. Portable Linux Processor Affinity (PLPA). Retrieved 2015-06-08 from https:\/\/www.open-mpi.org\/projects\/plpa\/.  The Open MPI project. 2009. Portable Linux Processor Affinity (PLPA). Retrieved 2015-06-08 from https:\/\/www.open-mpi.org\/projects\/plpa\/."},{"key":"e_1_2_1_137_1","unstructured":"The Open MPI Project. 2013. mpirun(1) man page (version 1.6.4). Retrieved 2016-02-08 from http:\/\/www.open-mpi.org\/doc\/v1.6\/man1\/mpirun.1.php#sect9.  The Open MPI Project. 2013. mpirun(1) man page (version 1.6.4). Retrieved 2016-02-08 from http:\/\/www.open-mpi.org\/doc\/v1.6\/man1\/mpirun.1.php#sect9."},{"key":"e_1_2_1_138_1","doi-asserted-by":"crossref","unstructured":"Samuel Thibault Raymond Namyst and Pierre-Andr\u00e9 Wacrenier. 2007. Building portable thread schedulers for hierarchical multiprocessors: The bubblesched framework. In Euro-Par. 42--51.   Samuel Thibault Raymond Namyst and Pierre-Andr\u00e9 Wacrenier. 2007. Building portable thread schedulers for hierarchical multiprocessors: The bubblesched framework. In Euro-Par. 42--51.","DOI":"10.1007\/978-3-540-74466-5_6"},{"key":"e_1_2_1_139_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2004.64"},{"key":"e_1_2_1_140_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2008.05.006"},{"key":"e_1_2_1_141_1","doi-asserted-by":"publisher","DOI":"10.5555\/762761.762767"},{"key":"e_1_2_1_142_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid.2011.83"},{"key":"e_1_2_1_144_1","volume-title":"International Convention on Information 8 Communication Technology Electronics 8 Microelectronics (MIPRO\u201913)","author":"Velkoski Goran","year":"2013","unstructured":"Goran Velkoski , Sasko Ristov , and Marjan Gusev . 2013 . Loosely or tightly coupled affinity for matrix-vector multiplication . In International Convention on Information 8 Communication Technology Electronics 8 Microelectronics (MIPRO\u201913) . 228--233. Goran Velkoski, Sasko Ristov, and Marjan Gusev. 2013. Loosely or tightly coupled affinity for matrix-vector multiplication. In International Convention on Information 8 Communication Technology Electronics 8 Microelectronics (MIPRO\u201913). 228--233."},{"key":"e_1_2_1_145_1","doi-asserted-by":"publisher","DOI":"10.1145\/237090.237205"},{"key":"e_1_2_1_147_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.65"},{"key":"e_1_2_1_148_1","doi-asserted-by":"publisher","DOI":"10.1145\/143369.143377"},{"key":"e_1_2_1_149_1","doi-asserted-by":"publisher","DOI":"10.1145\/1400097.1400102"},{"key":"e_1_2_1_150_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.117"},{"key":"e_1_2_1_151_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2011.49"},{"key":"e_1_2_1_152_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2009.40"},{"key":"e_1_2_1_153_1","doi-asserted-by":"publisher","DOI":"10.1145\/2379776.2379780"},{"key":"e_1_2_1_154_1","doi-asserted-by":"publisher","DOI":"10.1109\/HOTI.2010.24"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3006385","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3006385","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:39:38Z","timestamp":1750217978000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3006385"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,12,5]]},"references-count":138,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2017,12,31]]}},"alternative-id":["10.1145\/3006385"],"URL":"https:\/\/doi.org\/10.1145\/3006385","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,12,5]]},"assertion":[{"value":"2016-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-12-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}