{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:51:09Z","timestamp":1750308669423,"version":"3.41.0"},"reference-count":14,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2012,10,8]],"date-time":"2012-10-08T00:00:00Z","timestamp":1349654400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGMETRICS Perform. Eval. Rev."],"published-print":{"date-parts":[[2012,10,8]]},"abstract":"<jats:p>Multicore multiprocessors use a Non Uniform Memory Architecture (NUMA) to improve their scalability. However, NUMA introduces performance penalties due to remote memory accesses. Without efficiently managing data layout and thread mapping to cores, scientific applications may suffer performance loss, even if they are optimized for NUMA. In this paper, we present algorithms and a runtime system that optimize the execution of OpenMP applications on NUMA architectures. By collecting information from hardware counters, the runtime system directs thread placement and reduces performance penalties by minimizing the critical path of OpenMP parallel regions. The runtime system uses a scalable algorithm that derives placement decisions with negligible overhead. We evaluate our algorithms and the runtime system with four NPB applications implemented in OpenMP. On average the algorithms achieve between 8.13% and 25.68% performance improvement, compared to the default Linux thread placement scheme. The algorithms miss the optimal thread placement in only 8.9% of the cases.<\/jats:p>","DOI":"10.1145\/2381056.2381079","type":"journal-article","created":{"date-parts":[[2012,10,11]],"date-time":"2012-10-11T14:55:16Z","timestamp":1349967316000},"page":"106-112","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["Critical path-based thread placement for NUMA systems"],"prefix":"10.1145","volume":"40","author":[{"given":"ChunYi","family":"Su","sequence":"first","affiliation":[{"name":"Virginia Tech, Blacksburg, VA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dong","family":"Li","sequence":"additional","affiliation":[{"name":"Oak Ridge National Lab, Oak Ridge, TN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dimitrios S.","family":"Nikolopoulos","sequence":"additional","affiliation":[{"name":"FORTH-ICS, Heraklion, Crete, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Matthew","family":"Grove","sequence":"additional","affiliation":[{"name":"Virginia Tech, Blacksburg, VA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kirk","family":"Cameron","sequence":"additional","affiliation":[{"name":"Virginia Tech, Blacksburg, VA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bronis R.","family":"de Supinski","sequence":"additional","affiliation":[{"name":"LLNL, Livermore, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,10,8]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"AMD","author":"AMD.","year":"2010","unstructured":"AMD. BIOS and Kernel Developer's Guide (BKDG) For AMD Family 10h Processors . AMD , 2010 . AMD. BIOS and Kernel Developer's Guide (BKDG) For AMD Family 10h Processors. AMD, 2010."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854350"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1454115.1454151"},{"key":"e_1_2_1_4_1","unstructured":"Klug T. Ott M. Weidendorfer J. Trinitis C. and M\u00fcnchen T. U. autopin -- Automated Optimization of Thread-to-Core Pinning on Multicore Systems.  Klug T. Ott M. Weidendorfer J. Trinitis C. and M\u00fcnchen T. U. autopin -- Automated Optimization of Thread-to-Core Pinning on Multicore Systems."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2010.5470463"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1993478.1993481"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.53"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2010.5452060"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5161108"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/SBAC-PAD.2009.16"},{"volume-title":"Proceedings of the 16th International Euro-Par Conference on Parallel Processing: Part I.","author":"Singh K.","key":"e_1_2_1_11_1","unstructured":"Singh , K. , Curtis-Maury , M. , McKee , S. A. , Blagojevi ?, F., Nikolopoulos , D. S. , de Supinski , B. R. , and Schulz , M . Comparing Scalability Prediction Strategies on an SMP of CMPs . In Proceedings of the 16th International Euro-Par Conference on Parallel Processing: Part I. Singh, K., Curtis-Maury, M., McKee, S. A., Blagojevi?, F., Nikolopoulos, D. S., de Supinski, B. R., and Schulz, M. Comparing Scalability Prediction Strategies on an SMP of CMPs. In Proceedings of the 16th International Euro-Par Conference on Parallel Processing: Part I."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1366219.1366222"},{"key":"e_1_2_1_13_1","first-page":"1","volume-title":"High Performance Computer Architecture (HPCA), 2010 IEEE 16th International Symposium on (Jan.","author":"Ware M.","year":"2010","unstructured":"Ware , M. , Rajamani , K. , Floyd , M. , Brock , B. , Rubio , J. , Rawson , F. , and Carter , J . Architecting for Power Management: The IBM POWER7 Approach . In High Performance Computer Architecture (HPCA), 2010 IEEE 16th International Symposium on (Jan. 2010 ), pp. 1 -- 11 . Ware, M., Rajamani, K., Floyd, M., Brock, B., Rubio, J., Rawson, F., and Carter, J. Architecting for Power Management: The IBM POWER7 Approach. In High Performance Computer Architecture (HPCA), 2010 IEEE 16th International Symposium on (Jan. 2010), pp. 1--11."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1736020.1736036"}],"container-title":["ACM SIGMETRICS Performance Evaluation Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2381056.2381079","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2381056.2381079","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:00:40Z","timestamp":1750276840000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2381056.2381079"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,10,8]]},"references-count":14,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2012,10,8]]}},"alternative-id":["10.1145\/2381056.2381079"],"URL":"https:\/\/doi.org\/10.1145\/2381056.2381079","relation":{},"ISSN":["0163-5999"],"issn-type":[{"type":"print","value":"0163-5999"}],"subject":[],"published":{"date-parts":[[2012,10,8]]},"assertion":[{"value":"2012-10-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}