{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T11:01:38Z","timestamp":1783076498687,"version":"3.54.6"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2018,9,4]],"date-time":"2018-09-04T00:00:00Z","timestamp":1536019200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2018,9,30]]},"abstract":"<jats:p>To exploit parallelism and scalability of multiple GPUs in a system, it is critical to place compute and data together. However, two key techniques that have been used to hide memory latency and improve thread-level parallelism (TLP), memory interleaving, and thread block scheduling, in traditional GPU systems are at odds with efficient use of multiple GPUs. Distributing data across multiple GPUs to improve overall memory bandwidth utilization incurs high remote traffic when the data and compute are misaligned. Nondeterministic thread block scheduling to improve compute resource utilization impedes co-placement of compute and data. Our goal in this work is to enable co-placement of compute and data in the presence of fine-grained interleaved memory with a low-cost approach.<\/jats:p>\n          <jats:p>To this end, we propose a mechanism that identifies exclusively accessed data and place the data along with the thread block that accesses it in the same GPU. The key ideas are (1) the amount of data exclusively used by a thread block can be estimated, and that exclusive data (of any size) can be localized to one GPU with coarse-grained interleaved pages; (2) using the affinity-based thread block scheduling policy, we can co-place compute and data together; and (3) by using dual address mode with lightweight changes to virtual to physical page mappings, we can selectively choose different interleaved memory pages for each data structure. Our evaluations across a wide range of workloads show that the proposed mechanism improves performance by 31% and reduces 38% remote traffic over a baseline system.<\/jats:p>","DOI":"10.1145\/3232521","type":"journal-article","created":{"date-parts":[[2018,9,4]],"date-time":"2018-09-04T12:37:30Z","timestamp":1536064650000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":24,"title":["CODA"],"prefix":"10.1145","volume":"15","author":[{"given":"Hyojong","family":"Kim","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8731-1084","authenticated-orcid":false,"given":"Ramyad","family":"Hadidi","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8801-9384","authenticated-orcid":false,"given":"Lifeng","family":"Nai","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hyesoon","family":"Kim","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nuwan","family":"Jayasena","sequence":"additional","affiliation":[{"name":"Advanced Micro Devices, Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yasuko","family":"Eckert","sequence":"additional","affiliation":[{"name":"Advanced Micro Devices, Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Onur","family":"Kayiran","sequence":"additional","affiliation":[{"name":"Advanced Micro Devices, Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gabriel","family":"Loh","sequence":"additional","affiliation":[{"name":"Advanced Micro Devices, Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2018,9,4]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750386"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750385"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750397"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080231"},{"key":"e_1_2_1_5_1","volume-title":"High Bandwidth Memory (HBM) DRAM. JESD235A (November","author":"JEDEC Solid State Technology Association","year":"2015"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3123975"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628109"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/195473.195485"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the 1st Workshop on Near-Data Processing (WoNDP\u201913)","author":"Chu Michael","year":"2013"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the IEEE Custom Integrated Circuits Conference","author":"Elliott Duncan","year":"1992"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/54.748803"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.12"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2989081.2989102"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.375174"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the 20th International Conference on Compiler Construction: Part of the Joint European Conferences on Theory and Practice of Software (CC\/ETAPS\u201911)","author":"Grewe Dominik"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2017.8167757"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2018.00018"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3155287"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/331532.331589"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555779"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.27"},{"key":"e_1_2_1_23_1","unstructured":"Intel Corporation. 2007. Intel\u00ae64 and IA-32 Architectures Software Developer's Manual.  Intel Corporation. 2007. Intel\u00ae64 and IA-32 Architectures Software Developer's Manual."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD.2012.6378608"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.5555\/2523721.2523744"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.55"},{"key":"e_1_2_1_27_1","unstructured":"Hyesoon Kim Jaekyu Lee Nagesh B. Lakshminarayana Jaewoong Sim Jieun Lim Tri Pho Hyojong Kim and Ramyad Hadidi. 2012. MacSim: A CPU-GPU Heterogeneous Simulation Framework User Guide.  Hyesoon Kim Jaekyu Lee Nagesh B. Lakshminarayana Jaewoong Sim Jieun Lim Tri Pho Hyojong Kim and Ramyad Hadidi. 2012. MacSim: A CPU-GPU Heterogeneous Simulation Framework User Guide."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1941553.1941591"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/977395.977673"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/2523721.2523756"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2798725"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2008.31"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669121"},{"key":"e_1_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Richard C. Murphy Peter M. Kogge and Arun Rodrigues. 2000. The characterization of data intensive memory workloads on distributed PIM systems. In Revised Papers from the Second International Workshop on Intelligent Memory Systems (IMS\u201900).   Richard C. Murphy Peter M. Kogge and Arun Rodrigues. 2000. The characterization of data intensive memory workloads on distributed PIM systems. In Revised Papers from the Second International Workshop on Intelligent Memory Systems (IMS\u201900).","DOI":"10.1007\/3-540-44570-6_6"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.54"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2018.00077"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807626"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1147\/JRD.2015.2409732"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2155620.2155656"},{"key":"e_1_2_1_40_1","unstructured":"NVIDIA Corp. 2009. NVIDIA\u2019s Next Generation CUDA\u2122 Compute Architecture: Fermi\u2122. Retrieved from https:\/\/www.nvidia.com\/content\/PDF\/fermi_white_papers\/NVIDIA_Fermi_Compute_Architecture_Whitepaper.pdf.  NVIDIA Corp. 2009. NVIDIA\u2019s Next Generation CUDA\u2122 Compute Architecture: Fermi\u2122. Retrieved from https:\/\/www.nvidia.com\/content\/PDF\/fermi_white_papers\/NVIDIA_Fermi_Compute_Architecture_Whitepaper.pdf."},{"key":"e_1_2_1_41_1","unstructured":"NVIDIA Corp. 2016. NVIDIA Tesla P100. Retrieved from https:\/\/images.nvidia.com\/content\/pdf\/tesla\/whitepaper\/pascal-architecture-whitepaper.pdf.  NVIDIA Corp. 2016. NVIDIA Tesla P100. Retrieved from https:\/\/images.nvidia.com\/content\/pdf\/tesla\/whitepaper\/pascal-architecture-whitepaper.pdf."},{"key":"e_1_2_1_42_1","unstructured":"NVIDIA Corp. 2017. NVIDIA Tesla V100. Retrieved from http:\/\/www.nvidia.com\/object\/volta-architecture-whitepaper.html.  NVIDIA Corp. 2017. NVIDIA Tesla V100. Retrieved from http:\/\/www.nvidia.com\/object\/volta-architecture-whitepaper.html."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/279358.279387"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.1997.585348"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/40.592312"},{"key":"e_1_2_1_46_1","volume-title":"Proceedings of the 25th USENIX Security Symposium (USENIX Security\u201916)","author":"Pessl Peter","year":"2016"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541942"},{"key":"e_1_2_1_48_1","volume-title":"Proceedings of the 2014 IEEE 20th International Symposium on High Performance Computer Architecture (HPCA\u201914)","author":"Power Jason"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2014.6844483"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2544100"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/174223.158909"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/1964218.1964225"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/L-CA.2011.4"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854336"},{"key":"e_1_2_1_56_1","volume-title":"Proceedings of the 2016 Design, Automation Test in Europe Conference Exhibition (DATE\u201916)","author":"Thanh-Hoang Tung"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2600212.2600213"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/2492408.2492418"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/360128.360134"},{"key":"e_1_2_1_60_1","volume-title":"Proceedings of the 2016 IEEE 22nd International Symposium on High Performance Computer Architecture (HPCA\u201916)","author":"Zheng Tianhao"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/2996190"},{"key":"e_1_2_1_62_1","volume-title":"Lemmon","author":"Zurawski John H.","year":"1995"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3232521","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3232521","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:39:42Z","timestamp":1750210782000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3232521"}},"subtitle":["Enabling Co-location of Computation and Data for Multiple GPU Systems"],"short-title":[],"issued":{"date-parts":[[2018,9,4]]},"references-count":61,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2018,9,30]]}},"alternative-id":["10.1145\/3232521"],"URL":"https:\/\/doi.org\/10.1145\/3232521","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,9,4]]},"assertion":[{"value":"2018-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-09-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}