{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T15:30:05Z","timestamp":1781883005950,"version":"3.54.5"},"reference-count":34,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2018,3,22]],"date-time":"2018-03-22T00:00:00Z","timestamp":1521676800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["1618509"],"award-info":[{"award-number":["1618509"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2018,3,31]]},"abstract":"<jats:p>Graphics Processing Units (GPUs) leverage massive thread-level parallelism (TLP) to achieve high computation throughput and hide long memory latency. However, recent studies have shown that the GPU performance does not scale with the GPU occupancy or the degrees of TLP that a GPU supports, especially for memory-intensive workloads. The current understanding points to L1 D-cache contention or off-chip memory bandwidth. In this article, we perform a novel scalability analysis from the perspective of throughput utilization of various GPU components, including off-chip DRAM, multiple levels of caches, and the interconnect between L1 D-caches and L2 partitions. We show that the interconnect bandwidth is a critical bound for GPU performance scalability.<\/jats:p>\n          <jats:p>For the applications that do not have saturated throughput utilization on a particular resource, their performance scales well with increased TLP. To improve TLP for such applications efficiently, we propose a fast context switching approach. When a warp\/thread block (TB) is stalled by a long latency operation, the context of the warp\/TB is spilled to spare on-chip resource so that a new warp\/TB can be launched. The switched-out warp\/TB is switched back when another warp\/TB is completed or switched out. With this fine-grain fast context switching, higher TLP can be supported without increasing the sizes of critical resources like the register file. Our experiment shows that the performance can be improved by up to 47% and a geometric mean of 22% for a set of applications with unsaturated throughput utilization. Compared to the state-of-the-art TLP improvement scheme, our proposed scheme achieves 12% higher performance on average and 16% for unsaturated benchmarks.<\/jats:p>","DOI":"10.1145\/3177964","type":"journal-article","created":{"date-parts":[[2018,3,23]],"date-time":"2018-03-23T12:29:49Z","timestamp":1521808189000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["GPU Performance vs. Thread-Level Parallelism"],"prefix":"10.1145","volume":"15","author":[{"given":"Zhen","family":"Lin","sequence":"first","affiliation":[{"name":"North Carolina State University, Raleigh, North Carolina"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michael","family":"Mantor","sequence":"additional","affiliation":[{"name":"Advanced Micro Devices, Orlando, Florida"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huiyang","family":"Zhou","sequence":"additional","affiliation":[{"name":"North Carolina State University, Orlando, Florida"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2018,3,22]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"OpenCL\u2014The open standard for parallel programming","year":"2010","unstructured":"2010. OpenCL\u2014The open standard for parallel programming . KHRONOS Group ( 2010 ). 2010. OpenCL\u2014The open standard for parallel programming. KHRONOS Group (2010)."},{"key":"e_1_2_1_2_1","volume-title":"CCUDA C programming guide","year":"2011","unstructured":"2011. CCUDA C programming guide . NVIDIA Corporation ( 2011 ). 2011. CCUDA C programming guide. NVIDIA Corporation (2011)."},{"key":"e_1_2_1_3_1","unstructured":"2011. CUDA C\/C++ SDK CODE samples. NVIDIA Corporation (2011).  2011. CUDA C\/C++ SDK CODE samples. NVIDIA Corporation (2011)."},{"key":"e_1_2_1_4_1","volume-title":"AMD GRAPHICS CORES NEXT (GCN) ARCHITECTURE White Paper","year":"2012","unstructured":"2012. AMD GRAPHICS CORES NEXT (GCN) ARCHITECTURE White Paper . AMD Corporation ( 2012 ). 2012. AMD GRAPHICS CORES NEXT (GCN) ARCHITECTURE White Paper. AMD Corporation (2012)."},{"key":"e_1_2_1_5_1","volume-title":"NVIDIA tesla P100 whitepaper","year":"2016","unstructured":"2016. NVIDIA tesla P100 whitepaper . NVIDIA Corporation ( 2016 ). 2016. NVIDIA tesla P100 whitepaper. NVIDIA Corporation (2016)."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2012.6168946"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the 2016 IEEE International Symposium on Workload Characterization (IISWC\u201916)","author":"Dublish S.","unstructured":"S. Dublish , V. Nagarajan , and N. Topham . 2016. Characterizing memory bottlenecks in GPGPU workloads . In Proceedings of the 2016 IEEE International Symposium on Workload Characterization (IISWC\u201916) . 1--2. S. Dublish, V. Nagarajan, and N. Topham. 2016. Characterizing memory bottlenecks in GPGPU workloads. In Proceedings of the 2016 IEEE International Symposium on Workload Characterization (IISWC\u201916). 1--2."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1543753.1543756"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.18"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the Conference on Innovative Parallel Computing (InPar\u201912)","author":"Grauer-Gray S.","unstructured":"S. Grauer-Gray , L. Xu , R. Searles , S. Ayalasomayajula , and J. Cavazos . 2012. Auto-tuning a high-level language targeted to GPU codes . In Proceedings of the Conference on Innovative Parallel Computing (InPar\u201912) . 1--10. S. Grauer-Gray, L. Xu, R. Searles, S. Ayalasomayajula, and J. Cavazos. 2012. Auto-tuning a high-level language targeted to GPU codes. In Proceedings of the Conference on Innovative Parallel Computing (InPar\u201912). 1--10."},{"key":"e_1_2_1_12_1","volume-title":"Presented as Part of the 4th USENIX Workshop on Hot Topics in Parallelism. USENIX","author":"Gregg Chris","unstructured":"Chris Gregg , Jonathan Dorn , Kim Hazelwood , and Kevin Skadron . 2012. Fine-grained resource sharing for concurrent GPGPU kernels . In Presented as Part of the 4th USENIX Workshop on Hot Topics in Parallelism. USENIX , Berkeley, CA . Retrieved from https:\/\/www.usenix.org\/conference\/hotpar12\/fine-grained-resource-sharing-concurrent-gpgpu-kernels. Chris Gregg, Jonathan Dorn, Kim Hazelwood, and Kevin Skadron. 2012. Fine-grained resource sharing for concurrent GPGPU kernels. In Presented as Part of the 4th USENIX Workshop on Hot Topics in Parallelism. USENIX, Berkeley, CA. Retrieved from https:\/\/www.usenix.org\/conference\/hotpar12\/fine-grained-resource-sharing-concurrent-gpgpu-kernels."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCOM.1987.1096719"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the 22nd International Conference on Parallel Architectures and Compilation Techniques. 157--166","author":"Kay\u00c4s\u0300ran O.","unstructured":"O. Kay\u00c4s\u0300ran , A. Jog , M. T. Kandemir , and C. R. Das . 2013. Neither more nor less: Optimizing thread-level parallelism for GPGPUs . In Proceedings of the 22nd International Conference on Parallel Architectures and Compilation Techniques. 157--166 . O. Kay\u00c4s\u0300ran, A. Jog, M. T. Kandemir, and C. R. Das. 2013. Neither more nor less: Optimizing thread-level parallelism for GPGPUs. In Proceedings of the 22nd International Conference on Parallel Architectures and Compilation Techniques. 157--166."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2014.6835937"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750418"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628107"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915)","author":"Li D.","unstructured":"D. Li , M. Rhu , D. R. Johnson , M. O\u2019Connor , M. Erez , D. Burger , D. S. Fussell , and S. W. Redder . 2015. Priority-based cache allocation in throughput processors . In Proceedings of the 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915) . 89--100. D. Li, M. Rhu, D. R. Johnson, M. O\u2019Connor, M. Erez, D. Burger, D. S. Fussell, and S. W. Redder. 2015. Priority-based cache allocation in throughput processors. In Proceedings of the 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915). 89--100."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/3014904.3015007"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2830772.2830822"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2925426.2926267"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2015.22"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2016.2549523"},{"key":"e_1_2_1_24_1","volume-title":"Diving Deeper: The Maxwell 2 Memory Crossbar 8 ROP Partitions.","year":"2016","unstructured":"Online:. 2016 . Diving Deeper: The Maxwell 2 Memory Crossbar 8 ROP Partitions. Retrieved from http:\/\/www.anandtech.com\/show\/8935\/geforce-gtx-970-correcting-the-specs-exploring-memory-allocation\/2. Online:. 2016. Diving Deeper: The Maxwell 2 Memory Crossbar 8 ROP Partitions. Retrieved from http:\/\/www.anandtech.com\/show\/8935\/geforce-gtx-970-correcting-the-specs-exploring-memory-allocation\/2."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2451116.2451160"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2694344.2694346"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.16"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915)","author":"Sethia A.","unstructured":"A. Sethia , D. A. Jamshidi , and S. Mahlke . 2015. Mascar: Speeding up GPU warps by reducing memory pitstops . In Proceedings of the 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915) . 174--185. A. Sethia, D. A. Jamshidi, and S. Mahlke. 2015. Mascar: Speeding up GPU warps by reducing memory pitstops. In Proceedings of the 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915). 174--185."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195656"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201916)","author":"Wang Z.","unstructured":"Z. Wang , J. Yang , R. Melhem , B. Childers , Y. Zhang , and M. Guo . 2016. Simultaneous multikernel GPU: Multi-tasking throughput processors via fine-grained sharing . In Proceedings of the 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201916) . 358--369. Z. Wang, J. Yang, R. Melhem, B. Childers, Y. Zhang, and M. Guo. 2016. Simultaneous multikernel GPU: Multi-tasking throughput processors via fine-grained sharing. In Proceedings of the 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201916). 358--369."},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of the 2010 IEEE International Symposium on Performance Analysis of Systems Software (ISPASS\u201910)","author":"Wong H.","unstructured":"H. Wong , M. M. Papadopoulou , M. Sadooghi-Alvandi , and A. Moshovos . 2010. Demystifying GPU microarchitecture through microbenchmarking . In Proceedings of the 2010 IEEE International Symposium on Performance Analysis of Systems Software (ISPASS\u201910) . 235--246. H. Wong, M. M. Papadopoulou, M. Sadooghi-Alvandi, and A. Moshovos. 2010. Demystifying GPU microarchitecture through microbenchmarking. In Proceedings of the 2010 IEEE International Symposium on Performance Analysis of Systems Software (ISPASS\u201910). 235--246."},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the 2014 IEEE 20th International Symposium on High Performance Computer Architecture (HPCA\u201914)","author":"Xiang P.","unstructured":"P. Xiang , Y. Yang , and H. Zhou . 2014. Warp-level divergence in GPUs: Characterization, impact, and mitigation . In Proceedings of the 2014 IEEE 20th International Symposium on High Performance Computer Architecture (HPCA\u201914) . 284--295. P. Xiang, Y. Yang, and H. Zhou. 2014. Warp-level divergence in GPUs: Characterization, impact, and mitigation. In Proceedings of the 2014 IEEE 20th International Symposium on High Performance Computer Architecture (HPCA\u201914). 284--295."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.29"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.59"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3177964","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3177964","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3177964","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:02:55Z","timestamp":1750215775000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3177964"}},"subtitle":["Scalability Analysis and a Novel Way to Improve TLP"],"short-title":[],"issued":{"date-parts":[[2018,3,22]]},"references-count":34,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2018,3,31]]}},"alternative-id":["10.1145\/3177964"],"URL":"https:\/\/doi.org\/10.1145\/3177964","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,3,22]]},"assertion":[{"value":"2017-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-03-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}