{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,26]],"date-time":"2025-09-26T13:12:31Z","timestamp":1758892351919,"version":"3.41.0"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2015,9,8]],"date-time":"2015-09-08T00:00:00Z","timestamp":1441670400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004963","name":"Seventh Framework Programme","doi-asserted-by":"publisher","award":["288653"],"award-info":[{"award-number":["288653"]}],"id":[{"id":"10.13039\/501100004963","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2015,10,6]]},"abstract":"<jats:p>Temporal SIMT (TSIMT) has been suggested as an alternative to conventional (spatial) SIMT for improving GPU performance on branch-intensive code. Although TSIMT has been briefly mentioned before, it was not evaluated. We present a complete design and evaluation of TSIMT GPUs, along with the inclusion of scalarization and a combination of temporal and spatial SIMT, named Spatiotemporal SIMT (STSIMT). Simulations show that TSIMT alone results in a performance reduction, but a combination of scalarization and STSIMT yields a mean performance enhancement of 19.6% and improves the energy-delay product by 26.2% compared to SIMT.<\/jats:p>","DOI":"10.1145\/2811402","type":"journal-article","created":{"date-parts":[[2015,9,15]],"date-time":"2015-09-15T12:09:15Z","timestamp":1442318955000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Spatiotemporal SIMT and Scalarization for Improving GPU Efficiency"],"prefix":"10.1145","volume":"12","author":[{"given":"Jan","family":"Lucas","sequence":"first","affiliation":[{"name":"Technische Universit\u00e4t Berlin, Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"Andersch","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Berlin, Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mauricio","family":"Alvarez-Mesa","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Berlin, Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ben","family":"Juurlink","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Berlin, Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,9,8]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1383422.1383443"},{"volume-title":"AMD Graphics Core Next GCN Architecture White Paper","author":"MD.","key":"e_1_2_2_2_1","unstructured":"A MD. 2012. AMD Graphics Core Next GCN Architecture White Paper . Sunnyvale, CA . AMD. 2012. AMD Graphics Core Next GCN Architecture White Paper. Sunnyvale, CA."},{"volume-title":"Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS).","author":"Bakhoda A.","key":"e_1_2_2_3_1","unstructured":"A. Bakhoda , G. L. Yuan , W. W. L. Fung , H. Wong , and T. M. Aamodt . 2009. Analyzing CUDA workloads using a detailed GPU simulator . In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). A. Bakhoda, G. L. Yuan, W. W. L. Fung, H. Wong, and T. M. Aamodt. 2009. Analyzing CUDA workloads using a detailed GPU simulator. In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)."},{"volume-title":"Proceedings of the 39th International Symposium on Computer Architecture (ISCA).","author":"Brunie N.","key":"e_1_2_2_4_1","unstructured":"N. Brunie , S. Collange , and G. Diamos . 2012. Simultaneous branch and warp interweaving for sustained GPU performance . In Proceedings of the 39th International Symposium on Computer Architecture (ISCA). N. Brunie, S. Collange, and G. Diamos. 2012. Simultaneous branch and warp interweaving for sustained GPU performance. In Proceedings of the 39th International Symposium on Computer Architecture (ISCA)."},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.63"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.13"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854318"},{"volume-title":"Proceedings of the 17th International Symposium on High Performance Computer Architecture (HPCA).","author":"Fung W. W. L.","key":"e_1_2_2_10_1","unstructured":"W. W. L. Fung and T. M. Aamodt . 2011. Thread block compaction for efficient SIMT control flow . In Proceedings of the 17th International Symposium on High Performance Computer Architecture (HPCA). W. W. L. Fung and T. M. Aamodt. 2011. Thread block compaction for efficient SIMT control flow. In Proceedings of the 17th International Symposium on High Performance Computer Architecture (HPCA)."},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.12"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1543753.1543756"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.18"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2038037.1941590"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2004.10007"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2011.89"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485934"},{"key":"e_1_2_2_18_1","volume-title":"Patent No. US 2013\/0042090 A1. Filed","author":"Krashinsky R. M.","year":"2011","unstructured":"R. M. Krashinsky . 2011 . Temporal SIMT Execution Optimization. (Aug. 2011) . Patent No. US 2013\/0042090 A1. Filed August 12, 2011, Issued February 14, 2013. R. M. Krashinsky. 2011. Temporal SIMT Execution Optimization. (Aug. 2011). Patent No. US 2013\/0042090 A1. Filed August 12, 2011, Issued February 14, 2013."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2000064.2000080"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2013.6494995"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2008.31"},{"key":"e_1_2_2_22_1","unstructured":"J. E. Lindholm M. Y. Siu S. S. Moy S. Liu and J. R. Nickolls. 2008b. Simulating Multiported Memories using Lower Port Count Memories. Patent No. US 7339592 B2 Filed July 2004 Issued March 2008.  J. E. Lindholm M. Y. Siu S. S. Moy S. Liu and J. R. Nickolls. 2008b. Simulating Multiported Memories using Lower Port Count Memories. Patent No. US 7339592 B2 Filed July 2004 Issued March 2008."},{"key":"e_1_2_2_23_1","unstructured":"A. Lumsdaine and D. Gregor. 2004. Boost Graph Library: Sequential Vertex Coloring. http:\/\/www.boost.org\/doc\/libs\/1_57_0\/libs\/graph\/doc\/sequential_vertex_coloring.html.  A. Lumsdaine and D. Gregor. 2004. Boost Graph Library: Sequential Vertex Coloring. http:\/\/www.boost.org\/doc\/libs\/1_57_0\/libs\/graph\/doc\/sequential_vertex_coloring.html."},{"volume-title":"Advanced Compiler Design and Implementation. Morgan Kaufmann","author":"Muchnick S. S.","key":"e_1_2_2_24_1","unstructured":"S. S. Muchnick . 1997. Advanced Compiler Design and Implementation. Morgan Kaufmann , Burlington, MA . S. S. Muchnick. 1997. Advanced Compiler Design and Implementation. Morgan Kaufmann, Burlington, MA."},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2155620.2155656"},{"key":"e_1_2_2_26_1","unstructured":"NVIDIA. 2011. NVidia GPU Computing SDK 3.1.  NVIDIA. 2011. NVidia GPU Computing SDK 3.1."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/339647.339693"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485954"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-36949-0_18"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCST.2016.2557221"},{"volume-title":"Proceedings of the IEEE International Symposium on Performance Analysis of Systems Software (ISPASS).","author":"Wong H.","key":"e_1_2_2_31_1","unstructured":"H. Wong , M.-M. Papadopoulou , M. Sadooghi-Alvandi , and A. Moshovos . 2010. Demystifying GPU microarchitecture through microbenchmarking . In Proceedings of the IEEE International Symposium on Performance Analysis of Systems Software (ISPASS). H. Wong, M.-M. Papadopoulou, M. Sadooghi-Alvandi, and A. Moshovos. 2010. Demystifying GPU microarchitecture through microbenchmarking. In Proceedings of the IEEE International Symposium on Performance Analysis of Systems Software (ISPASS)."},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2464996.2465022"},{"volume-title":"Proceedings of the 20th International Symposium on High Performance Computer Architecture (HPCA).","author":"Xiang P.","key":"e_1_2_2_33_1","unstructured":"P. Xiang , Y. Yang , and H. Zhou . 2014. Warp-level divergence in GPUs: Characterization, impact, and mitigation . In Proceedings of the 20th International Symposium on High Performance Computer Architecture (HPCA). P. Xiang, Y. Yang, and H. Zhou. 2014. Warp-level divergence in GPUs: Characterization, impact, and mitigation. In Proceedings of the 20th International Symposium on High Performance Computer Architecture (HPCA)."},{"key":"e_1_2_2_34_1","volume-title":"Retrieved","author":"Ziegler G.","year":"2011","unstructured":"G. Ziegler . 2011 . Analysis-Driven Optimization . Retrieved August 17, 2015 from http:\/\/www.nvidia.de\/content\/PDF\/isc-2011\/Ziegler.pdf. G. Ziegler. 2011. Analysis-Driven Optimization. Retrieved August 17, 2015 from http:\/\/www.nvidia.de\/content\/PDF\/isc-2011\/Ziegler.pdf."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2811402","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2811402","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T05:48:49Z","timestamp":1750225729000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2811402"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,9,8]]},"references-count":33,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2015,10,6]]}},"alternative-id":["10.1145\/2811402"],"URL":"https:\/\/doi.org\/10.1145\/2811402","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2015,9,8]]},"assertion":[{"value":"2015-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-09-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}