{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T12:09:40Z","timestamp":1763726980191,"version":"3.41.0"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2016,1,4]],"date-time":"2016-01-04T00:00:00Z","timestamp":1451865600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Research Project of National University of Defense Technology","award":["GC-14-06-02"],"award-info":[{"award-number":["GC-14-06-02"]}]},{"DOI":"10.13039\/501100001809","name":"National Science Foundation of China","doi-asserted-by":"crossref","award":["61402493 and 61433007"],"award-info":[{"award-number":["61402493 and 61433007"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2016,1,7]]},"abstract":"<jats:p>The efficacy of single instruction, multiple data (SIMD) architectures is limited when handling divergent control flows. This circumstance results in SIMD fragments using only a subset of the available lanes. We propose an iteration interleaving--based SIMD lane partition (IISLP) architecture that interleaves the execution of consecutive iterations and dynamically partitions SIMD lanes into branch paths with comparable execution time. The benefits are twofold: SIMD fragments under divergent branches can execute in parallel, and the pathology of fragment starvation can also be well eliminated. Our experiments show that IISLP doubles the performance of a baseline mechanism and provides a speedup of 28% versus instruction shuffle.<\/jats:p>","DOI":"10.1145\/2847253","type":"journal-article","created":{"date-parts":[[2016,1,7]],"date-time":"2016-01-07T14:04:54Z","timestamp":1452175494000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Iteration Interleaving--Based SIMD Lane Partition"],"prefix":"10.1145","volume":"12","author":[{"given":"Yaohua","family":"Wang","sequence":"first","affiliation":[{"name":"National University of Defense Technology, Hunan Province, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dong","family":"Wang","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Hunan Province, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuming","family":"Chen","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Hunan Province, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zonglin","family":"Liu","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Hunan Province, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shenggang","family":"Chen","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Hunan Province, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaowen","family":"Chen","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Hunan Province, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xu","family":"Zhou","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Hunan Province, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,1,4]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/PROC.1972.8647"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2013.129"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1147\/rd.515.0559"},{"volume-title":"Proceedings of the 2014 IEEE 20th International Symposium on High Performance Computing Architecture (HPCA-20)","author":"Tantawy A.","key":"e_1_2_2_4_1"},{"volume-title":"Proceedings of the 2001 IEEE 17th International Symposium on High Performance Computer Architecture (HPCA\u201911)","author":"Fung W. W. L.","key":"e_1_2_2_5_1"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.12"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1543753.1543756"},{"key":"e_1_2_2_8_1","unstructured":"Q. He. 2006. The Principle of the Computer Graphics. Tsinghua University Press Beijing China.  Q. He. 2006. The Principle of the Computer Graphics. Tsinghua University Press Beijing China."},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485952"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/360128.360145"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/40.918001"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2004.90"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2000064.2000080"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815992"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2155620.2155656"},{"key":"e_1_2_2_16_1","unstructured":"NVIDIA Corporation. 2008. GeForce Gtx 280 Specifications. Available at http:\/\/www.geforce.com\/hardware\/desktop-gpus\/geforce-gtx-280\/specifications.  NVIDIA Corporation. 2008. GeForce Gtx 280 Specifications. Available at http:\/\/www.geforce.com\/hardware\/desktop-gpus\/geforce-gtx-280\/specifications."},{"key":"e_1_2_2_17_1","unstructured":"NVIDIA Corporation. 2009. Nvidia's Next Generation CUDA Compute Architecture: Fermi. Available at http:\/\/www.nvidia.com.  NVIDIA Corporation. 2009. Nvidia's Next Generation CUDA Compute Architecture: Fermi. Available at http:\/\/www.nvidia.com."},{"key":"e_1_2_2_18_1","unstructured":"NVIDIA Corporation. 2012. Nvidia's Next Generation CUDA Compute Architecture: Kepler GK110. Available at http:\/\/www.nvidia.com.  NVIDIA Corporation. 2012. Nvidia's Next Generation CUDA Compute Architecture: Kepler GK110. Available at http:\/\/www.nvidia.com."},{"key":"e_1_2_2_19_1","unstructured":"OPCODE. 2003. OPCODE Optimized Collision Detection Library (OPCODE). Retrieved December 7 2015 from http:\/\/www.codercorner.com\/Opcode.htm.  OPCODE. 2003. OPCODE Optimized Collision Detection Library (OPCODE). Retrieved December 7 2015 from http:\/\/www.codercorner.com\/Opcode.htm."},{"volume-title":"Proceedings of the 40th International Symposium on Computer Architecture (ISCA\u201912)","author":"Rhu M.","key":"e_1_2_2_20_1"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485953"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2013.6522352"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2006.74"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1399504.1360617"},{"key":"e_1_2_2_25_1","unstructured":"K. Suhring. 2015 H.264 Joint Model (JM)-h.264\/AVC Reference Software. Available at http:\/\/iphome.hhi.de\/suehring\/tml\/.  K. Suhring. 2015 H.264 Joint Model (JM)-h.264\/AVC Reference Software. Available at http:\/\/iphome.hhi.de\/suehring\/tml\/."},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2013.6522353"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/L-CA.2011.34"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2010.8"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1250662.1250689"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2847253","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2847253","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T05:48:23Z","timestamp":1750225703000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2847253"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,1,4]]},"references-count":29,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2016,1,7]]}},"alternative-id":["10.1145\/2847253"],"URL":"https:\/\/doi.org\/10.1145\/2847253","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2016,1,4]]},"assertion":[{"value":"2015-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-01-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}