{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:07:40Z","timestamp":1750306060812,"version":"3.41.0"},"reference-count":36,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2017,5,26]],"date-time":"2017-05-26T00:00:00Z","timestamp":1495756800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"RoMoL ERC Advanced","award":["321253"],"award-info":[{"award-number":["321253"]}]},{"name":"Royal Society Newton International Fellowship"},{"name":"Management of University and Research Grants","award":["AGAUR - FI-DGR 2014"],"award-info":[{"award-number":["AGAUR - FI-DGR 2014"]}]},{"DOI":"10.13039\/501100000780","name":"European Union","doi-asserted-by":"crossref","award":["TIN2015-65316-P"],"award-info":[{"award-number":["TIN2015-65316-P"]}],"id":[{"id":"10.13039\/501100000780","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2017,6,30]]},"abstract":"<jats:p>In the low-end mobile processor market, power, energy, and area budgets are significantly lower than in the server\/desktop\/laptop\/high-end mobile markets. It has been shown that vector processors are a highly energy-efficient way to increase performance; however, adding support for them incurs area and power overheads that would not be acceptable for low-end mobile processors. In this work, we propose an integrated vector-scalar design for the ARM architecture that mostly reuses scalar hardware to support the execution of vector instructions. The key element of the design is our proposed block-based model of execution that groups vector computational instructions together to execute them in a coordinated manner. We implemented a classic vector unit and compare its results against our integrated design. Our integrated design improves the performance (more than 6\u00d7) and energy consumption (up to 5\u00d7) of a scalar in-order core with negligible area overhead (only 4.7% when using a vector register with 32 elements). In contrast, the area overhead of the classic vector unit can be significant (around 44%) if a dedicated vector floating-point unit is incorporated. Our block-based vector execution outperforms the classic vector unit for all kernels with floating-point data and also consumes less energy. We also complement the integrated design with three energy\/performance-efficient techniques that further reduce power and increase performance. The first proposal covers the design and implementation of chaining logic that is optimized to work with the cache hierarchy through vector memory instructions, the second proposal reduces the number of reads\/writes from\/to the vector register file, and the third idea optimizes complex memory access patterns with the memory shape instruction and unified indexed vector load.<\/jats:p>","DOI":"10.1145\/3075618","type":"journal-article","created":{"date-parts":[[2017,5,31]],"date-time":"2017-05-31T19:32:40Z","timestamp":1496259160000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["An Integrated Vector-Scalar Design on an In-Order ARM Core"],"prefix":"10.1145","volume":"14","author":[{"given":"Milan","family":"Stanic","sequence":"first","affiliation":[{"name":"Barcelona Supercomputing Center, Veldhoven, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Oscar","family":"Palomar","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Timothy","family":"Hayes","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ivan","family":"Ratkovic","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adrian","family":"Cristal","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Osman","family":"Unsal","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mateo","family":"Valero","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center, Veldhoven, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,5,26]]},"reference":[{"volume-title":"University of California","author":"Asanovic Krste","key":"e_1_2_2_1_1","unstructured":"Krste Asanovic . May 1998. Vector Microprocessors . Ph.D. Dissertation . University of California , Berkeley . Krste Asanovic. May 1998. Vector Microprocessors. Ph.D. Dissertation. University of California, Berkeley."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2024716.2024718"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/956417.956540"},{"key":"e_1_2_2_5_1","volume-title":"PTX: Parallel thread execution ISA version 2.3. 1.","author":"Compute NVIDIA","year":"2010","unstructured":"NVIDIA Compute . 2010 . PTX: Parallel thread execution ISA version 2.3. 1. NVIDIA Compute. 2010. PTX: Parallel thread execution ISA version 2.3. 1."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.5555\/545215.545247"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/525424.822650"},{"key":"e_1_2_2_8_1","unstructured":"Nadeem Firasta Mark Buxton Paula Jinbo Kaveh Nasri and Shihjong Kuo. 2008. Intel AVX: New Frontiers in Performance Improvements and Energy Efficiency. White Paper. (2008).  Nadeem Firasta Mark Buxton Paula Jinbo Kaveh Nasri and Shihjong Kuo. 2008. Intel AVX: New Frontiers in Performance Improvements and Energy Efficiency. White Paper. (2008)."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/2014698.2014884"},{"key":"e_1_2_2_11_1","volume-title":"Proceedings of the 22nd International Conference on Parallel Architectures and Compilation Techniques (PACT\u201913)","author":"Govindaraju Venkatraman","year":"2013","unstructured":"Venkatraman Govindaraju , Tony Nowatzki , and Karthikeyan Sankaralingam . 2013 . Breaking simd shackles: Liberating accelerators by exposing flexible microarchitectural mechanisms . In Proceedings of the 22nd International Conference on Parallel Architectures and Compilation Techniques (PACT\u201913) . Venkatraman Govindaraju, Tony Nowatzki, and Karthikeyan Sankaralingam. 2013. Breaking simd shackles: Liberating accelerators by exposing flexible microarchitectural mechanisms. In Proceedings of the 22nd International Conference on Parallel Architectures and Compilation Techniques (PACT\u201913)."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1128022.1128023"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1147\/JRD.2016.2527418"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2006.41"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2155620.2155623"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.24"},{"key":"e_1_2_2_17_1","volume-title":"Hennessy and David A Patterson","author":"John","year":"2011","unstructured":"John L. Hennessy and David A Patterson . 2011 . Computer Architecture : A Quantitative Approach (5th ed.). Elsevier . John L. Hennessy and David A Patterson. 2011. Computer Architecture: A Quantitative Approach (5th ed.). Elsevier."},{"key":"e_1_2_2_18_1","volume-title":"Proceedings of WCED in Conjunction with ISCA","volume":"27","author":"Hu Zhigang","year":"2000","unstructured":"Zhigang Hu and Margaret Martonosi . 2000 . Reducing register file power consumption by exploiting value lifetime . In Proceedings of WCED in Conjunction with ISCA , Vol. 27 . Zhigang Hu and Margaret Martonosi. 2000. Reducing register file power consumption by exploiting value lifetime. In Proceedings of WCED in Conjunction with ISCA, Vol. 27."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/40.918001"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/859618.859664"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/998680.1006736"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2000064.2000080"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1188455.1188537"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669172"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-07518-1_13"},{"key":"e_1_2_2_26_1","volume-title":"Ang","author":"Murphy Richard C.","year":"2010","unstructured":"Richard C. Murphy , Kyle B. Wheeler , Brian W. Barrett , and James A . Ang . 2010 . Introducing the graph 500. Cray Users Group (CUG) ( 2010). Richard C. Murphy, Kyle B. Wheeler, Brian W. Barrett, and James A. Ang. 2010. Introducing the graph 500. Cray Users Group (CUG) (2010)."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/942806.943844"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/305138.305148"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/359327.359336"},{"volume-title":"Scientific Computing on Vector Computers","author":"Sch\u00f6nauer Willi","key":"e_1_2_2_30_1","unstructured":"Willi Sch\u00f6nauer . 1987. Scientific Computing on Vector Computers . Elsevier Science . Willi Sch\u00f6nauer. 1987. Scientific Computing on Vector Computers. Elsevier Science."},{"key":"e_1_2_2_31_1","volume-title":"ARM Architecture Reference Manual","author":"Seal David","unstructured":"David Seal . 2000. ARM Architecture Reference Manual ( 2 nd ed.). Addison-Wesley Longman Publishing Co. , Boston, MA . David Seal. 2000. ARM Architecture Reference Manual (2nd ed.). Addison-Wesley Longman Publishing Co., Boston, MA.","edition":"2"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2009.9"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2015.7477467"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICENCO.2011.6153927"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.31"},{"volume-title":"Xeon Phi. In Proceedings of the 2014 International Conference on High Performance Computing 8 Simulation (HPCS\u201914)","author":"Stanic Milan","key":"e_1_2_2_36_1","unstructured":"Milan Stanic , Oscar Palomar , Ivan Ratkovic , Milovan Duric , Osman Unsal , Adrian Cristal , and M. R. Valero . 2014. Evaluation of vectorization potential of Graph500 on Intel\u2019s Xeon Phi. In Proceedings of the 2014 International Conference on High Performance Computing 8 Simulation (HPCS\u201914) . IEEE, 47--54. Milan Stanic, Oscar Palomar, Ivan Ratkovic, Milovan Duric, Osman Unsal, Adrian Cristal, and M. R. Valero. 2014. Evaluation of vectorization potential of Graph500 on Intel\u2019s Xeon Phi. In Proceedings of the 2014 International Conference on High Performance Computing 8 Simulation (HPCS\u201914). IEEE, 47--54."},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.809248"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/846215.846797"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3075618","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3075618","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:03:42Z","timestamp":1750215822000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3075618"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,5,26]]},"references-count":36,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2017,6,30]]}},"alternative-id":["10.1145\/3075618"],"URL":"https:\/\/doi.org\/10.1145\/3075618","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2017,5,26]]},"assertion":[{"value":"2016-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-03-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-05-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}