{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"institution":[{"id":[{"id":"https:\/\/ror.org\/03mb6wj31","id-type":"ROR","asserted-by":"publisher"},{"id":"https:\/\/www.isni.org\/000000041937028X","id-type":"ISNI","asserted-by":"publisher"},{"id":"https:\/\/www.wikidata.org\/entity\/Q1640731","id-type":"wikidata","asserted-by":"publisher"}],"name":"Universitat Polit\u00e8cnica de Catalunya","acronym":["UPC"]}],"indexed":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T20:28:08Z","timestamp":1768940888057,"version":"3.49.0"},"reference-count":0,"publisher":"Universitat Polit\u00e8cnica de Catalunya","license":[{"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"abstract":"<jats:p>The rate of annual data generation grows exponentially. At the same time, there is a high demand to analyze that information quickly. In the past, every processor generation came with a substantial frequency increase, leading to higher application throughput. Nowadays, due to the cease of Dennard scaling, further performance must come from exploiting parallelism.\r\nVector architectures offer an efficient manner, in terms of performance and energy, of exploting parallelism at data-level by means of instructions that operate over multiple elements at the same time. This is popularly known as Single Instruction Multiple Data (SIMD). Traditionally, vector processors were employed to accelerate applications in research, and they were not industry-oriented. However, vector processors are becoming widely used for data processing in multimedia applications, and entering in new application domains such as machine learning and genomics. In this thesis, we study the circumstances that cause inefficiencies in vector processors, and new hardware\/software techniques are proposed to improve the performance and energy consumption of these processors.\r\nWe first analyze the behavior of predicated vector instructions in a real machine. We observe that their execution time is dependent on the vector register length and not on the source mask employed. Therefore, a hardware\/software mechanism is proposed to alleviate this situation, that will have a higher impact in future processors with wider vector register lengths.\r\nWe then study the impact of memory accesses to performance. We identify that an irregular memory access pattern prevents an efficient vectorization, which is automatically discarded by the compiler. For this reason, we propose a near-memory accelerator capable of rearranging data structures and transforming irregular memory accesses to dense ones. This operation may be performed by the devices as the host processor is computing other code regions.\r\nFinally, we observe that many applications with irregular memory access patterns just perform a simple operation on the data before it is evicted back to main memory. In these situations, there is a lack of data access locality, leading to an inefficient use of the memory hierarchy. For this reason, we propose to utilize the accelerators previously described to compute directly near memory.<\/jats:p>\n                <jats:p>La tasa de generaci\u00f3n de informaci\u00f3n aumenta cada a\u00f1o. Al mismo tiempo, existe una alta demanda para analizar dicha informaci\u00f3n en el menor tiempo posible. En el pasado, se recurr\u00eda a aumentar la frecuencia de los procesadores para conseguir una mayor velocidad de procesamiento de los datos. En la actualidad, debido al fin de la ley de Dennard, la frecuencia deja de ser una opci\u00f3n y se apunta al paralelismo como la mejor alternativa. Las arquitecturas vectoriales ofrecen una manera eficiente, en t\u00e9rminos de rendimiento y energ\u00eda, de explotar el paralelismo a nivel de datos a trav\u00e9s de instrucciones que operan sobre m\u00faltiples elementos al mismo tiempo, conocidas popularmente como SIMD. Tradicionalmente, los procesadores vectoriales se utilizaban para acelerar las aplicaciones en la investigaci\u00f3n y no estaban orientados a la industria. Sin embargo, dichos procesadores est\u00e1n siendo cada vez m\u00e1s utilizados para el procesamiento de datos en aplicaciones multimedia. En esta tesis doctoral, se investigan las causas que pueden suponer la ineficiencia de las arquitecturas vectoriales, y se proponen mejoras a nivel de hardware y software con el fin de mejorar el rendimiento y el consumo de estos procesadores. En primer lugar, se estudia el funcionamiento de las instrucciones vectoriales predicadas en una m\u00e1quina real. Como resultado, se observa que el tiempo de ejecuci\u00f3n y el consumo de dichas instrucciones es independiente de la m\u00e1scara empleada, mientras que s\u00ed es dependiente de la longitud de los registros vectoriales que contienen los datos. Por tanto, se propone un mecanismo hardware\/software para aliviar esta situaci\u00f3n, que se agravar\u00e1 en el futuro con la aparici\u00f3n de procesadores con la longitud de los registros vectoriales m\u00e1s alta. En segundo lugar, se analiza el impacto de los accesos a memoria por parte del procesador vectorial. En este caso, se comprueba que un acceso irregular a memoria impide una vectorizaci\u00f3n eficiente de las aplicaciones, que es descartada autom\u00e1ticamente por el compilador. Por tanto, en esta tesis se propone un acelerador cerca de memoria capaz de reordenar los datos y proporcionar accesos secuenciales a memoria mientras el procesador est\u00e1 computando otras regiones de la aplicaci\u00f3n. En tercer lugar, se propone utilizar los aceleradores previamente descritos como elementos de c\u00f3mputo, dado que muchas aplicaciones acceden a memoria de manera irregular para realizar un c\u00f3mputo muy sencillo en el procesador. Este movimiento de datos puede ser evitado si la operaci\u00f3n es realizada cerca de memoria. El rendimiento de estos aceleradores es evaluado en aplicaciones de computaci\u00f3n de altas prestaciones y en grafos, un campo de la ciencia muy afectado por esta situaci\u00f3n.<\/jats:p>","DOI":"10.5821\/dissertation-2117-351088","type":"dissertation","created":{"date-parts":[[2023,7,19]],"date-time":"2023-07-19T01:48:48Z","timestamp":1689731328000},"approved":{"date-parts":[[2021,7,19]]},"source":"Crossref","is-referenced-by-count":0,"title":["Novel techniques to improve the performance and the energy of vector architectures"],"prefix":"10.5821","author":[{"sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adri\u00e1n","family":"Barredo Ferreira","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"3865","container-title":[],"original-title":[],"deposited":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T06:34:20Z","timestamp":1768890860000},"score":1,"resource":{"primary":{"URL":"https:\/\/hdl.handle.net\/2117\/351088"}},"subtitle":[],"editor":[{"given":"Miquel","family":"Moret\u00f3 Planas","sequence":"first","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]},{"given":"Adri\u00e0","family":"Armejach Sanosa","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[null]]},"references-count":0,"URL":"https:\/\/doi.org\/10.5821\/dissertation-2117-351088","relation":{},"subject":[]}}