{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T23:50:33Z","timestamp":1783036233524,"version":"3.54.6"},"reference-count":17,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2009,8]]},"abstract":"<jats:p>The availability of huge system memory, even on standard servers, generated a lot of interest in main memory database engines. In data warehouse systems, highly compressed column-oriented data structures are quite prominent. In order to scale with the data volume and the system load, many of these systems are highly distributed with a shared-nothing approach. The fundamental principle of all systems is a full table scan over one or multiple compressed columns. Recent research proposed different techniques to speedup table scans like intelligent compression or using an additional hardware such as graphic cards or FPGAs. In this paper, we show that utilizing the embedded Vector Processing Units (VPUs) found in standard superscalar processors can speed up the performance of mainmemory full table scan by factors. This is achieved without changing the hardware architecture and thereby without additional power consumption. Moreover, as on-chip VPUs directly access the system's RAM, no additional costly copy operations are needed for using the new SIMD-scan approach in standard main memory database engines. Therefore, we propose this scan approach to be used as the standard scan operator for compressed column-oriented main memory storage. We then discuss how well our solution scales with the number of processor cores; consequently, to what degree it can be applied in multi-threaded environments. To verify the feasibility of our approach, we implemented the proposed techniques on a modern Intel multi-core processor using Intel\u00ae Streaming SIMD Extensions (Intel\u00ae SSE). In addition, we integrated the new SIMD-scan approach into SAP\u00ae Netweaver\u00ae Business Warehouse Accelerator. We conclude with describing the performance benefits of using our approach for processing and scanning compressed data using VPUs in column-oriented main memory database systems.<\/jats:p>","DOI":"10.14778\/1687627.1687671","type":"journal-article","created":{"date-parts":[[2014,6,24]],"date-time":"2014-06-24T12:17:57Z","timestamp":1403612277000},"page":"385-394","source":"Crossref","is-referenced-by-count":159,"title":["SIMD-scan"],"prefix":"10.14778","volume":"2","author":[{"given":"Thomas","family":"Willhalm","sequence":"first","affiliation":[{"name":"Intel GmbH, Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nicolae","family":"Popovici","sequence":"additional","affiliation":[{"name":"Intel GmbH, Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yazan","family":"Boshmaf","sequence":"additional","affiliation":[{"name":"SAP AG, Dietmar-Hopp-Allee, Walldorf, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hasso","family":"Plattner","sequence":"additional","affiliation":[{"name":"University of Potsdam, Potsdam, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexander","family":"Zeier","sequence":"additional","affiliation":[{"name":"University of Potsdam, Potsdam, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jan","family":"Schaffner","sequence":"additional","affiliation":[{"name":"University of Potsdam, Potsdam, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2009,8]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/362084.362137"},{"key":"e_1_2_1_2_1","first-page":"487","article-title":"Performance tradeoffs in read-optimized databases","author":"Harizopoulos S.","year":"2006","journal-title":"VLDB"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/PROC.1966.5273"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.44904"},{"key":"e_1_2_1_5_1","first-page":"22","article-title":"Data Compression and Database Performance","author":"Graefe G.","year":"1991","journal-title":"Applied Computing"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2006.150"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1247480.1247525"},{"key":"e_1_2_1_8_1","first-page":"610","article-title":"Main-memory scan sharing for multi-core CPUs","author":"Qiao L.","year":"2008","journal-title":"VLDB"},{"key":"e_1_2_1_9_1","first-page":"622","article-title":"Row-wise parallel predicate evaluation","author":"Johnson R.","year":"2008","journal-title":"VLDB"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/564691.564709"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1363189.1363195"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/163090.163096"},{"key":"e_1_2_1_13_1","unstructured":"Goldstein J. Ramakrishnan R. Shaft U. \"Compressing relations and indexes \" In ICDE 1998   Goldstein J. Ramakrishnan R. Shaft U. \"Compressing relations and indexes \" In ICDE 1998"},{"key":"e_1_2_1_14_1","article-title":"Applications Tuning for Streaming SIMD Extensions","author":"Abel J.","year":"1999","journal-title":"Intel Technology Journal Q2"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/40.755466"},{"key":"e_1_2_1_16_1","volume-title":"Gerber R., Bik A., Smith K., Tian X., \"The Software Optimization Cookbook,\"","edition":"2"},{"key":"e_1_2_1_17_1","unstructured":"SAP AG https:\/\/www.sdn.sap.com\/irj\/sdn\/bia  SAP AG https:\/\/www.sdn.sap.com\/irj\/sdn\/bia"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/1687627.1687671","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:27:53Z","timestamp":1672226873000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/1687627.1687671"}},"subtitle":["ultra fast in-memory table scan using on-chip vector processing units"],"short-title":[],"issued":{"date-parts":[[2009,8]]},"references-count":17,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2009,8]]}},"alternative-id":["10.14778\/1687627.1687671"],"URL":"https:\/\/doi.org\/10.14778\/1687627.1687671","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2009,8]]}}}