{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T02:46:52Z","timestamp":1783738012473,"version":"3.55.0"},"reference-count":20,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2022,8]]},"abstract":"<jats:p>The ad-hoc development of new specialized computation engines targeted to very specific data workloads has created a siloed data landscape. Commonly, these engines share little to nothing with each other and are hard to maintain, evolve, and optimize, and ultimately provide an inconsistent experience to data users. In order to address these issues, Meta has created Velox, a novel open source C++ database acceleration library. Velox provides reusable, extensible, high-performance, and dialect-agnostic data processing components for building execution engines, and enhancing data management systems. The library heavily relies on vectorization and adaptivity, and is designed from the ground up to support efficient computation over complex data types due to their ubiquity in modern workloads. Velox is currently integrated or being integrated with more than a dozen data systems at Meta, including analytical query engines such as Presto and Spark, stream processing platforms, message buses and data warehouse ingestion infrastructure, machine learning systems for feature engineering and data preprocessing (PyTorch), and more. It provides benefits in terms of (a) efficiency wins by democratizing optimizations previously only found in individual engines, (b) increased consistency for data users, and (c) engineering efficiency by promoting reusability.<\/jats:p>","DOI":"10.14778\/3554821.3554829","type":"journal-article","created":{"date-parts":[[2022,9,29]],"date-time":"2022-09-29T22:28:39Z","timestamp":1664490519000},"page":"3372-3384","source":"Crossref","is-referenced-by-count":47,"title":["Velox"],"prefix":"10.14778","volume":"15","author":[{"given":"Pedro","family":"Pedreira","sequence":"first","affiliation":[{"name":"Meta Platforms Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Orri","family":"Erling","sequence":"additional","affiliation":[{"name":"Meta Platforms Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Masha","family":"Basmanova","sequence":"additional","affiliation":[{"name":"Meta Platforms Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kevin","family":"Wilfong","sequence":"additional","affiliation":[{"name":"Meta Platforms Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Laith","family":"Sakka","sequence":"additional","affiliation":[{"name":"Meta Platforms Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Krishna","family":"Pai","sequence":"additional","affiliation":[{"name":"Meta Platforms Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei","family":"He","sequence":"additional","affiliation":[{"name":"Meta Platforms Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Biswapesh","family":"Chattopadhyay","sequence":"additional","affiliation":[{"name":"Meta Platforms Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,9,29]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Apache Arrow. [n.d.]. Apache Arrow C++ Compute Functions. https:\/\/github.com\/apache\/arrow\/tree\/master\/cpp\/src\/arrow\/compute. Accessed: 2022-02-23.  Apache Arrow. [n.d.]. Apache Arrow C++ Compute Functions. https:\/\/github.com\/apache\/arrow\/tree\/master\/cpp\/src\/arrow\/compute. Accessed: 2022-02-23."},{"key":"e_1_2_1_2_1","unstructured":"Apache Arrow. [n.d.]. Arrow Columnar Format. https:\/\/arrow.apache.org\/docs\/format\/Columnar.html. Accessed: 2022-02-23.  Apache Arrow. [n.d.]. Arrow Columnar Format. https:\/\/arrow.apache.org\/docs\/format\/Columnar.html. Accessed: 2022-02-23."},{"key":"e_1_2_1_3_1","unstructured":"Apache Arrow. [n.d.]. A cross-language development platform for in-memory analytics. https:\/\/arrow.apache.org\/. Accessed: 2022-02-23.  Apache Arrow. [n.d.]. A cross-language development platform for in-memory analytics. https:\/\/arrow.apache.org\/. Accessed: 2022-02-23."},{"key":"e_1_2_1_4_1","volume-title":"Conference on Innovative Data Systems Research, CIDR.","author":"Boncz Peter","year":"2005","unstructured":"Peter Boncz , Marcin Zukowski , and Niels Nes . 2005 . MonetDB\/X100: Hyper-pipelining query execution . In Conference on Innovative Data Systems Research, CIDR. Peter Boncz, Marcin Zukowski, and Niels Nes. 2005. MonetDB\/X100: Hyper-pipelining query execution. In Conference on Innovative Data Systems Research, CIDR."},{"key":"e_1_2_1_5_1","unstructured":"Nathan Bronson and Xiao Shi. [n.d.]. Open-sourcing F14 for faster more memory-efficient hash tables. https:\/\/engineering.fb.com\/2019\/04\/25\/developer-tools\/f14\/. Accessed: 2022-02-23.  Nathan Bronson and Xiao Shi. [n.d.]. Open-sourcing F14 for faster more memory-efficient hash tables. https:\/\/engineering.fb.com\/2019\/04\/25\/developer-tools\/f14\/. Accessed: 2022-02-23."},{"key":"e_1_2_1_6_1","unstructured":"The CXL Consortium. [n.d.]. Compute Express Link: The Breakthrough CPU-to-Device Interconnect. https:\/\/www.computeexpresslink.org\/. Accessed: 2022-02-23.  The CXL Consortium. [n.d.]. Compute Express Link: The Breakthrough CPU-to-Device Interconnect. https:\/\/www.computeexpresslink.org\/. Accessed: 2022-02-23."},{"key":"e_1_2_1_7_1","unstructured":"Dremio. [n.d.]. Introducing the Gandiva Initiative for Apache Arrow. https:\/\/www.dremio.com\/announcing-gandiva-initiative-for-apache-arrow. Accessed: 2022-02-23.  Dremio. [n.d.]. Introducing the Gandiva Initiative for Apache Arrow. https:\/\/www.dremio.com\/announcing-gandiva-initiative-for-apache-arrow. Accessed: 2022-02-23."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.273032"},{"key":"e_1_2_1_9_1","volume-title":"Adaptive Execution of Compiled Queries. In 2018 IEEE 34th International Conference on Data Engineering (ICDE).","author":"Kohn Andr\u00e9","year":"2018","unstructured":"Andr\u00e9 Kohn , Viktor Leis , and Thomas Neumann . 2018 . Adaptive Execution of Compiled Queries. In 2018 IEEE 34th International Conference on Data Engineering (ICDE). Andr\u00e9 Kohn, Viktor Leis, and Thomas Neumann. 2018. Adaptive Execution of Compiled Queries. In 2018 IEEE 34th International Conference on Data Engineering (ICDE)."},{"key":"e_1_2_1_10_1","unstructured":"Wes McKinney. [n.d.]. Adding new columnar memory layouts to Arrow. https:\/\/lists.apache.org\/thread\/49qzofswg1r5z7zh39pjvd1m2ggz2kdq. Accessed: 2022-02-23.  Wes McKinney. [n.d.]. Adding new columnar memory layouts to Arrow. https:\/\/lists.apache.org\/thread\/49qzofswg1r5z7zh39pjvd1m2ggz2kdq. Accessed: 2022-02-23."},{"key":"e_1_2_1_11_1","volume-title":"Umbra: A Disk-Based System with In-Memory Performance. In 10th Conference on Innovative Data Systems Research, CIDR","author":"Neumann Thomas","year":"2020","unstructured":"Thomas Neumann and Michael J. Freitag . 2020 . Umbra: A Disk-Based System with In-Memory Performance. In 10th Conference on Innovative Data Systems Research, CIDR 2020 . www.cidrdb.org. Thomas Neumann and Michael J. Freitag. 2020. Umbra: A Disk-Based System with In-Memory Performance. In 10th Conference on Innovative Data Systems Research, CIDR 2020. www.cidrdb.org."},{"key":"e_1_2_1_12_1","unstructured":"OAP. [n.d.]. Gazelle Plugin - A Native Engine for Spark SQL with vectorized SIMD optimizations. https:\/\/oap-project.github.io\/gazelle_plugin\/latest\/. Accessed: 2022-02-23.  OAP. [n.d.]. Gazelle Plugin - A Native Engine for Spark SQL with vectorized SIMD optimizations. https:\/\/oap-project.github.io\/gazelle_plugin\/latest\/. Accessed: 2022-02-23."},{"key":"e_1_2_1_13_1","unstructured":"OAP. [n.d.]. Optimized Analytics Package. https:\/\/oap-project.github.io\/latest\/. Accessed: 2022-02-23.  OAP. [n.d.]. Optimized Analytics Package. https:\/\/oap-project.github.io\/latest\/. Accessed: 2022-02-23."},{"key":"e_1_2_1_14_1","volume-title":"Ziqi Wang, Yingjun Wu, Ran Xian, and Tieying Zhang.","author":"Pavlo Andrew","year":"2017","unstructured":"Andrew Pavlo , Gustavo Angulo , Joy Arulraj , Haibin Lin , Jiexi Lin , Lin Ma , Prashanth Menon , Todd C. Mowry , Matthew Perron , Ian Quah , Siddharth Santurkar , Anthony Tomasic , Skye Toor , Dana Van Aken , Ziqi Wang, Yingjun Wu, Ran Xian, and Tieying Zhang. 2017 . Self-Driving Database Management Systems. In CIDR. Andrew Pavlo, Gustavo Angulo, Joy Arulraj, Haibin Lin, Jiexi Lin, Lin Ma, Prashanth Menon, Todd C. Mowry, Matthew Perron, Ian Quah, Siddharth Santurkar, Anthony Tomasic, Skye Toor, Dana Van Aken, Ziqi Wang, Yingjun Wu, Ran Xian, and Tieying Zhang. 2017. Self-Driving Database Management Systems. In CIDR."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3320212"},{"key":"e_1_2_1_16_1","unstructured":"Greg Rahn Alexander Behm and Ala Luszczak. [n.d.]. Photon: The next-generation query engine for the lakehouse. https:\/\/databricks.com\/product\/photon. Accessed: 2022-02-23.  Greg Rahn Alexander Behm and Ala Luszczak. [n.d.]. Photon: The next-generation query engine for the lakehouse. https:\/\/databricks.com\/product\/photon. Accessed: 2022-02-23."},{"key":"e_1_2_1_17_1","volume-title":"Presto: SQL on Everything. In 2019 IEEE 35th International Conference on Data Engineering (ICDE). 1802--1813","author":"Sethi Raghav","year":"2019","unstructured":"Raghav Sethi , Martin Traverso , Dain Sundstrom , David Phillips , Wenlei Xie , Yutian Sun , Nezih Yegitbasi , Haozhun Jin , Eric Hwang , Nileema Shingte , and Christopher Berner . 2019 . Presto: SQL on Everything. In 2019 IEEE 35th International Conference on Data Engineering (ICDE). 1802--1813 . Raghav Sethi, Martin Traverso, Dain Sundstrom, David Phillips, Wenlei Xie, Yutian Sun, Nezih Yegitbasi, Haozhun Jin, Eric Hwang, Nileema Shingte, and Christopher Berner. 2019. Presto: SQL on Everything. In 2019 IEEE 35th International Conference on Data Engineering (ICDE). 1802--1813."},{"key":"e_1_2_1_18_1","unstructured":"Apache Spark. [n.d.]. Apache Spark - Unified Engine for large-scale data analytics. https:\/\/spark.apache.org\/. Accessed: 2022-02-23.  Apache Spark. [n.d.]. Apache Spark - Unified Engine for large-scale data analytics. https:\/\/spark.apache.org\/. Accessed: 2022-02-23."},{"key":"e_1_2_1_19_1","unstructured":"Substrait. [n.d.]. Cross-Language Serialization for Relational Algebra. https:\/\/substrait.io\/. Accessed: 2022-02-23.  Substrait. [n.d.]. Cross-Language Serialization for Relational Algebra. https:\/\/substrait.io\/. Accessed: 2022-02-23."},{"key":"e_1_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Mark Zhao Niket Agarwal Aarti Basant Bugra Gedik Satadru Pan Mustafa Ozdal Rakesh Komuravelli Jerry Pan Tianshu Bao Haowei Lu Sundaram Narayanan Jack Langman Kevin Wilfong Harsha Rastogi Carole-Jean Wu Christos Kozyrakis and Parik Pol. 2022. Understanding Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training. arXiv:2108.09373 [cs.DC]  Mark Zhao Niket Agarwal Aarti Basant Bugra Gedik Satadru Pan Mustafa Ozdal Rakesh Komuravelli Jerry Pan Tianshu Bao Haowei Lu Sundaram Narayanan Jack Langman Kevin Wilfong Harsha Rastogi Carole-Jean Wu Christos Kozyrakis and Parik Pol. 2022. Understanding Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training. arXiv:2108.09373 [cs.DC]","DOI":"10.1145\/3470496.3533044"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3554821.3554829","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:23:32Z","timestamp":1672226612000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3554821.3554829"}},"subtitle":["meta's unified execution engine"],"short-title":[],"issued":{"date-parts":[[2022,8]]},"references-count":20,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2022,8]]}},"alternative-id":["10.14778\/3554821.3554829"],"URL":"https:\/\/doi.org\/10.14778\/3554821.3554829","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2022,8]]}}}