{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,30]],"date-time":"2025-10-30T07:14:37Z","timestamp":1761808477698},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,7]]},"abstract":"<jats:p>Microsoft recently introduced Azure Synapse Analytics, which offers an integrated experience across data ingestion, storage, and querying in Apache Spark and T-SQL over data in the lake, including files and warehouse tables. In this paper, we present our experiences with designing and implementing Hyperspace, the indexing subsystem underlying Synapse. Hyperspace enables users to build multiple types of secondary indexes on their data, maintain them through a multi-user concurrency model, and leverage them automatically---without any change to their application code---for query\/workload acceleration. Many requirements of Hyperspace are based on feedback from several enterprise customers. We present the details of Hyperspace's underlying design, the user-facing APIs, its concurrency control protocol for index access, its index-aware query processing techniques, and its maintenance mechanisms for handling index updates. Evaluations over standard industry benchmarks and real customer workloads show that Hyperspace can accelerate query execution by up to 10x and in certain real-world workloads, even up to two orders of magnitude.<\/jats:p>","DOI":"10.14778\/3476311.3476382","type":"journal-article","created":{"date-parts":[[2021,10,28]],"date-time":"2021-10-28T22:48:43Z","timestamp":1635461323000},"page":"3043-3055","source":"Crossref","is-referenced-by-count":9,"title":["Hyperspace"],"prefix":"10.14778","volume":"14","author":[{"given":"Rahul","family":"Potharaju","sequence":"first","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Terry","family":"Kim","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eunjin","family":"Song","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wentao","family":"Wu","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lev","family":"Novik","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Apoorve","family":"Dave","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrew","family":"Fogarty","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pouria","family":"Pirzadeh","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vidip","family":"Acharya","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gurleen","family":"Dhody","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiying","family":"Li","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sinduja","family":"Ramanujam","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nicolas","family":"Bruno","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"C\u00e9sar A.","family":"Galindo-Legaria","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vivek","family":"Narasayya","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Surajit","family":"Chaudhuri","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anil K.","family":"Nori","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tomas","family":"Talius","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Raghu","family":"Ramakrishnan","sequence":"additional","affiliation":[{"name":"Microsoft Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,28]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"[n.d.]. Apache Parquet. https:\/\/parquet.apache.org\/.  [n.d.]. Apache Parquet. https:\/\/parquet.apache.org\/."},{"key":"e_1_2_1_2_1","unstructured":"[n.d.]. Azure Synapse Analytics. https:\/\/azure.microsoft.com\/en-us\/services\/synapse-analytics\/.  [n.d.]. Azure Synapse Analytics. https:\/\/azure.microsoft.com\/en-us\/services\/synapse-analytics\/."},{"key":"e_1_2_1_3_1","unstructured":"[n.d.]. Bucketing in Spark. https:\/\/jaceklaskowski.gitbooks.io\/mastering-spark-sql\/spark-sql-bucketing.html.  [n.d.]. Bucketing in Spark. https:\/\/jaceklaskowski.gitbooks.io\/mastering-spark-sql\/spark-sql-bucketing.html."},{"key":"e_1_2_1_4_1","unstructured":"[n.d.]. Deep Dive into GPU Support in Apache Spark 3.x. https:\/\/databricks.com\/session_na20\/deep-dive-into-gpu-support-in-apache-spark-3-x.  [n.d.]. Deep Dive into GPU Support in Apache Spark 3.x. https:\/\/databricks.com\/session_na20\/deep-dive-into-gpu-support-in-apache-spark-3-x."},{"key":"e_1_2_1_5_1","unstructured":"[n.d.]. FPGA-Based Acceleration Architecture for Spark SQL. https:\/\/databricks.com\/session\/fpga-based-acceleration-architecture-for-spark-sql.  [n.d.]. FPGA-Based Acceleration Architecture for Spark SQL. https:\/\/databricks.com\/session\/fpga-based-acceleration-architecture-for-spark-sql."},{"key":"e_1_2_1_6_1","unstructured":"[n.d.]. Microsoft SQL Server Columnstore Indexes. https:\/\/docs.microsoft.com\/en-us\/sql\/relational-databases\/indexes\/columnstore-indexes-overview?view=sql-server-ver15.  [n.d.]. Microsoft SQL Server Columnstore Indexes. https:\/\/docs.microsoft.com\/en-us\/sql\/relational-databases\/indexes\/columnstore-indexes-overview?view=sql-server-ver15."},{"key":"e_1_2_1_7_1","unstructured":"[n.d.]. Parquet MR. https:\/\/github.com\/apache\/parquet-mr.  [n.d.]. Parquet MR. https:\/\/github.com\/apache\/parquet-mr."},{"key":"e_1_2_1_8_1","unstructured":"[n.d.]. Scala Pattern Matching. https:\/\/docs.scala-lang.org\/tour\/pattern-matching.html.  [n.d.]. Scala Pattern Matching. https:\/\/docs.scala-lang.org\/tour\/pattern-matching.html."},{"key":"e_1_2_1_9_1","unstructured":"[n.d.]. The TPC-DS Benchmark. http:\/\/www.tpc.org\/tpcds\/.  [n.d.]. The TPC-DS Benchmark. http:\/\/www.tpc.org\/tpcds\/."},{"key":"e_1_2_1_10_1","unstructured":"[n.d.]. The TPC-H Benchmark. http:\/\/www.tpc.org\/tpch\/.  [n.d.]. The TPC-H Benchmark. http:\/\/www.tpc.org\/tpch\/."},{"key":"e_1_2_1_11_1","unstructured":"[n.d.]. Windows Azure Storage Blob (WASB). https:\/\/gerardnico.com\/azure\/wasb.  [n.d.]. Windows Azure Storage Blob (WASB). https:\/\/gerardnico.com\/azure\/wasb."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/3236187.3236195"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415545"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415560"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2742797"},{"key":"e_1_2_1_16_1","volume-title":"Merge strategies: from merge sort to Timsort. URL https:\/\/hal-upec-upem.archives-ouvertes.fr\/hal-01212839, working paper or preprint","author":"Auger Nicolas","year":"2015","unstructured":"Nicolas Auger , Cyril Nicaud , and Carine Pivoteau . 2015. Merge strategies: from merge sort to Timsort. URL https:\/\/hal-upec-upem.archives-ouvertes.fr\/hal-01212839, working paper or preprint ( 2015 ). Nicolas Auger, Cyril Nicaud, and Carine Pivoteau. 2015. Merge strategies: from merge sort to Timsort. URL https:\/\/hal-upec-upem.archives-ouvertes.fr\/hal-01212839, working paper or preprint (2015)."},{"key":"e_1_2_1_17_1","unstructured":"Azure. 2021. Synapse Link. https:\/\/docs.microsoft.com\/en-us\/azure\/cosmos-db\/synapse-link.  Azure. 2021. Synapse Link. https:\/\/docs.microsoft.com\/en-us\/azure\/cosmos-db\/synapse-link."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2018.2850339"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2043556.2043571"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/645923.673646"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/276305.276337"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/356770.356776"},{"key":"e_1_2_1_23_1","unstructured":"DataBricks. 2015. Deep Dive into Spark SQL's Catalyst Optimizer. https:\/\/goo.gl\/GZtSWC.  DataBricks. 2015. Deep Dive into Spark SQL's Catalyst Optimizer. https:\/\/goo.gl\/GZtSWC."},{"key":"e_1_2_1_24_1","unstructured":"Dremio. 2020. Reflections. https:\/\/docs.dremio.com\/acceleration\/reflections.html.  Dremio. 2020. Reflections. https:\/\/docs.dremio.com\/acceleration\/reflections.html."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2814710.2814713"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/320083.320092"},{"key":"e_1_2_1_27_1","unstructured":"Apache Foundation. 2020. Apache Hudi. https:\/\/github.com\/apache\/hudi.  Apache Foundation. 2020. Apache Hudi. https:\/\/github.com\/apache\/hudi."},{"key":"e_1_2_1_28_1","unstructured":"Apache Foundation. 2020. Apache Iceberg. https:\/\/iceberg.apache.org\/spark\/.  Apache Foundation. 2020. Apache Iceberg. https:\/\/iceberg.apache.org\/spark\/."},{"key":"e_1_2_1_29_1","unstructured":"Apache Foundation. 2021. Apache Spark. https:\/\/github.com\/apache\/spark.  Apache Foundation. 2021. Apache Spark. https:\/\/github.com\/apache\/spark."},{"key":"e_1_2_1_30_1","unstructured":"Linux Foundation. 2020. Delta Lake. https:\/\/github.com\/delta-io\/delta.  Linux Foundation. 2020. Delta Lake. https:\/\/github.com\/delta-io\/delta."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/971697.602266"},{"key":"e_1_2_1_32_1","volume-title":"Hadoop Performance Models. CoRR abs\/1106.0940","author":"Herodotou Herodotos","year":"2011","unstructured":"Herodotos Herodotou . 2011. Hadoop Performance Models. CoRR abs\/1106.0940 ( 2011 ). Herodotos Herodotou. 2011. Hadoop Performance Models. CoRR abs\/1106.0940 (2011)."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/3402707.3402746"},{"key":"e_1_2_1_34_1","first-page":"5","article-title":"A What-if Engine for Cost-based MapReduce Optimization","volume":"36","author":"Herodotou Herodotos","year":"2013","unstructured":"Herodotos Herodotou and Shivnath Babu . 2013 . A What-if Engine for Cost-based MapReduce Optimization . IEEE Data Eng. Bull. 36 , 1 (2013), 5 -- 14 . Herodotos Herodotou and Shivnath Babu. 2013. A What-if Engine for Cost-based MapReduce Optimization. IEEE Data Eng. Bull. 36, 1 (2013), 5--14.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.14778\/3192965.3192971"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3190656"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407832"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2903744"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/1286887.1286911"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.14778\/3055540.3055551"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.14778\/3231751.3231765"},{"key":"e_1_2_1_42_1","volume-title":"https:\/\/goo.gl\/2KkwMv","author":"Compliance GDPR","year":"2017","unstructured":"Microsoft. 2017. GDPR Compliance . https:\/\/goo.gl\/2KkwMv ( 2017 ). Microsoft. 2017. GDPR Compliance. https:\/\/goo.gl\/2KkwMv (2017)."},{"key":"e_1_2_1_43_1","unstructured":"Microsoft. 2020. Hyperspace. https:\/\/github.com\/microsoft\/hyperspace.  Microsoft. 2020. Hyperspace. https:\/\/github.com\/microsoft\/hyperspace."},{"key":"e_1_2_1_44_1","unstructured":"Microsoft. 2020. Hyperspace in Azure Synapse Analytics. https:\/\/aka.ms\/synapse\/hyperspace.  Microsoft. 2020. Hyperspace in Azure Synapse Analytics. https:\/\/aka.ms\/synapse\/hyperspace."},{"key":"e_1_2_1_45_1","unstructured":"Microsoft. 2020. Time Series Insights. https:\/\/azure.microsoft.com\/en-us\/services\/time-series-insights\/.  Microsoft. 2020. Time Series Insights. https:\/\/azure.microsoft.com\/en-us\/services\/time-series-insights\/."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.5555\/645924.671173"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/253262.253268"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.14778\/3368289.3368292"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/s002360050048"},{"key":"e_1_2_1_51_1","unstructured":"Rahul Potharaju. 2021. NVIDIA GPU Acceleration for Apache Spark\u2122 in Azure Synapse Analytics. http:\/\/aka.ms\/synapse-spark-gpu (2021).  Rahul Potharaju. 2021. NVIDIA GPU Acceleration for Apache Spark\u2122 in Azure Synapse Analytics. http:\/\/aka.ms\/synapse-spark-gpu (2021)."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415547"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370036.2145836"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.14778\/3339490.3339495"},{"key":"e_1_2_1_55_1","volume-title":"Hyperspace for Delta Lake. https:\/\/aka.ms\/sais2021-hyperspace-for-delta-lake","author":"Rahul Potharaju Eunjin Song","year":"2021","unstructured":"Eunjin Song Rahul Potharaju , Terry Kim . 2021. Hyperspace for Delta Lake. https:\/\/aka.ms\/sais2021-hyperspace-for-delta-lake ( 2021 ). Eunjin Song Rahul Potharaju, Terry Kim. 2021. Hyperspace for Delta Lake. https:\/\/aka.ms\/sais2021-hyperspace-for-delta-lake (2021)."},{"key":"e_1_2_1_56_1","volume-title":"Hyperspace: An Indexing Sub-system for Apache Spark. https:\/\/aka.ms\/sais2020-hyperspace","author":"Rahul Potharaju Terry Kim","year":"2020","unstructured":"Terry Kim Rahul Potharaju . 2020 . Hyperspace: An Indexing Sub-system for Apache Spark. https:\/\/aka.ms\/sais2020-hyperspace (2020). Terry Kim Rahul Potharaju. 2020. Hyperspace: An Indexing Sub-system for Apache Spark. https:\/\/aka.ms\/sais2020-hyperspace (2020)."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3056100"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.5555\/647221.718601"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSST.2010.5496972"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.5555\/1083592.1083658"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687553.1687609"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3320227"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.5555\/846219.847390"},{"key":"e_1_2_1_64_1","volume-title":"Move Fast and Meet Deadlines: Fine-grained Real-time Stream Processing with Cameo. In NSDI.","author":"Xu Le","year":"2021","unstructured":"Le Xu , Shivaram Venkataraman , Indranil Gupta , Luo Mai , and Rahul Potharaju . 2021 . Move Fast and Meet Deadlines: Fine-grained Real-time Stream Processing with Cameo. In NSDI. Le Xu, Shivaram Venkataraman, Indranil Gupta, Luo Mai, and Rahul Potharaju. 2021. Move Fast and Meet Deadlines: Fine-grained Real-time Stream Processing with Cameo. In NSDI."},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10707-018-0330-9"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.5555\/2228298.2228301"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3476311.3476382","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:37:40Z","timestamp":1672227460000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3476311.3476382"}},"subtitle":["the indexing subsystem of azure synapse"],"short-title":[],"issued":{"date-parts":[[2021,7]]},"references-count":65,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2021,7]]}},"alternative-id":["10.14778\/3476311.3476382"],"URL":"https:\/\/doi.org\/10.14778\/3476311.3476382","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,7]]}}}