{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,12,29]],"date-time":"2022-12-29T05:21:22Z","timestamp":1672291282122},"reference-count":12,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2019,8]]},"abstract":"<jats:p>\n            The last decade has witnessed a huge increase in data being ingested into the cloud from a variety of data sources. The ingested data takes various forms such as JSON, CSV, and binary formats. Traditionally, data is either ingested into storage in raw form, indexed ad-hoc using range indices, or cooked into analytics-friendly columnar formats. None of these solutions is able to handle modern requirements on storage: making the data available immediately for ad-hoc and streaming queries while ingesting at extremely high throughputs. We demonstrate FishStore, our open-source concurrent latch-free storage layer for data with flexible schema. FishStore builds on recent advances in parsing and indexing techniques, and is based on multi-chain hash indexing of dynamically registered\n            <jats:italic>predicated subsets<\/jats:italic>\n            of data. We find predicated subset hashing to be a powerful primitive that supports a broad range of queries on ingested data and admits a higher performance (by up to an order of magnitude) implementation than current alternatives.\n          <\/jats:p>","DOI":"10.14778\/3352063.3352100","type":"journal-article","created":{"date-parts":[[2019,9,18]],"date-time":"2019-09-18T18:36:11Z","timestamp":1568831771000},"page":"1922-1925","source":"Crossref","is-referenced-by-count":1,"title":["FishStore"],"prefix":"10.14778","volume":"12","author":[{"given":"Badrish","family":"Chandramouli","sequence":"first","affiliation":[{"name":"Microsoft Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dong","family":"Xie","sequence":"additional","affiliation":[{"name":"University of Utah and Microsoft Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yinan","family":"Li","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Donald","family":"Kossmann","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,8]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Apache Parquet. https:\/\/parquet.apache.org\/.  Apache Parquet. https:\/\/parquet.apache.org\/."},{"key":"e_1_2_1_2_1","unstructured":"Apache Thrift. https:\/\/thrift.apache.org\/.  Apache Thrift. https:\/\/thrift.apache.org\/."},{"key":"e_1_2_1_3_1","unstructured":"FishStore. https:\/\/github.com\/microsoft\/FishStore.  FishStore. https:\/\/github.com\/microsoft\/FishStore."},{"key":"e_1_2_1_4_1","unstructured":"GH Archive. https:\/\/www.gharchive.org\/.  GH Archive. https:\/\/www.gharchive.org\/."},{"key":"e_1_2_1_5_1","unstructured":"Microsoft PowerBI. https:\/\/powerbi.microsoft.com\/.  Microsoft PowerBI. https:\/\/powerbi.microsoft.com\/."},{"key":"e_1_2_1_6_1","unstructured":"RocksDB. http:\/\/rocksdb.org\/.  RocksDB. http:\/\/rocksdb.org\/."},{"key":"e_1_2_1_7_1","unstructured":"simdjson. https:\/\/github.com\/lemire\/simdjson.  simdjson. https:\/\/github.com\/lemire\/simdjson."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213864"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196898"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.14778\/3402755.3402799"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.14778\/3115404.3115416"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319896"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3352063.3352100","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:42:55Z","timestamp":1672224175000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3352063.3352100"}},"subtitle":["fast ingestion and indexing of raw data"],"short-title":[],"issued":{"date-parts":[[2019,8]]},"references-count":12,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2019,8]]}},"alternative-id":["10.14778\/3352063.3352100"],"URL":"https:\/\/doi.org\/10.14778\/3352063.3352100","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2019,8]]}}}