{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T03:47:50Z","timestamp":1783568870989,"version":"3.55.0"},"reference-count":18,"publisher":"Association for Computing Machinery (ACM)","issue":"13","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2016,9]]},"abstract":"<jats:p>\n            We present LocationSpark, a spatial data processing system built on top of Apache Spark, a widely used distributed data processing system. LocationSpark offers a rich set of spatial query operators, e.g., range search,\n            <jats:italic>k<\/jats:italic>\n            NN, spatio-textual operation, spatial-join, and\n            <jats:italic>k<\/jats:italic>\n            NN-join. To achieve high performance, LocationSpark employs various spatial indexes for in-memory data, and guarantees that immutable spatial indexes have low overhead with fault tolerance. In addition, we build two new layers over Spark, namely a query scheduler and a query executor. The query scheduler is responsible for mitigating skew in spatial queries, while the query executor selects the best plan based on the indexes and the nature of the spatial queries. Furthermore, to avoid unnecessary network communication overhead when processing overlapped spatial data, We embed an efficient spatial Bloom filter into LocationSpark's indexes. Finally, LocationSpark tracks frequently accessed spatial data, and dynamically flushes less frequently accessed data into disk. We evaluate our system on real workloads and demonstrate that it achieves an order of magnitude performance gain over a baseline framework.\n          <\/jats:p>","DOI":"10.14778\/3007263.3007310","type":"journal-article","created":{"date-parts":[[2016,11,1]],"date-time":"2016-11-01T13:47:47Z","timestamp":1478008067000},"page":"1565-1568","source":"Crossref","is-referenced-by-count":124,"title":["LocationSpark"],"prefix":"10.14778","volume":"9","author":[{"given":"Mingjie","family":"Tang","sequence":"first","affiliation":[{"name":"Purdue University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yongyang","family":"Yu","sequence":"additional","affiliation":[{"name":"Purdue University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qutaibah M.","family":"Malluhi","sequence":"additional","affiliation":[{"name":"Qatar University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mourad","family":"Ouzzani","sequence":"additional","affiliation":[{"name":"Qatar Computing Research Institute, HBKU"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Walid G.","family":"Aref","sequence":"additional","affiliation":[{"name":"Purdue University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2016,9]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Geotrellis. https:\/\/github.com\/geotrellis\/geotrellis.  Geotrellis. https:\/\/github.com\/geotrellis\/geotrellis."},{"key":"e_1_2_1_2_1","unstructured":"Magellan. https:\/\/github.com\/harsha2010\/magellan.  Magellan. https:\/\/github.com\/harsha2010\/magellan."},{"key":"e_1_2_1_3_1","unstructured":"Spatialspark. http:\/\/simin.me\/projects\/spatialspark\/.  Spatialspark. http:\/\/simin.me\/projects\/spatialspark\/."},{"key":"e_1_2_1_4_1","volume-title":"Optimizing joins in a map-reduce environment. Technical report","author":"Afrati F. N.","year":"2009","unstructured":"F. N. Afrati and J. D. Ullman . Optimizing joins in a map-reduce environment. Technical report , National Technical University of Athens , Stanford University , December 2009 . F. N. Afrati and J. D. Ullman. Optimizing joins in a map-reduce environment. Technical report, National Technical University of Athens, Stanford University, December 2009."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536222.2536227"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.14778\/2831360.2831361"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2742797"},{"key":"e_1_2_1_8_1","volume-title":"OSDI'04","author":"Dean J.","year":"2004","unstructured":"J. Dean and S. Ghemawat . Mapreduce: simplified data processing on large clusters . In OSDI'04 . USENIX Association , 2004 . J. Dean and S. Ghemawat. Mapreduce: simplified data processing on large clusters. In OSDI'04. USENIX Association, 2004."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2015.7113382"},{"key":"e_1_2_1_10_1","first-page":"599","volume-title":"OSDI'14","author":"Gonzalez J. E.","year":"2014","unstructured":"J. E. Gonzalez , R. S. Xin , A. Dave , D. Crankshaw , M. J. Franklin , and I. Stoica . Graphx: Graph processing in a distributed dataflow framework . In OSDI'14 , pages 599 -- 613 , Broomfield, CO , Oct. 2014 . USENIX Association. J. E. Gonzalez, R. S. Xin, A. Dave, D. Crankshaw, M. J. Franklin, and I. Stoica. Graphx: Graph processing in a distributed dataflow framework. In OSDI'14, pages 599--613, Broomfield, CO, Oct. 2014. USENIX Association."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2820783.2820860"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213840"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.14778\/2336664.2336674"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/MDM.2011.41"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2756547"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/2556549.2556570"},{"key":"e_1_2_1_17_1","first-page":"15","volume-title":"NSDI'12","author":"Zaharia M.","year":"2012","unstructured":"M. Zaharia , M. Chowdhury , T. Das , A. Dave , J. Ma , M. McCauly , M. J. Franklin , S. Shenker , and I. Stoica . Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing . In NSDI'12 , pages 15 -- 28 , San Jose, CA , 2012 . USENIX. M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauly, M. J. Franklin, S. Shenker, and I. Stoica. Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. In NSDI'12, pages 15--28, San Jose, CA, 2012. USENIX."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2522737"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3007263.3007310","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T09:40:52Z","timestamp":1672220452000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3007263.3007310"}},"subtitle":["a distributed in-memory data management system for big spatial data"],"short-title":[],"issued":{"date-parts":[[2016,9]]},"references-count":18,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2016,9]]}},"alternative-id":["10.14778\/3007263.3007310"],"URL":"https:\/\/doi.org\/10.14778\/3007263.3007310","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2016,9]]}}}