{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,1,7]],"date-time":"2023-01-07T15:14:54Z","timestamp":1673104494219},"reference-count":9,"publisher":"Association for Computing Machinery (ACM)","issue":"13","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2016,9]]},"abstract":"<jats:p>Large quantities of raw data are being generated by many different sources in different formats. Private and public sectors alike acclaim the valuable information and insights that can be mined from such data to better understand the dynamics of everyday life, such as traffic, worldwide logistics, and social behavior. For this reason, storing, managing, and analyzing \"Big Data\" at scale is getting a tremendous amount of attention, both in academia and industry. In this paper, we demonstrate the power of a parallel connection that we have built between Apache Spark and Apache AsterixDB (Incubating) to enable complex analytics such as machine learning and graph analysis on data drawn from large semi-structured data collections. The integration of these two systems allows researchers and data scientists to leverage AsterixDB capabilities, including fast ingestion and indexing of semi-structured data and efficient answering of geo-spatial and fuzzy text queries. Complex data analytics can then be performed on the resulting AsterixDB query output in order to obtain additional insights by leveraging the power of Spark's machine learning and graph libraries.<\/jats:p>","DOI":"10.14778\/3007263.3007315","type":"journal-article","created":{"date-parts":[[2016,11,1]],"date-time":"2016-11-01T13:47:47Z","timestamp":1478008067000},"page":"1585-1588","source":"Crossref","is-referenced-by-count":3,"title":["Large-scale complex analytics on semi-structured datasets using asterixDB and spark"],"prefix":"10.14778","volume":"9","author":[{"given":"Wail Y.","family":"Alkowaileet","sequence":"first","affiliation":[{"name":"Center for Complex Engineering Systems at KACST and MIT"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sattam","family":"Alsubaiee","sequence":"additional","affiliation":[{"name":"Center for Complex Engineering Systems at KACST and MIT"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael J.","family":"Carey","sequence":"additional","affiliation":[{"name":"University of California and Couchbase"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Till","family":"Westmann","sequence":"additional","affiliation":[{"name":"Couchbase"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yingyi","family":"Bu","sequence":"additional","affiliation":[{"name":"Couchbase"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,9]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.14778\/2733085.2733096"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732951.2732958"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767921"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920881"},{"key":"e_1_2_1_5_1","volume-title":"The research space: using the career paths of scholars to predict the evolution of the research output of individuals, institutions, and nations","author":"Guevara M. R.","year":"2016","unstructured":"M. R. Guevara The research space: using the career paths of scholars to predict the evolution of the research output of individuals, institutions, and nations . 2016 . arXiv:1602.08409 {cs.DL}. M. R. Guevara et al. The research space: using the career paths of scholars to predict the evolution of the research output of individuals, institutions, and nations. 2016. arXiv:1602.08409 {cs.DL}."},{"key":"e_1_2_1_6_1","volume-title":"Proc. HotCloud","author":"Zaharia M.","year":"2010","unstructured":"M. Zaharia : Cluster computing with working sets . In Proc. HotCloud , 2010 . M. Zaharia et al. Spark: Cluster computing with working sets. In Proc. HotCloud, 2010."},{"key":"e_1_2_1_7_1","volume-title":"NSDI","author":"Zaharia M.","year":"2012","unstructured":"M. Zaharia Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing . NSDI , 2012 . M. Zaharia et al. Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. NSDI, 2012."},{"key":"e_1_2_1_8_1","unstructured":"MongoDB-Spark Connector: White Paper. https:\/\/www.mongodb.com\/collateral\/apache-spark-and-mongodb-turning-analytics-into-real-time-action.  MongoDB-Spark Connector: White Paper. https:\/\/www.mongodb.com\/collateral\/apache-spark-and-mongodb-turning-analytics-into-real-time-action."},{"key":"e_1_2_1_9_1","unstructured":"Apache Zeppelin. https:\/\/zeppelin.incubator.apache.org\/.  Apache Zeppelin. https:\/\/zeppelin.incubator.apache.org\/."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3007263.3007315","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T09:41:06Z","timestamp":1672220466000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3007263.3007315"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,9]]},"references-count":9,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2016,9]]}},"alternative-id":["10.14778\/3007263.3007315"],"URL":"https:\/\/doi.org\/10.14778\/3007263.3007315","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2016,9]]}}}