{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T15:32:06Z","timestamp":1781105526516,"version":"3.54.1"},"reference-count":37,"publisher":"IGI Global Scientific Publishing","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,1,1]]},"abstract":"<p>Streaming data join is a critical process in the field of near-real-time data warehousing. For this purpose, an adaptive semi-stream join algorithm called CACHEJOIN (Cache Join) focusing non-uniform stream data is provided in the literature. However, this algorithm cannot exploit the memory and CPU resources optimally and consequently it leaves its service rate suboptimal due to sequential execution of both of its phases, called stream-probing (SP) phase and disk-probing (DP) phase. By integrating the advantages of CACHEJOIN, this article presents two modifications for it. The first is called P-CACHEJOIN (Parallel Cache Join) that enables the parallel processing of two phases in CACHEJOIN. This increases number of joined stream records and therefore improves throughput considerably. The second is called OP-CACHEJOIN (Optimized Parallel Cache Join) that implements a parallel loading of stored data into memory while the DP phase is executing. This research presents the performance analysis of both of the approaches defined within the paper existing CACHEJOIN empirically using synthetic skewed dataset.<\/p>","DOI":"10.4018\/jdm.2020010102","type":"journal-article","created":{"date-parts":[[2019,12,12]],"date-time":"2019-12-12T15:16:11Z","timestamp":1576163771000},"page":"20-37","source":"Crossref","is-referenced-by-count":8,"title":["Optimizing Semi-Stream CACHEJOIN for Near-Real- Time Data Warehousing"],"prefix":"10.4018","volume":"31","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6785-7875","authenticated-orcid":true,"given":"M. Asif","family":"Naeem","sequence":"first","affiliation":[{"name":"School of Engineering, Computer and Mathematical Sciences, Auckland University of Technology, Auckland, New Zealand"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9424-8274","authenticated-orcid":true,"given":"Erum","family":"Mehmood","sequence":"additional","affiliation":[{"name":"School of Science and Technology, University of Management and Technology, Lahore, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"M. G. Abbas","family":"Malik","sequence":"additional","affiliation":[{"name":"Universal College of Learning, Palmerston North, New Zealand"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Noreen","family":"Jamil","sequence":"additional","affiliation":[{"name":"National University FAST, Islamabad, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"2432","reference":[{"key":"JDM.2020010102-0","doi-asserted-by":"publisher","DOI":"10.4018\/jdm.2012100102"},{"key":"JDM.2020010102-1","doi-asserted-by":"crossref","first-page":"159","DOI":"10.1109\/ICDE.2011.5767906","article-title":"Semi-streamed index join for near-real time execution of ETL transformations.","author":"M. A.Bornea","year":"2011","journal-title":"Proceedings of the 2011 IEEE 27th International Conference on Data Engineering"},{"key":"JDM.2020010102-2","doi-asserted-by":"crossref","unstructured":"Candea, G., Polyzotis, N., & Vingralek, R. (2011). Predictable performance and high query concurrency for data analytics. The VLDB Journal\u2014The International Journal on Very Large Data Bases, 20(2), 227-248.","DOI":"10.1007\/s00778-011-0221-2"},{"key":"JDM.2020010102-3","first-page":"1","article-title":"A partition-based approach to support streaming updates over persistent data in an active datawarehouse.","author":"A.Chakraborty","year":"2009","journal-title":"Proceedings of the 2009 IEEE International Symposium on Parallel & Distributed Processing"},{"key":"JDM.2020010102-4","doi-asserted-by":"publisher","DOI":"10.4018\/jdm.2014070102"},{"key":"JDM.2020010102-5","doi-asserted-by":"publisher","DOI":"10.4018\/jdm.2009040104"},{"key":"JDM.2020010102-6","author":"A.Chris","year":"2006","journal-title":"The long tail: Why the future of business is selling less of more"},{"key":"JDM.2020010102-7","doi-asserted-by":"publisher","DOI":"10.4018\/JDM.2015010101"},{"key":"JDM.2020010102-8","doi-asserted-by":"crossref","unstructured":"Dittrich, J. P., Seeger, B., Taylor, D. S., & Widmayer, P. (2002, August). Progressive merge join: A generic and non-blocking sort-based join algorithm. Proceedings of the 28th international conference on Very Large Data Bases (pp. 299-310). VLDB Endowment.","DOI":"10.1016\/B978-155860869-6\/50034-2"},{"key":"JDM.2020010102-9","doi-asserted-by":"crossref","first-page":"1188","DOI":"10.1109\/ICNC.2013.6818158","article-title":"The algorithm of the join data stream with diskresident relation.","author":"W.Du","year":"2013","journal-title":"Proceedings of the 2013 Ninth International Conference on Natural Computation (ICNC)"},{"key":"JDM.2020010102-10","doi-asserted-by":"crossref","unstructured":"Golfarelli, M., & Rizzi, S. (2009). A survey on temporal data warehousing. International Journal of Data Warehousing, 5.","DOI":"10.4018\/jdwm.2009010101"},{"key":"JDM.2020010102-11","doi-asserted-by":"publisher","DOI":"10.4018\/JDM.2015070103"},{"key":"JDM.2020010102-12","doi-asserted-by":"publisher","DOI":"10.4018\/jdm.2005040103"},{"key":"JDM.2020010102-13","doi-asserted-by":"publisher","DOI":"10.4018\/JDM.2019070103"},{"key":"JDM.2020010102-14","author":"R.Kimball","year":"2011","journal-title":"The Data Warehouse? ETL Toolkit: Practical Techniques for Extracting, Cleaning, Conforming, and Delivering Data"},{"issue":"4","key":"JDM.2020010102-15","doi-asserted-by":"crossref","first-page":"373","DOI":"10.1023\/A:1024940629314","article-title":"Bursty and hierarchical structure in streams.","volume":"7","author":"J.Kleinberg","year":"2003","journal-title":"Data Mining and Knowledge Discovery"},{"key":"JDM.2020010102-16","volume":"Vol. 3","author":"D. E.Knuth","year":"1998","journal-title":"The art of computer programming: sorting and searching"},{"key":"JDM.2020010102-17","doi-asserted-by":"publisher","DOI":"10.4018\/JDM.2015040102"},{"key":"JDM.2020010102-18","doi-asserted-by":"publisher","DOI":"10.1109\/ICCCBDA.2017.7951887"},{"key":"JDM.2020010102-19","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1109\/ICDE.2004.1320002","article-title":"Hash-merge join: A non-blocking join algorithm for producing fast and early join results.","author":"M. F.Mokbel","year":"2004","journal-title":"Proceedings of the 20th International Conference on Data Engineering"},{"key":"JDM.2020010102-20","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-24577-0_5"},{"key":"JDM.2020010102-21","doi-asserted-by":"publisher","DOI":"10.4018\/jdwm.2011100102"},{"key":"JDM.2020010102-22","first-page":"21","article-title":"Optimised X-HYBRIDJOIN for near-real-time data warehousing.","author":"M. A.Naeem","year":"2012","journal-title":"Proceedings of the Twenty-Third Australasian Database Conference"},{"key":"JDM.2020010102-23","doi-asserted-by":"crossref","first-page":"431","DOI":"10.1007\/978-3-642-32584-7_35","article-title":"A lightweight stream-based join with limited resource consumption.","author":"M. A.Naeem","year":"2012","journal-title":"Proceedings of the International Conference on Data Warehousing and Knowledge Discovery"},{"key":"JDM.2020010102-24","doi-asserted-by":"publisher","DOI":"10.1145\/1871940.1871952"},{"key":"JDM.2020010102-25","doi-asserted-by":"crossref","first-page":"236","DOI":"10.1007\/978-3-642-40131-2_20","article-title":"SSCJ: A semi-stream cache join using a front-stage cache module.","author":"M. A.Naeem","year":"2013","journal-title":"Proceedings of the International Conference on Data Warehousing and Knowledge Discovery"},{"key":"JDM.2020010102-26","doi-asserted-by":"publisher","DOI":"10.4018\/jdm.2007010104"},{"key":"JDM.2020010102-27","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2008.27"},{"key":"JDM.2020010102-28","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2007.367893"},{"key":"JDM.2020010102-29","author":"R.Ramakrishnan","year":"2000","journal-title":"Database management systems"},{"key":"JDM.2020010102-30","doi-asserted-by":"crossref","first-page":"74","DOI":"10.1007\/11546849_8","article-title":"A survey of open source tools for business intelligence.","author":"C.Thomsen","year":"2005","journal-title":"Proceedings of the International Conference on Data Warehousing and Knowledge Discovery"},{"key":"JDM.2020010102-31","doi-asserted-by":"publisher","DOI":"10.4018\/jdm.2004070105"},{"key":"JDM.2020010102-32","doi-asserted-by":"publisher","DOI":"10.4018\/jdm.2004010102"},{"key":"JDM.2020010102-33","doi-asserted-by":"publisher","DOI":"10.1109\/TLA.2018.8789552"},{"key":"JDM.2020010102-34","doi-asserted-by":"publisher","DOI":"10.4018\/jdwm.2009070101"},{"key":"JDM.2020010102-35","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2017.2756932"},{"key":"JDM.2020010102-36","doi-asserted-by":"publisher","DOI":"10.4018\/jdm.2007070104"}],"container-title":["Journal of Database Management"],"original-title":[],"language":"ng","link":[{"URL":"https:\/\/www.igi-global.com\/viewtitle.aspx?TitleId=245298","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,5,6]],"date-time":"2022-05-06T16:25:05Z","timestamp":1651854305000},"score":1,"resource":{"primary":{"URL":"https:\/\/services.igi-global.com\/resolvedoi\/resolve.aspx?doi=10.4018\/JDM.2020010102"}},"subtitle":[""],"short-title":[],"issued":{"date-parts":[[2020,1,1]]},"references-count":37,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,1]]}},"URL":"https:\/\/doi.org\/10.4018\/jdm.2020010102","relation":{},"ISSN":["1063-8016","1533-8010"],"issn-type":[{"value":"1063-8016","type":"print"},{"value":"1533-8010","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,1,1]]}}}