{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T04:40:13Z","timestamp":1785300013227,"version":"3.55.0"},"reference-count":44,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2022,1,14]],"date-time":"2022-01-14T00:00:00Z","timestamp":1642118400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Data"],"abstract":"<jats:p>Frequent itemset mining (FIM) is a common approach for discovering hidden frequent patterns from transactional databases used in prediction, association rules, classification, etc. Apriori is an FIM elementary algorithm with iterative nature used to find the frequent itemsets. Apriori is used to scan the dataset multiple times to generate big frequent itemsets with different cardinalities. Apriori performance descends when data gets bigger due to the multiple dataset scan to extract the frequent itemsets. Eclat is a scalable version of the Apriori algorithm that utilizes a vertical layout. The vertical layout has many advantages; it helps to solve the problem of multiple datasets scanning and has information that helps to find each itemset support. In a vertical layout, itemset support can be achieved by intersecting transaction ids (tidset\/tids) and pruning irrelevant itemsets. However, when tids become too big for memory, it affects algorithms efficiency. In this paper, we introduce SHFIM (spark-based hybrid frequent itemset mining), which is a three-phase algorithm that utilizes both horizontal and vertical layout diffset instead of tidset to keep track of the differences between transaction ids rather than the intersections. Moreover, some improvements are developed to decrease the number of candidate itemsets. SHFIM is implemented and tested over the Spark framework, which utilizes the RDD (resilient distributed datasets) concept and in-memory processing that tackles MapReduce framework problem. We compared the SHFIM performance with Spark-based Eclat and dEclat algorithms for the four benchmark datasets. Experimental results proved that SHFIM outperforms Eclat and dEclat Spark-based algorithms in both dense and sparse datasets in terms of execution time.<\/jats:p>","DOI":"10.3390\/data7010011","type":"journal-article","created":{"date-parts":[[2022,1,14]],"date-time":"2022-01-14T12:34:04Z","timestamp":1642163644000},"page":"11","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":23,"title":["An Efficient Spark-Based Hybrid Frequent Itemset Mining Algorithm for Big Data"],"prefix":"10.3390","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3837-7781","authenticated-orcid":false,"given":"Mohamed Reda","family":"Al-Bana","sequence":"first","affiliation":[{"name":"Department of Information Systems, Faculty of Computers and Artificial Intelligence, Helwan University, Cairo 11795, Egypt"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6340-2547","authenticated-orcid":false,"given":"Marwa Salah","family":"Farhan","sequence":"additional","affiliation":[{"name":"Department of Information Systems, Faculty of Computers and Artificial Intelligence, Helwan University, Cairo 11795, Egypt"},{"name":"Faculty of Informatics and Computer Science, British University in Egypt, Cairo 11837, Egypt"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5419-3739","authenticated-orcid":false,"given":"Nermin Abdelhakim","family":"Othman","sequence":"additional","affiliation":[{"name":"Department of Information Systems, Faculty of Computers and Artificial Intelligence, Helwan University, Cairo 11795, Egypt"},{"name":"Faculty of Informatics and Computer Science, British University in Egypt, Cairo 11837, Egypt"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,1,14]]},"reference":[{"key":"ref_1","unstructured":"Jiawei, H., and Kamber, M. (2021, December 13). Data Mining Concepts and Techniques, 550. Available online: https:\/\/www.researchgate.net\/publication\/235902451_Data_Mining_Concept_and_Techniques."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1016\/j.bdr.2017.06.006","article-title":"Frequent Itemsets Mining for Big Data: A Comparative Analysis","volume":"9","author":"Apiletti","year":"2017","journal-title":"Big Data Res."},{"key":"ref_3","unstructured":"(2022, January 04). Big Data Tutorial|All You Need to Know about Big Data|Edureka. Available online: https:\/\/www.edureka.co\/blog\/big-data-tutorial."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"24","DOI":"10.1186\/s40537-015-0032-1","article-title":"A survey of open source tools for machine learning with big data in the Hadoop ecosystem","volume":"2","author":"Landset","year":"2015","journal-title":"J. Big Data"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2610","DOI":"10.1007\/s10489-020-01677-5","article-title":"K-PbC: An Improved Cluster Center Initialization for Categorical Data Clustering","volume":"50","author":"Tai","year":"2020","journal-title":"Applied Intelligence"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"216","DOI":"10.1093\/bib\/bbt074","article-title":"A primer to frequent itemset mining for bioinformatics","volume":"16","author":"Naulaerts","year":"2015","journal-title":"Brief. Bioinform."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"586","DOI":"10.1016\/j.procs.2015.10.040","article-title":"Efficient Data Mining Method to Predict the Risk of Heart Diseases through Frequent Itemsets","volume":"70","author":"Ilayaraja","year":"2015","journal-title":"Procedia Comput. Sci."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Loshin, D. (2013). Knowledge Discovery and Data Mining for Predictive Analytics. Bus. Intell., 271\u2013286.","DOI":"10.1016\/B978-0-12-385889-4.00017-X"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"e1329","DOI":"10.1002\/widm.1329","article-title":"Frequent itemset mining: A 25 years review","volume":"9","author":"Luna","year":"2019","journal-title":"Wiley Interdiscip. Rev. Data Min. Knowl. Discov."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Apiletti, D., Baralis, E., Cerquitelli, T., Chiusano, S., and Grimaudo, L. (2013, January 16\u201318). SeaRum: A Cloud-Based Service for Association Rule Mining. Proceedings of the 2013 12th IEEE International Conference on Trust, Security and Privacy in Computing and Communications, Washington, DC, USA.","DOI":"10.1109\/TrustCom.2013.153"},{"key":"ref_11","unstructured":"Gao, C., Tung, A.K.H., Xu, X., Pan, F., and Yang, J. (2004, January 13\u201318). FARMER: Finding interesting rule groups in microarray datasets. Proceedings of the 2004 ACM SIGMOD International Conference on Management of Data, Paris, France."},{"key":"ref_12","unstructured":"Tania, C., and Di Corso, E. (2022, January 09). Characterizing Thermal Energy Consumption through Exploratory Data Mining Algorithms. Available online: https:\/\/iris.polito.it\/handle\/11583\/2639284."},{"key":"ref_13","unstructured":"Antonie, M., Zaiane, O.R., and Coman, A. (2001, January 26). Application of Data Mining Techniques for Medical Image Classification. Proceedings of the Second International Conference on Multimedia Data Mining, San Francisco, CA, USA."},{"key":"ref_14","unstructured":"Rakesh, A., and Srikant, R. (2022, January 09). Fast Algorithms for Mining Association Rules. Available online: https:\/\/dl.acm.org\/doi\/10.5555\/645920.672836."},{"key":"ref_15","unstructured":"(2022, January 04). Apriori Algorithm\u2014GeeksforGeeks. Available online: https:\/\/www.geeksforgeeks.org\/apriori-algorithm\/."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"372","DOI":"10.1109\/69.846291","article-title":"Scalable algorithms for association mining","volume":"12","author":"Zaki","year":"2000","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_17","unstructured":"(2022, January 04). ML|ECLAT Algorithm\u2014GeeksforGeeks. Available online: https:\/\/www.geeksforgeeks.org\/ml-eclat-algorithm\/."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zaki, M.J., and Gouda, K. (2003, January 24\u201327). Fast vertical mining using diffsets. Proceedings of the Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining\u2014KDD \u201903, Washington, DC, USA.","DOI":"10.1145\/956755.956788"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1165","DOI":"10.1007\/s10115-018-1248-0","article-title":"The big data system, components, tools, and technologies: A survey","volume":"60","author":"Rao","year":"2019","journal-title":"Knowl. Inf. Syst."},{"key":"ref_20","unstructured":"(2020, November 29). Big Data Analysis Using Apache Hadoop. Available online: https:\/\/www.researchgate.net\/publication\/261309523_Big_data_analysis_using_Apache_Hadoop."},{"key":"ref_21","unstructured":"(2020, November 28). Apache Hadoop. Available online: http:\/\/hadoop.apache.org\/."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Weets, J.-F., Kakhani, M.K., and Kumar, A. Limitations and challenges of HDFS and MapReduce. Proceedings of the 2015 International Conference on Green Computing and Internet of Things (ICGCIoT), NW Washington, DC, USA, 8\u201315 October 2015.","DOI":"10.1109\/ICGCIoT.2015.7380524"},{"key":"ref_23","unstructured":"(2020, December 23). Frequent Pattern Mining\u2014RDD-Based API\u2014Spark 2.2.0 Documentation. Available online: https:\/\/spark.apache.org\/docs\/2.2.0\/mllib-frequent-pattern-mining.html."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1007\/s41060-016-0027-9","article-title":"Big data analytics on Apache Spark","volume":"1","author":"Salloum","year":"2016","journal-title":"Int. J. Data Sci. Anal."},{"key":"ref_25","unstructured":"(2020, December 22). Frequent Pattern Mining\u2014Spark 3.0.1 Documentation. Available online: https:\/\/spark.apache.org\/docs\/latest\/ml-frequent-pattern-mining.html."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Cai, B.Z., Zhu, X., Zheng, Y., Liu, D., and Xu, L. (2018). A Caching-Based Parallel FP-Growth in Apache Spark, Springer International Publishing.","DOI":"10.1007\/978-3-030-05057-3_39"},{"key":"ref_27","unstructured":"(2021, May 27). BloomFilter (Spark 2.1.0 JavaDoc). Available online: https:\/\/spark.apache.org\/docs\/2.1.0\/api\/java\/org\/apache\/spark\/util\/sketch\/BloomFilter.html."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"3565","DOI":"10.1007\/s10115-020-01464-1","article-title":"EAFIM: Efficient apriori-based frequent itemset mining algorithm on Spark for big transactional data","volume":"62","author":"Raj","year":"2020","journal-title":"Knowl. Inf. Syst."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"6","DOI":"10.1186\/s40537-018-0112-0","article-title":"Adaptive-Miner: An efficient distributed association rule mining algorithm on Spark","volume":"5","author":"Rathee","year":"2018","journal-title":"J. Big Data"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"3652","DOI":"10.1007\/s11227-017-1963-4","article-title":"HFIM: A Spark-based hybrid frequent itemset mining algorithm for big data processing","volume":"73","author":"Sethi","year":"2017","journal-title":"J. Supercomput."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1493","DOI":"10.1007\/s10586-015-0477-1","article-title":"A distributed frequent itemset mining algorithm using Spark for Big Data analytics","volume":"18","author":"Zhang","year":"2015","journal-title":"Clust. Comput."},{"key":"ref_32","unstructured":"Li, H., Wang, Y., Zhang, D., Zhang, M., and Chang, E.Y. (2008, January 23\u201325). RecSys \u201908. Proceedings of the 2008 ACM Conference on Recommender Systems, Lausanne, Switzerland."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Rathee, S., Kaul, M., and Kashyap, A. (2015). R-Apriori. Proceedings of the 8th Workshop on Ph.D. Workshop in Information and Knowledge Management, ACM Press.","DOI":"10.1145\/2809890.2809893"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Qiu, H., Gu, R., Yuan, C., and Huang, Y. (2014, January 19\u201323). YAFIM: A Parallel Frequent Itemset Mining Algorithm with Spark. Proceedings of the 2014 IEEE International Parallel & Distributed Processing Symposium Workshops, Phoenix, AZ, USA.","DOI":"10.1109\/IPDPSW.2014.185"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"135144","DOI":"10.1109\/ACCESS.2021.3115514","article-title":"A Distributed Method for Fast Mining Frequent Patterns From Big Data","volume":"9","author":"Huang","year":"2021","journal-title":"IEEE Access"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"755","DOI":"10.1007\/978-3-030-37051-0_85","article-title":"RDD-Eclat: Approaches to Parallelize Eclat Algorithm on Spark RDD Framework","volume":"Volume 44","author":"Singh","year":"2020","journal-title":"Lecture Notes on Data Engineering and Communications Technologies"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Leung, C.K., Zhang, H., Souza, J., and Lee, W. (2018). Scalable Vertical Mining for Big Data Analytics of Frequent Itemsets. Lecture Notes in Computer Science (In-cluding Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Springer International Publishing.","DOI":"10.1007\/978-3-319-98809-2_1"},{"key":"ref_38","first-page":"401","article-title":"Parallel Eclat for Opportunistic Mining of Frequent Itemsets","volume":"Volume 9261","author":"Liu","year":"2015","journal-title":"Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Moens, S., Aksehirli, E., and Goethals, B. (2013). Frequent Itemset Mining for Big Data. 2013 IEEE Int. Conf. Big Data, 111\u2013118.","DOI":"10.1109\/BigData.2013.6691742"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1016\/j.future.2019.09.041","article-title":"Map-optimize-reduce: CAN tree assisted FP-growth algorithm for clusters based FP mining on Hadoop","volume":"103","author":"Ragaventhiran","year":"2020","journal-title":"Futur. Gener. Comput. Syst."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Shi, X., Chen, S., and Yang, H. (2017, January 25\u201326). DFPS: Distributed FP-growth algorithm based on Spark. Proceedings of the 2017 IEEE 2nd Advanced Information Technology, Electronic and Automation Control Conference (IAEAC), Chongqing, China.","DOI":"10.1109\/IAEAC.2017.8054308"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/335191.335372","article-title":"Mining frequent patterns without candidate generation","volume":"29","author":"Han","year":"2000","journal-title":"ACM SIGMOD Rec."},{"key":"ref_43","unstructured":"(2014). Frequent Pattern Mining. Freq. Pattern Min., 9783319078212, 1\u2013471."},{"key":"ref_44","unstructured":"(2020, December 12). Frequent Itemset Mining Dataset Repository. Available online: http:\/\/fimi.uantwerpen.be\/data\/."}],"container-title":["Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2306-5729\/7\/1\/11\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T14:15:01Z","timestamp":1760364901000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2306-5729\/7\/1\/11"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,14]]},"references-count":44,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,1]]}},"alternative-id":["data7010011"],"URL":"https:\/\/doi.org\/10.3390\/data7010011","relation":{},"ISSN":["2306-5729"],"issn-type":[{"value":"2306-5729","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,14]]}}}