{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T17:17:21Z","timestamp":1782407841052,"version":"3.54.5"},"reference-count":44,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2021,7,21]],"date-time":"2021-07-21T00:00:00Z","timestamp":1626825600000},"content-version":"vor","delay-in-days":201,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Complexity"],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:p>The amount of data produced in scientific and commercial fields is growing dramatically. Correspondingly, big data technologies, such as Hadoop and Spark, have emerged to tackle the challenges of collecting, processing, and storing such large\u2010scale data. Unfortunately, big data applications usually have performance issues and do not fully exploit a hardware infrastructure. One reason is that applications are developed using high\u2010level programming languages that do not provide low\u2010level system control in terms of performance of highly parallel programming models like message passing interface (MPI). Moreover, big data is considered a barrier of parallel programming models or accelerators (e.g., CUDA and OpenCL). Therefore, the aim of this study is to investigate how the performance of big data applications can be enhanced without sacrificing the power consumption of a hardware infrastructure. A Hybrid Spark MPI OpenACC (HSMO) system is proposed for integrating Spark as a big data programming model, with MPI and OpenACC as parallel programming models. Such integration brings together the advantages of each programming model and provides greater effectiveness. To enhance performance without sacrificing power consumption, the integration approach needs to exploit the hardware infrastructure in an intelligent manner. For achieving this performance enhancement, a mapping technique is proposed that is built based on the application\u2019s virtual topology as well as the physical topology of the undelaying resources. To the best of our knowledge, there is no existing method in big data applications related to utilizing graphics processing units (GPUs), which are now an essential part of high\u2010performance computing (HPC) as a powerful resource for fast computation.<\/jats:p>","DOI":"10.1155\/2021\/9943289","type":"journal-article","created":{"date-parts":[[2021,7,21]],"date-time":"2021-07-21T17:52:02Z","timestamp":1626889922000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Accelerating Spark\u2010Based Applications with MPI and OpenACC"],"prefix":"10.1155","volume":"2021","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0499-7170","authenticated-orcid":false,"given":"Saeed","family":"Alshahrani","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2862-4897","authenticated-orcid":false,"given":"Waleed","family":"Al Shehri","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6288-6474","authenticated-orcid":false,"given":"Jameel","family":"Almalki","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7644-5039","authenticated-orcid":false,"given":"Ahmed M.","family":"Alghamdi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8714-4749","authenticated-orcid":false,"given":"Abdullah M.","family":"Alammari","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2021,7,21]]},"reference":[{"key":"e_1_2_11_1_2","doi-asserted-by":"crossref","unstructured":"BergamaschiS. CavazzoniC. CurioniA. andFoxG. New opportunities in high performance data analytics (HPDA) and high performance computing (HPC) Proceedings of 2014 International Conference on High Performance Computing & Simulation (HPCS) July 2014 Bologna Italy https:\/\/doi.org\/10.1109\/HPCSim.2014.6903656.","DOI":"10.1109\/HPCSim.2014.6903656"},{"key":"e_1_2_11_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2011.2155130"},{"key":"e_1_2_11_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2018.01.034"},{"key":"e_1_2_11_4_2","doi-asserted-by":"crossref","unstructured":"DeRoseL. The path to delivering programable exascale systems Proceedings of 2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS) May 2019 Rio de Janerio Brazil https:\/\/doi.org\/10.1109\/IPDPS.2019.00081.","DOI":"10.1109\/IPDPS.2019.00081"},{"key":"e_1_2_11_5_2","doi-asserted-by":"publisher","DOI":"10.25046\/aj040105"},{"key":"e_1_2_11_6_2","doi-asserted-by":"crossref","unstructured":"UtaA. VarbanescuA. L. MusaafirA. LemaireC. andIosupA. Exploring HPC and big data convergence: a graph processing study on intel knights landing Proceedings of 2018 IEEE International Conference on Cluster Computing (CLUSTER) 2018-September Belfast UK 66\u201377 https:\/\/doi.org\/10.1109\/CLUSTER.2018.00019 2-s2.0-85057283772.","DOI":"10.1109\/CLUSTER.2018.00019"},{"key":"e_1_2_11_7_2","doi-asserted-by":"crossref","unstructured":"GittensA. DevarakondaA. RacahE.et al. Matrix factorizations at scale: a comparison of scientific data analytics in spark and C+MPI using three case studies Proceedings of 2016 IEEE International Conference on Big Data (Big Data) December 2016 Washington DC USA 204\u2013213 https:\/\/doi.org\/10.1109\/BigData.2016.7840606 2-s2.0-85015251232.","DOI":"10.1109\/BigData.2016.7840606"},{"key":"e_1_2_11_8_2","first-page":"149","article-title":"Evaluation of high-performance computing techniques for big data applications","volume":"31","author":"Al Shehri W.","year":"2019","journal-title":"Science International"},{"key":"e_1_2_11_9_2","first-page":"561","volume-title":"Big Data and HPC Convergence for Smart Infrastructures: A Review and Proposed Architecture","author":"Usman S.","year":"2020"},{"key":"e_1_2_11_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/2699414"},{"key":"e_1_2_11_11_2","first-page":"81","article-title":"A hybrid spark MPI OpenACC system","volume":"19","author":"Al Shehri W.","year":"2019","journal-title":"International Journal of Computer Science and Network Society"},{"key":"e_1_2_11_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2703149"},{"key":"e_1_2_11_13_2","doi-asserted-by":"crossref","unstructured":"LeeG. ToliaN. RanganathanP. andKatzR. H. Topology-aware resource allocation for data-intensive workloads Proceedings of the First ACM Asia-Pacific Workshop on Workshop on Systems-APSys \u203210 August 2010 Hangzhou China https:\/\/doi.org\/10.1145\/1851276.1851278 2-s2.0-78149326439.","DOI":"10.1145\/1851276.1851278"},{"key":"e_1_2_11_14_2","doi-asserted-by":"crossref","unstructured":"RamanR. LivnyM. andSolomonM. Matchmaking: distributed resource management for high throughput computing Proceedings of the Seventh International Symposium on High Performance Distributed Computing July 1998 Chicago IL USA 140\u2013146 https:\/\/doi.org\/10.1109\/HPDC.1998.709966 2-s2.0-84974666914.","DOI":"10.1109\/HPDC.1998.709966"},{"key":"e_1_2_11_15_2","doi-asserted-by":"crossref","unstructured":"MohanamuralyP.andStaffelbachG. Hardware locality-aware partitioning and dynamic load-balancing of unstructured meshes for large-scale scientific applications Proceedings of the Platform for Advanced Scientific Computing Conference June 2020 Geneva Switzerland https:\/\/doi.org\/10.1145\/3394277.3401851.","DOI":"10.1145\/3394277.3401851"},{"key":"e_1_2_11_16_2","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.6265"},{"key":"e_1_2_11_17_2","doi-asserted-by":"crossref","unstructured":"WangK. ZhouX. LiT. ZhaoD. LangM. andRaicuI. Optimizing load balancing and data-locality with data-aware scheduling Proceedings of 2014 IEEE International Conference on Big Data (Big Data) October 2014 Washington DC USA 119\u2013128 https:\/\/doi.org\/10.1109\/BigData.2014.7004220 2-s2.0-84921748797.","DOI":"10.1109\/BigData.2014.7004220"},{"key":"e_1_2_11_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10586-018-2634-9"},{"key":"e_1_2_11_19_2","doi-asserted-by":"crossref","unstructured":"Di StefanoA. Di StefanoA. MoranaG. andZitoD. Coope4M: a deployment framework for communication-intensive applications on mesos Proceedings of 2018 IEEE 27th International Conference on Enabling Technologies: Infrastructure for Collaborative Enterprises (WETICE) June 2018 Paris France 36\u201341 https:\/\/doi.org\/10.1109\/WETICE.2018.00014 2-s2.0-85057301886.","DOI":"10.1109\/WETICE.2018.00014"},{"key":"e_1_2_11_20_2","doi-asserted-by":"crossref","unstructured":"IsardM. PrabhakaranV. CurreyJ. WiederU. TalwarK. andGoldbergA. Quincy Proceedings of the ACM SIGOPS 22nd Symposium on Operating Systems Principles-SOSP\u201909 October 2009 Big Sky MT USA https:\/\/doi.org\/10.1145\/1629575.1629601 2-s2.0-72249118633.","DOI":"10.1145\/1629575.1629601"},{"key":"e_1_2_11_21_2","doi-asserted-by":"crossref","unstructured":"YinJ. ForanA. andWangJ. DL-MPI: enabling data locality computation for MPI-based data-intensive applications Proceedings of 2013 IEEE International Conference on Big Data October 2013 Silicon Valley CA USA 506\u2013511 https:\/\/doi.org\/10.1109\/BigData.2013.6691614 2-s2.0-84893218367.","DOI":"10.1109\/BigData.2013.6691614"},{"key":"e_1_2_11_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2011.99"},{"key":"e_1_2_11_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-019-02826-5"},{"key":"e_1_2_11_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2018.108162940"},{"key":"e_1_2_11_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/J.JPDC.2018.02.028"},{"key":"e_1_2_11_26_2","doi-asserted-by":"publisher","DOI":"10.1016\/J.JOCS.2019.01.003"},{"key":"e_1_2_11_27_2","doi-asserted-by":"crossref","unstructured":"MastorasA.andGrossT. R. Understanding parallelization tradeoffs for linear pipelines Proceedings of the 9th International Workshop on Programming Models and Applications for Multicores and Manycores February 2018 Vienna Austria 1\u201310 https:\/\/doi.org\/10.1145\/3178442.3178443 2-s2.0-85051584972.","DOI":"10.1145\/3178442.3178443"},{"key":"e_1_2_11_28_2","doi-asserted-by":"crossref","unstructured":"MeadeA. BuckleyJ. andCollinsJ. J. Challenges of evolving sequential to parallel code Proceedings of the 12th International Workshop and the 7th Annual ERCIM Workshop on Principles on Software Evolution and Software Evolution-IWPSE-EVOL\u201911 September 2011 Szeged Hungary https:\/\/doi.org\/10.1145\/2024445.2024447 2-s2.0-80053199787.","DOI":"10.1145\/2024445.2024447"},{"key":"e_1_2_11_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/J.PROCS.2015.07.286"},{"key":"e_1_2_11_30_2","doi-asserted-by":"crossref","unstructured":"JhaS. QiuJ. LuckowA. ManthaP. andFoxG. C. A tale of two data-intensive paradigms: applications abstractions and architectures Proceedings of 2014 IEEE International Congress on Big Data June 2014 Anchorage AK USA 645\u2013652 https:\/\/doi.org\/10.1109\/BigData.Congress.2014.137 2-s2.0-84923884968.","DOI":"10.1109\/BigData.Congress.2014.137"},{"key":"e_1_2_11_31_2","doi-asserted-by":"crossref","unstructured":"SatishN. SundaramN. PatwaryM. M. A.et al. Navigating the maze of graph analytics frameworks using massive graph datasets Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data June 2014 Snowbird UT USA 979\u2013990 https:\/\/doi.org\/10.1145\/2588555.2610518 2-s2.0-84904339615.","DOI":"10.1145\/2588555.2610518"},{"key":"e_1_2_11_32_2","doi-asserted-by":"crossref","unstructured":"SlotaG. M. RajamanickamS. andMadduriK. A case study of complex graph analysis in distributed memory: implementation and optimization Proceedings of 2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS) May 2016 Chicago IL USA 293\u2013302 https:\/\/doi.org\/10.1109\/IPDPS.2016.93 2-s2.0-84983234050.","DOI":"10.1109\/IPDPS.2016.93"},{"key":"e_1_2_11_33_2","unstructured":"OusterhoutK. RastiR. RatnasamyS. ShenkerS. andChunB.-G. Making sense of performance in data analytics frameworks Proceedings of the 12th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201915) March 2015 Santa Clara CA USA 293\u2013307."},{"key":"e_1_2_11_34_2","doi-asserted-by":"publisher","DOI":"10.14778\/3090163.3090168"},{"key":"e_1_2_11_35_2","doi-asserted-by":"crossref","unstructured":"LuX. LiangF. WangB. ZhaL. andXuZ. DataMPI: extending MPI to hadoop-like big data computing Proceedings of 2014 IEEE 28th International Parallel and Distributed Processing Symposium May 2014 Phoenix AZ USA 829\u2013838 https:\/\/doi.org\/10.1109\/IPDPS.2014.90 2-s2.0-84906654880.","DOI":"10.1109\/IPDPS.2014.90"},{"key":"e_1_2_11_36_2","unstructured":"\u201cH2O.ai. Sparkling Water.\u201dhttps:\/\/github.com\/h2oai\/sparkling-water (accessed Aug. 14 2018)."},{"key":"e_1_2_11_37_2","unstructured":"\u201cdeeplearning4j. Deep Learning for Java. Open-Source Distributed Deep Learning Library for the JVM.\u201dhttps:\/\/deeplearning4j.org\/ (accessed Aug. 14 2018)."},{"key":"e_1_2_11_38_2","doi-asserted-by":"crossref","unstructured":"GrossmanM.andSarkarV. SWAT Proceedings of the 25th ACM International Symposium on High-Performance Parallel and Distributed Computing June 2016 Kyoto Japan 81\u201392 https:\/\/doi.org\/10.1145\/2907294.2907307 2-s2.0-84978496078.","DOI":"10.1145\/2907294.2907307"},{"key":"e_1_2_11_39_2","doi-asserted-by":"crossref","unstructured":"GittensA. RothaugeK. WangS.et al. Accelerating large-scale data analysis by offloading to high-performance computing libraries using alchemist Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining August 2018 London UK 293\u2013301 https:\/\/doi.org\/10.1145\/3219819.3219927 2-s2.0-85051537878.","DOI":"10.1145\/3219819.3219927"},{"key":"e_1_2_11_40_2","unstructured":"L. C. Platforms P. Strategies and Y. Yan \u201cIntroduction to OpenACC Directives \u201d 2012."},{"key":"e_1_2_11_41_2","unstructured":"Databricks \u201cIntro to Apache Spark \u201d 2015.https:\/\/stanford.edu\/\u223crezab\/sparkclass\/slides\/itas_workshop.pdf."},{"key":"e_1_2_11_42_2","unstructured":"\u201cMPI_Cart_create(MPI_Comm Comm_old Int Ndims Int \u2217dims Int \u2217periods Int Reorder MPI_Comm \u2217comm_cart) Function.\u201dhttps:\/\/mpi.deino.net\/mpi_functions\/MPI_Cart_create.html (accessed Oct. 18 2019)."},{"key":"e_1_2_11_43_2","unstructured":"\u201cSNAP: network datasets: social circles.\u201dhttps:\/\/snap.stanford.edu\/data\/ego-Twitter.html (accessed Dec. 01 2019)."},{"key":"e_1_2_11_44_2","unstructured":"\u201cUbuntu Manpage: Powerstat-a tool to measure power consumption.\u201dhttp:\/\/manpages.ubuntu.com\/manpages\/xenial\/man8\/powerstat.8.html (accessed Oct. 18 2019)."}],"container-title":["Complexity"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/complexity\/2021\/9943289.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/complexity\/2021\/9943289.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1155\/2021\/9943289","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,9]],"date-time":"2024-08-09T21:44:37Z","timestamp":1723239877000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1155\/2021\/9943289"}},"subtitle":[],"editor":[{"given":"Adil Mehmood","family":"Khan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2021,1]]},"references-count":44,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["10.1155\/2021\/9943289"],"URL":"https:\/\/doi.org\/10.1155\/2021\/9943289","archive":["Portico"],"relation":{},"ISSN":["1076-2787","1099-0526"],"issn-type":[{"value":"1076-2787","type":"print"},{"value":"1099-0526","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,1]]},"assertion":[{"value":"2021-03-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-07-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-07-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"9943289"}}