{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T17:34:41Z","timestamp":1754156081219,"version":"3.41.2"},"reference-count":36,"publisher":"Emerald","issue":"4","license":[{"start":{"date-parts":[[2021,8,6]],"date-time":"2021-08-06T00:00:00Z","timestamp":1628208000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.emerald.com\/insight\/site-policies"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJWIS"],"published-print":{"date-parts":[[2021,9,6]]},"abstract":"<jats:sec><jats:title content-type=\"abstract-subheading\">Purpose<\/jats:title><jats:p>This paper aims to evaluate different approaches for the parallelization of compute-intensive tasks. The study compares a Java multi-threaded algorithm, distributed computing solutions with MapReduce (Apache Hadoop) and resilient distributed data set (RDD) (Apache Spark) paradigms and a graphics processing unit (GPU) approach with Numba for compute unified device architecture (CUDA).<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Design\/methodology\/approach<\/jats:title><jats:p>The paper uses a simple but computationally intensive puzzle as a case study for experiments. To find all solutions using brute force search, 15! permutations had to be computed and tested against the solution rules. The experimental application comprises a Java multi-threaded algorithm, distributed computing solutions with MapReduce (Apache Hadoop) and RDD (Apache Spark) paradigms and a GPU approach with Numba for CUDA. The implementations were benchmarked on Amazon-EC2 instances for performance and scalability measurements.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Findings<\/jats:title><jats:p>The comparison of the solutions with Apache Hadoop and Apache Spark under Amazon EMR showed that the processing time measured in CPU minutes with Spark was up to 30% lower, while the performance of Spark especially benefits from an increasing number of tasks. With the CUDA implementation, more than 16 times faster execution is achievable for the same price compared to the Spark solution. Apart from the multi-threaded implementation, the processing times of all solutions scale approximately linearly. Finally, several application suggestions for the different parallelization approaches are derived from the insights of this study.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Originality\/value<\/jats:title><jats:p>There are numerous studies that have examined the performance of parallelization approaches. Most of these studies deal with processing large amounts of data or mathematical problems. This work, in contrast, compares these technologies on their ability to implement computationally intensive distributed algorithms.<\/jats:p><\/jats:sec>","DOI":"10.1108\/ijwis-03-2021-0032","type":"journal-article","created":{"date-parts":[[2021,8,5]],"date-time":"2021-08-05T05:37:33Z","timestamp":1628141853000},"page":"377-402","source":"Crossref","is-referenced-by-count":3,"title":["Performance evaluation of GPU- and cluster-computing for parallelization of compute-intensive tasks"],"prefix":"10.1108","volume":"17","author":[{"given":"Alexander","family":"D\u00f6schl","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Max-Emanuel","family":"Keller","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peter","family":"Mandl","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","published-online":{"date-parts":[[2021,8,6]]},"reference":[{"key":"key2021100813195334600_ref001","unstructured":"Amazon (2021a), \u201cAmazon s3\u201d, available at: https:\/\/aws.amazon.com\/de\/s3\/"},{"key":"key2021100813195334600_ref002","unstructured":"Amazon (2021b), \u201cAmazon EMR\u201d, available at: https:\/\/aws.amazon.com\/de\/emr\/"},{"key":"key2021100813195334600_ref003","unstructured":"Amazon (2021c), \u201cAmazon EC2\u201d, available at: https:\/\/aws.amazon.com\/de\/ec2\/"},{"key":"key2021100813195334600_ref004","unstructured":"Amazon (2021d), \u201cAmazon web services\u201d, available at: https:\/\/aws.amazon.com\/de\/"},{"key":"key2021100813195334600_ref006","unstructured":"Apache (2021), \u201cApache hadoop\u201d, available at: http:\/\/hadoop.apache.org\/"},{"key":"key2021100813195334600_ref007","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1145\/2939672.2939675","article-title":"Matrix computations and optimization in apache spark","volume-title":"In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","year":"2016"},{"key":"key2021100813195334600_ref008","unstructured":"CUDA (2021a), \u201cCUDA \u2013 thomas-krenn-wiki\u201d, available at: www.thomas-krenn.com\/de\/wiki\/CUDA"},{"key":"key2021100813195334600_ref009","unstructured":"CUDA (2021b), \u201cCUDA occupancy calculator:: CUDA toolkit documentation\u201d, available at: https:\/\/docs.nvidia.com\/cuda\/cuda-occupancy-calculator\/index.html"},{"key":"key2021100813195334600_ref010","unstructured":"Dask (2021), \u201cDask: Scalable analytics in python\u201d, available at: https:\/\/dask.org\/"},{"key":"key2021100813195334600_ref011","doi-asserted-by":"publisher","first-page":"313","DOI":"10.1145\/3428757.3429121","article-title":"Performance evaluation of apache hadoop and apache spark for parallelization of compute-intensive tasks","volume-title":"In Proceedings of the 22nd International Conference on Information Integration and Web-based Applications and ServicesISBN 978-1-4503-8922-8","year":"2020"},{"year":"2021","key":"key2021100813195334600_ref011a","article-title":"Permutation-games"},{"issue":"4","key":"key2021100813195334600_ref012","doi-asserted-by":"publisher","first-page":"13","DOI":"10.1109\/MM.2008.57","article-title":"Parallel computing experiences with CUDA","volume":"28","year":"2008","journal-title":"IEEE Micro"},{"key":"key2021100813195334600_ref013","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1109\/IC3INA.2017.8251736","article-title":"Performance factors of a CUDA GPU parallel program: a case study on a PDF password cracking brute-force algorithm","volume-title":"In 2017 International Conference on Computer, Control, Informatics and its Applications (IC3INA)ISBN 978-1-5386-3978-8","year":"2017"},{"key":"key2021100813195334600_ref014","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/ICASI.2016.7539833","article-title":"Detection DDoS attacks based on neural-network using apache spark","volume-title":"In 2016 International Conference on Applied System Innovation (ICASI)ISBN 978-1-4673-9888-6","year":"2016"},{"key":"key2021100813195334600_ref015","doi-asserted-by":"publisher","first-page":"420","DOI":"10.1109\/TELFOR.2018.8611982","article-title":"How CUDA powers the machine learning revolution","volume-title":"In 2018 26th Telecommunications Forum (TELFOR)","year":"2018"},{"key":"key2021100813195334600_ref016","unstructured":"Jao Cerreia (2021), \u201cJao cerreia\u201d, available at: www.joaocorreia.de"},{"key":"key2021100813195334600_ref017","doi-asserted-by":"crossref","unstructured":"Keller, M.E., Mandl, P., D\u00f6schl, A., Kailer, D. and Grimm, M. (2017), \u201cVerarbeitung komplexer XML-basierter massendaten in BigData-anwendungen\u201d, p. 6, ISSN 2296-4592, available at: https:\/\/ojs-hslu.ch\/ojs302\/index.php\/AKWI\/article\/view\/93","DOI":"10.26034\/lu.akwi.2017.3181"},{"issue":"8","key":"key2021100813195334600_ref018","doi-asserted-by":"publisher","first-page":"4353","DOI":"10.1007\/s00521-018-3354-z","article-title":"CPU versus GPU: which can perform matrix computation faster \u2013 performance comparison for basic linear algebra subprograms","volume":"31","year":"2018","journal-title":"Neural Computing and Applications"},{"key":"key2021100813195334600_ref019","unstructured":"McDonald, C., Evans, R. and Lowe, J. (2020), \u201cAccelerating apache spark 3.0 with GPUs and RAPIDS\u201d, available at: https:\/\/developer.nvidia.com\/blog\/accelerating-apache-spark-3-0-with-gpus-and-rapids\/"},{"key":"key2021100813195334600_ref020","unstructured":"Mandl, P. and D\u00f6schl, A. (2017), \u201cKlassisches multi-threading versus MapReduce zur parallelisierung rechenintensiver tasks in der amazon cloud\u201d, ISSN 1436-3011, 2198-2775, doi: 10.1365\/s40702-017-0360-z, available at: http:\/\/link.springer.com\/10.1365\/s40702-017-0360-z"},{"key":"key2021100813195334600_ref021","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3110355.3110356","article-title":"Benchmarking OpenCL, OpenACC, OpenMP, and CUDA: Programming productivity, performance, and energy consumption","volume-title":"In Proceedings of the 2017 Workshop on Adaptive Resource Management and Scheduling for Cloud Computing \u2013 ARMS-CC \u201817ISBN 978-1-4503-5116-4","year":"2017"},{"issue":"1","key":"key2021100813195334600_ref022","first-page":"1235","article-title":"MLlib: Machine learning in apache spark","volume":"17","year":"2016","journal-title":"The Journal of Machine Learning ResearchISSN 1532-4435"},{"key":"key2021100813195334600_ref023","doi-asserted-by":"publisher","first-page":"188","DOI":"10.1109\/BIBE.2017.00-57","article-title":"Streaming distributed DNA sequence alignment using apache spark","volume-title":"In 2017 IEEE 17th International Conference on Bioinformatics and Bioengineering (BIBE)ISBN 978-1-5386-1324-5","year":"2017"},{"issue":"2","key":"key2021100813195334600_ref024","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1145\/1365490.1365500","article-title":"Scalable parallel programming with CUDA: is CUDA the parallel programming model that application developers have been waiting for?","volume":"6","year":"2008","journal-title":"Queue"},{"key":"key2021100813195334600_ref025","unstructured":"Ouellet, E. and Saad, O. (2018), \u201cPermutations: fast implementations and a new indexing algorithm allowing multithreading\u201d, available at: www.codeproject.com\/Articles\/1250925\/Permutations-Fast-implementations-and-a-new-indexi"},{"issue":"5","key":"key2021100813195334600_ref026","doi-asserted-by":"publisher","first-page":"879","DOI":"10.1109\/JPROC.2008.917757","article-title":"GPU computing","volume":"96","year":"2008","journal-title":"Proceedings of the IEEEISSN 0018-9219, 1558-2256"},{"key":"key2021100813195334600_ref027","unstructured":"Programming guide (2021), \u201cProgramming guide: CUDA toolkit documentation\u201d, available at: https:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/index.html"},{"issue":"12","key":"key2021100813195334600_ref028","doi-asserted-by":"publisher","first-page":"e4367","DOI":"10.1002\/cpe.4367","article-title":"Performance comparison between hadoop and spark frameworks using HiBench benchmarks: performance comparison between hadoop and spark frameworks using HiBench benchmarks","volume":"30","year":"2017","journal-title":"Concurrency and Computation: Practice and Experience"},{"year":"2008","key":"key2021100813195334600_ref029","article-title":"Multicore processors \u2013 a necessity"},{"key":"key2021100813195334600_ref030","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2541583.2541588","article-title":"Clotho: an elastic MapReduce workload\/runtime co-design","volume-title":"In Proceedings of the 12th International Workshop on Adaptive and Reflective Middleware \u2013 ARM \u201813","year":"2013"},{"key":"key2021100813195334600_ref031","unstructured":"Springer, M. (2019), \u201cMemory-efficient object-oriented programming on GPUs\u201d, available at: http:\/\/arxiv.org\/abs\/1908.05845"},{"key":"key2021100813195334600_ref032","doi-asserted-by":"publisher","first-page":"5600462","DOI":"10.1109\/PCI.2010.47","article-title":"Parallel collection of live data using hadoop","volume-title":"In 2010 14th Panhellenic Conference on Informatics","year":"2010"},{"volume-title":"Hadoop: zuverl\u00e4ssige, Verteilte Und Skalierbare Big-Data-Anwendungen","year":"2012","key":"key2021100813195334600_ref033"},{"issue":"3","key":"key2021100813195334600_ref034","doi-asserted-by":"publisher","first-page":"1","DOI":"10.7236\/IJIBC.2019.11.3.1","volume":"11","year":"2019","journal-title":"Performance Comparison of Parallel Programming Frameworks in Digital Image Transformation"},{"key":"key2021100813195334600_ref005","unstructured":"Word Aligned (2009), \u201cNext permutation: When c++ gets it right\u201d, available at: http:\/\/wordaligned.org\/articles\/next-permutation"},{"key":"key2021100813195334600_ref035","first-page":"15","article-title":"Resilient distributed datasets: a fault-tolerant abstraction for in-memory cluster computing","volume-title":"In Presented as part of the 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12)ISBN 978-931971-92-8","year":"2012"}],"container-title":["International Journal of Web Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/IJWIS-03-2021-0032\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/IJWIS-03-2021-0032\/full\/html","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T22:23:51Z","timestamp":1753395831000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.emerald.com\/ijwis\/article\/17\/4\/377-402\/165696"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,6]]},"references-count":36,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2021,8,6]]},"published-print":{"date-parts":[[2021,9,6]]}},"alternative-id":["10.1108\/IJWIS-03-2021-0032"],"URL":"https:\/\/doi.org\/10.1108\/ijwis-03-2021-0032","relation":{},"ISSN":["1744-0084","1744-0084"],"issn-type":[{"type":"print","value":"1744-0084"},{"type":"print","value":"1744-0084"}],"subject":[],"published":{"date-parts":[[2021,8,6]]}}}