{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,5]],"date-time":"2025-10-05T14:37:14Z","timestamp":1759675034614,"version":"3.37.3"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2017,9,13]],"date-time":"2017-09-13T00:00:00Z","timestamp":1505260800000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"funder":[{"DOI":"10.13039\/501100001602","name":"Science Foundation Ireland","doi-asserted-by":"publisher","award":["14\/IA\/2474"],"award-info":[{"award-number":["14\/IA\/2474"]}],"id":[{"id":"10.13039\/501100001602","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2018,2]]},"DOI":"10.1007\/s11227-017-2141-4","type":"journal-article","created":{"date-parts":[[2017,9,13]],"date-time":"2017-09-13T00:57:05Z","timestamp":1505264225000},"page":"551-568","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Out-of-core implementation for accelerator kernels on heterogeneous clouds"],"prefix":"10.1007","volume":"74","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4070-7468","authenticated-orcid":false,"given":"Hamidreza","family":"Khaleghzadeh","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ziming","family":"Zhong","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ravi","family":"Reddy","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alexey","family":"Lastovetsky","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2017,9,13]]},"reference":[{"key":"2141_CR1","doi-asserted-by":"crossref","unstructured":"Filelis-Papadopoulos CK, Grylonakis ENG, Kyziropoulos PE, Gravvanis GA, Morrison JP (2016) Characterization of hardware in self-managing self-organizing cloud environment. In: Proceedings of the 20th Pan-Hellenic Conference on Informatics, Series PCI \u201916. ACM, pp 56:1\u201356:6","DOI":"10.1145\/3003733.3003749"},{"key":"2141_CR2","doi-asserted-by":"crossref","unstructured":"Lynn T, Xiong H, Dong D, Momani B, Gravvanis GA, Filelis-Papadopoulos CK, Elster AC, Khan MM, Tzovaras D, Giannoutakis KM, Petcu D, Neagul M, Dragon I, Kuppudayar P, Natarajan S, McGrath M, Gaydadjiev G, Becker T, Gourinovitch A, Kenny D, Morrison J (2016) CLOUDLIGHTNING: a framework for a self-organising and self-managing heterogeneous cloud. In: Proceedings of the 6th International Conference on Cloud Computing and Services Science, vols 1 and 2, Series CLOSER 2016. SCITEPRESS - Science and Technology Publications, Lda pp 333\u2013338","DOI":"10.5220\/0005921503330338"},{"key":"2141_CR3","first-page":"1","volume":"99","author":"CH Hong","year":"2017","unstructured":"Hong CH, Spence I, Nikolopoulos D (2017) FairGV: fair and fast GPU virtualization. IEEE Trans Parallel Distrib Syst 99:1\u20131","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"2141_CR4","unstructured":"CUBLAS-XT (2016) CUBLAS-XT: multi-GPU version of CUBLAS library supporting out-of-core routines. \n                        https:\/\/developer.nvidia.com\/cublas"},{"issue":"5\u20136","key":"2141_CR5","doi-asserted-by":"crossref","first-page":"232","DOI":"10.1016\/j.parco.2009.12.005","volume":"36","author":"S Tomov","year":"2010","unstructured":"Tomov S, Dongarra J, Baboulin M (2010) Towards dense linear algebra for hybrid GPU accelerated manycore systems. Parallel Comput 36(5\u20136):232\u2013240","journal-title":"Parallel Comput"},{"key":"2141_CR6","unstructured":"Khronos OpenCL Registry (2017) OpenCL command queues. \n                        https:\/\/www.khronos.org\/registry\/OpenCL\/specs\/opencl-2.2.pdf"},{"key":"2141_CR7","unstructured":"Intel (2017) Programming for Intel MIC architecture. \n                        https:\/\/software.intel.com\/en-us\/node\/684368"},{"key":"2141_CR8","unstructured":"NVIDIA (2016) CUDA C programming guide. \n                        https:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/"},{"key":"2141_CR9","unstructured":"NVIDIA (2013) Tesla K40 GPU accelerator. \n                        http:\/\/www.nvidia.com\/content\/PDF\/kepler\/Tesla-K40-PCIe-Passive-Board-Spec-BD-06902-001_v05.pdf"},{"key":"2141_CR10","unstructured":"Ostermann S, Iosup A, Yigitbasi N, Prodan R, Fahringer T, Epema D (2009) A performance analysis of EC2 cloud computing services for scientific computing. In: International Conference on Cloud Computing. Springer, pp 115\u2013131"},{"issue":"6","key":"2141_CR11","doi-asserted-by":"crossref","first-page":"931","DOI":"10.1109\/TPDS.2011.66","volume":"22","author":"A Iosup","year":"2011","unstructured":"Iosup A, Ostermann S, Yigitbasi MN, Prodan R, Fahringer T, Epema D (2011) Performance analysis of cloud computing services for many-tasks scientific computing. IEEE Trans Parallel Distrib Syst 22(6):931\u2013945","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"2141_CR12","doi-asserted-by":"crossref","unstructured":"Gupta A, Kal\u00e9 LV, Milojicic D, Faraboschi P, Balle SM (2013) HPC-aware VM placement in infrastructure clouds. In: 2013 IEEE International Conference on Cloud Engineering (IC2E), Mar 2013, pp 11\u201320","DOI":"10.1109\/IC2E.2013.38"},{"issue":"4","key":"2141_CR13","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1109\/MCSE.2013.49","volume":"15","author":"M Parashar","year":"2013","unstructured":"Parashar M, AbdelBaky M, Rodero I, Devarakonda A (2013) Cloud paradigms and practices for computational and data-enabled science and engineering. Comput Sci Eng 15(4):10\u201318","journal-title":"Comput Sci Eng"},{"issue":"6","key":"2141_CR14","doi-asserted-by":"crossref","first-page":"1408","DOI":"10.1016\/j.future.2012.03.011","volume":"29","author":"V Mauch","year":"2013","unstructured":"Mauch V, Kunze M, Hillenbrand M (2013) High performance cloud computing. Future Gener Comput Syst 29(6):1408\u20131416","journal-title":"Future Gener Comput Syst"},{"key":"2141_CR15","volume-title":"A GPGPU transparent virtualization component for high performance computing clouds","author":"G Giunta","year":"2010","unstructured":"Giunta G, Montella R, Agrillo G, Coviello G (2010) A GPGPU transparent virtualization component for high performance computing clouds. Springer, Berlin"},{"key":"2141_CR16","doi-asserted-by":"crossref","unstructured":"Byma S, Steffan JG, Bannazadeh H, Garcia AL, Chow P (2014) FPGAs in the cloud: booting virtualized hardware accelerators with OpenStack. In: 2014 IEEE 22nd Annual International Symposium on Field-Programmable Custom Computing Machines, May 2014, pp 109\u2013116","DOI":"10.1109\/FCCM.2014.42"},{"issue":"3","key":"2141_CR17","first-page":"35","volume":"50","author":"C-H Hong","year":"2017","unstructured":"Hong C-H, Spence I, Nikolopoulos DS (2017) GPU virtualization and scheduling methods: a comprehensive survey. ACM Comput Surv (CSUR) 50(3):35","journal-title":"ACM Comput Surv (CSUR)"},{"key":"2141_CR18","doi-asserted-by":"crossref","unstructured":"Gu L, Siegel J, Li X (2011) Using GPUs to compute large out-of-card FFTs. In: Proceedings of the International Conference on Supercomputing, Series ICS \u201911. ACM, pp 255\u2013264","DOI":"10.1145\/1995896.1995937"},{"issue":"11","key":"2141_CR19","doi-asserted-by":"crossref","first-page":"5634","DOI":"10.1109\/TAP.2014.2350536","volume":"62","author":"X Mu","year":"2014","unstructured":"Mu X, Zhou H-X, Chen K, Hong W (2014) Higher order method of moments with a parallel out-of-core LU solver on GPU\/CPU platform. IEEE Trans Antennas Propag 62(11):5634\u20135646","journal-title":"IEEE Trans Antennas Propag"},{"key":"2141_CR20","doi-asserted-by":"crossref","unstructured":"Zhong Z, Rychkov V, Lastovetsky A (2012) Data partitioning on heterogeneous multicore and multi-GPU systems using functional performance models of data-parallel applications. In: 2012 IEEE International Conference on Cluster Computing (Cluster 2012), 24\u201328 Sept 2012, pp 191\u2013199","DOI":"10.1109\/CLUSTER.2012.34"},{"key":"2141_CR21","unstructured":"Zhong Z (2014) Optimization of data-parallel scientific applications on highly heterogeneous modern HPC platforms. Ph.D. dissertation, University College Dublin"},{"issue":"02","key":"2141_CR22","doi-asserted-by":"crossref","first-page":"1650007","DOI":"10.1142\/S0129626416500079","volume":"26","author":"J Wu","year":"2016","unstructured":"Wu J, Jaja J (2016) Achieving native GPU performance for out-of-card large dense matrix multiplication. Parallel Process Lett 26(02):1650007","journal-title":"Parallel Process Lett"},{"key":"2141_CR23","unstructured":"Edgar R (2009) SciGPU-GEMM. \n                        https:\/\/github.com\/YaohuiZeng\/scigpugemm"},{"key":"2141_CR24","unstructured":"Martin D (2010) High performance computing linpack benchmark for CUDA. \n                        https:\/\/github.com\/avidday\/hpl-cuda"},{"key":"2141_CR25","unstructured":"NVIDIA (2017) CUDA toolkit documentation. \n                        http:\/\/docs.nvidia.com\/cuda\/cublas\/index.html#axzz4kRVc2o6B"},{"key":"2141_CR26","unstructured":"Khaleghzadeh H, Zhong Z, Reddy R, Lastovetsky A (2017) ZZGemmOOC: multi-GPU out-of-core routines for dense matrix multiplization. \n                        https:\/\/git.ucd.ie\/hcl\/zzgemmooc.git"},{"key":"2141_CR27","unstructured":"Khaleghzadeh H, Zhong Z, Reddy R, Lastovetsky A (2017) XeonPhiOOC: out-of-core package for out-of-core DGEMM on Xeon Phi. \n                        https:\/\/git.ucd.ie\/manumachu\/xeonphiooc.git"},{"key":"2141_CR28","unstructured":"Khaleghzadeh H, Zhong Z, Reddy R, Lastovetsky A (2017) FPGAOOC: out-of-core package for out-of-core DGEMM on FPGA. \n                        https:\/\/git.ucd.ie\/hcl\/fpgagemm.git"},{"key":"2141_CR29","unstructured":"Intel MKL BLAS. \n                        https:\/\/software.intel.com\/en-us\/mkl"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11227-017-2141-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-017-2141-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-017-2141-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2018,1,30]],"date-time":"2018-01-30T10:01:16Z","timestamp":1517306476000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11227-017-2141-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,9,13]]},"references-count":29,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2018,2]]}},"alternative-id":["2141"],"URL":"https:\/\/doi.org\/10.1007\/s11227-017-2141-4","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"type":"print","value":"0920-8542"},{"type":"electronic","value":"1573-0484"}],"subject":[],"published":{"date-parts":[[2017,9,13]]}}}