{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,8]],"date-time":"2025-11-08T17:35:48Z","timestamp":1762623348771,"version":"3.38.0"},"reference-count":31,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2011,9,1]],"date-time":"2011-09-01T00:00:00Z","timestamp":1314835200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Comput. Sci. Technol."],"published-print":{"date-parts":[[2011,9]]},"DOI":"10.1007\/s11390-011-0184-1","type":"journal-article","created":{"date-parts":[[2011,9,23]],"date-time":"2011-09-23T21:20:46Z","timestamp":1316812846000},"page":"854-865","source":"Crossref","is-referenced-by-count":29,"title":["Optimizing Linpack Benchmark on GPU-Accelerated Petascale Supercomputer"],"prefix":"10.1007","volume":"26","author":[{"given":"Feng","family":"Wang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Can-Qun","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yun-Fei","family":"Du","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Juan","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hui-Zhan","family":"Yi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei-Xia","family":"Xu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2011,9,23]]},"reference":[{"issue":"3","key":"184_CR1","doi-asserted-by":"crossref","first-page":"523","DOI":"10.1006\/jpdc.1994.1108","volume":"22","author":"JJ Dongarra","year":"1994","unstructured":"Dongarra J\u00a0J, van de Geijn R A, Walker D W. Scalability issues affecting the design of a dense linear algebra library. J. Parallel Distrib. Comput., 1994, 22(3): 523\u2013537.","journal-title":"J. Parallel Distrib. Comput."},{"key":"184_CR2","unstructured":"http:\/\/www.top500.org , Nov. 10, 2010."},{"key":"184_CR3","doi-asserted-by":"crossref","unstructured":"Villarreal J, Najjar W. Compiled hardware acceleration of molecular dynamics code. In Proc. International Conference on Field Programmable Logic and Applications (FPL 2008), Heidelberg, Germany, Sept. 8\u201310, 2008, pp.667-670.","DOI":"10.1109\/FPL.2008.4630035"},{"key":"184_CR4","unstructured":"NVIDIA. Fermi compute architecture whitepaper, 2009."},{"key":"184_CR5","unstructured":"AMD. AMD stream computing user guide v 1.4.0, Feb. 2009."},{"key":"184_CR6","unstructured":"NVIDIA. CUDA programming guide, June 2007."},{"key":"184_CR7","unstructured":"Munshi A. Opencl parallel computing on the GPU and CPU. In Proc. ACM SIGGRAPH 2008, Los Angeles, USA, Aug. 11\u201315, 2008."},{"issue":"5","key":"184_CR8","doi-asserted-by":"crossref","first-page":"913","DOI":"10.1007\/s11390-009-9266-8","volume":"24","author":"G Falcao","year":"2009","unstructured":"Falcao G, Yamagiwa S, Silva V, Sousa L. Parallel LDPC decoding on GPUs using a stream-based computing approach. Journal of Computer Science and Technology, 2009, 24(5): 913\u2013924.","journal-title":"Journal of Computer Science and Technology"},{"key":"184_CR9","doi-asserted-by":"crossref","unstructured":"Roberts E, Stone J E, Sepulveda L, Mei W, Hwu W, Luthey-Schulten Z. Long time-scale simulations of in vivo diffusion using GPU hardware. In Proc. the 2009 IEEE International Symposium on Parallel&Distributed Processing (IPDPS 2009), Rome, Italy, May 23\u201329, 2009, pp.1-8.","DOI":"10.1109\/IPDPS.2009.5160930"},{"key":"184_CR10","doi-asserted-by":"crossref","unstructured":"Meng J, Skadron K. Performance modeling and automatic ghost zone optimization for iterative stencil loops on GPUs. In Proc. the 23\u00a0rd International Conference on Supercomputing (ICS 2009), Yorktown Heights, USA, Jun. 8\u201312, 2009, pp.256-265.","DOI":"10.1145\/1542275.1542313"},{"key":"184_CR11","doi-asserted-by":"crossref","unstructured":"Di P, Wan Q, Zhang X, Wu H, Xue J. Toward harnessing DOACROSS parallelism for multi-GPGPUs. In Proc. the 39th International Conference on Parallel Processing, San Diego, USA, Sept. 13\u201316, 2010, pp.40-50.","DOI":"10.1109\/ICPP.2010.13"},{"key":"184_CR12","unstructured":"Fan Z, Qiu F, Kaufman A, Yoakum-Stover S. GPU cluster for high performance computing. In Proc. the 2004 ACM\/IEEE Conference on Supercomputing (SC 2004), Pittsburgh, USA, Nov. 6\u201312, 2004, p.47."},{"key":"184_CR13","unstructured":"Sun J C, Yuan G X, Zhang L B, Zhang Y Q. 2009 China top100 list of high performance computer. http:\/\/124.16.137.70\/2009-China-HPC-TOP100-20091101-eng.htm , Nov. 2009."},{"key":"184_CR14","unstructured":"Petitet A, Whaley R C, Dongarra J J, Cleary A. HPL \u2014 A portable implementation of the high-performance linpack benchmark for distributed memory computers. http:\/\/www.netlib.org\/benchmark\/hpl\/ , 2006."},{"key":"184_CR15","doi-asserted-by":"crossref","unstructured":"Luk C K, Hong S, Kim H. Qilin: Exploiting parallelism on heterogeneous multiprocessors with adaptive mapping. In Proc. the 42nd Annual IEEE\/ACM International Symposium on Microarchitecture (Micro-42), New York, USA, Dec. 12\u201316, 2009, pp.45-55.","DOI":"10.1145\/1669112.1669121"},{"issue":"9","key":"184_CR16","doi-asserted-by":"crossref","first-page":"803","DOI":"10.1002\/cpe.728","volume":"15","author":"JJ Dongarra","year":"2003","unstructured":"Dongarra J\u00a0J, Luszczek P, Petitet A. The linpack benchmark: Past, present and future. Concurrency and Computation: Practice and Experience, 2003, 15(9): 803\u2013820.","journal-title":"Concurrency and Computation: Practice and Experience"},{"issue":"1","key":"184_CR17","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/77626.79170","volume":"16","author":"JJ Dongarra","year":"1990","unstructured":"Dongarra J\u00a0J, Du Croz J, Hammarling S, Duff I S. A set of level 3 basic linear algebra subprograms. ACM Trans. Math. Softw., 1990, 16(1): 1\u201317.","journal-title":"ACM Trans. Math. Softw."},{"key":"184_CR18","doi-asserted-by":"crossref","unstructured":"Kistler M, Gunnels J, Brokenshire D, Benton B. Petascale computing with accelerators. In Proc. the 14th ACM SIG-PLAN Symposium on Principles and Practice of Parallel Programming (PPoPP 2009), Raleigh, USA, Feb. 14\u201318, 2009, pp.241-250.","DOI":"10.1145\/1504176.1504212"},{"key":"184_CR19","first-page":"157","volume":"11","author":"H Baliga","year":"2008","unstructured":"Baliga H, Cooray N, Gamsaragan E, Smith P, Yoon K, Abel J, Valles A. Original 45\u00a0nm Intels Core2 processor performance. Intel Technology Journal, 2008, 11: 157\u2013168.","journal-title":"Intel Technology Journal"},{"key":"184_CR20","unstructured":"AMD. AMD core math library for graphic processors release notes for version 1.0, 2009."},{"issue":"5","key":"184_CR21","doi-asserted-by":"crossref","first-page":"575","DOI":"10.1147\/rd.395.0575","volume":"39","author":"R Agarwal","year":"1995","unstructured":"Agarwal R, Balle S M, Gustavson F\u00a0G, Joshi M, Palkar P. A three-dimensional approach to parallel matrix multiplication. IBM Journal of Research and Development, 1995, 39(5): 575\u2013582.","journal-title":"IBM Journal of Research and Development"},{"key":"184_CR22","doi-asserted-by":"crossref","unstructured":"Ryoo S, Rodrigues C I, Baghsorkhi S S, Stone S S, Kirk D B, Hwu W\u00a0M\u2009W. Optimization principles and application per- formance evaluation of a multithreaded GPU using CUDA. In Proc. the 13th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP 2008), Salt Lake City, Feb. 20\u201323, 2008, pp.73-82.","DOI":"10.1145\/1345206.1345220"},{"key":"184_CR23","doi-asserted-by":"crossref","unstructured":"Quintana-Ort\u00ed G, Igual F D, Quintana-Ort\u00ed E S, van de Geijn R A. Solving dense linear systems on platforms with multiple hardware accelerators. In Proc. the 14th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP 2009), Raleigh, USA, Feb. 14\u201318, 2009, pp.121-130.","DOI":"10.1145\/1594835.1504196"},{"issue":"2","key":"184_CR24","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1145\/1353535.1346318","volume":"42","author":"MD Linderman","year":"2008","unstructured":"Linderman M D, Collins J D, Wang H, Meng T H. Merge: A programming model for heterogeneous multi-core systems. SIGOPS Oper. Syst. Rev., 2008, 42(2): 287\u2013296.","journal-title":"SIGOPS Oper. Syst. Rev."},{"key":"184_CR25","doi-asserted-by":"crossref","unstructured":"Fatica M. Accelerating linpack with CUDA on heterogenous clusters. In Proc. 2nd Workshop on General Purpose Processing on Graphics Processing Units (GPGPU-2), Washington DC, USA, 2009, pp.46-51.","DOI":"10.1145\/1513895.1513901"},{"issue":"5","key":"184_CR26","doi-asserted-by":"crossref","first-page":"503","DOI":"10.1147\/rd.515.0503","volume":"51","author":"CR Johns","year":"2007","unstructured":"Johns C R, Brokenshire D A. Introduction to the cell broadband engine architecture. IBM J. Res. Dev., 2007, 51(5): 503\u2013519.","journal-title":"IBM J. Res. Dev."},{"key":"184_CR27","unstructured":"ATI Radeon rv770. http:\/\/en.wikipedia.org\/wiki\/Radeon_R700 ."},{"key":"184_CR28","doi-asserted-by":"crossref","unstructured":"Hamano T, Endo T, Matsuoka S. Power-aware dynamic task scheduling for heterogeneous accelerated clusters. In Proc. Int. Parallel and Distributed Processing Symposium, Rome, Italy, May 23\u201329, 2009, pp.1-8.","DOI":"10.1109\/IPDPS.2009.5160977"},{"key":"184_CR29","unstructured":"Clearspeed Technology Inc. http:\/\/www.clearspeed.com\/ ."},{"key":"184_CR30","unstructured":"NVIDIA. http:\/\/www.nvidia.com\/object\/product_tesla_s1070_us.html , Nov. 10, 2010."},{"key":"184_CR31","doi-asserted-by":"crossref","unstructured":"Endo T, Matsuoka S. Massive supercomputing coping with heterogeneity of modern accelerators. In Proc. the 2008 IEEE International Symposium on Parallel&Distributed Processing (IPDPS 2008), Miami, USA, Apr. 14\u201318, 2008, pp.1-10.","DOI":"10.1109\/IPDPS.2008.4536251"}],"container-title":["Journal of Computer Science and Technology"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11390-011-0184-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11390-011-0184-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11390-011-0184-1","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,3,11]],"date-time":"2025-03-11T23:01:17Z","timestamp":1741734077000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11390-011-0184-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,9]]},"references-count":31,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2011,9]]}},"alternative-id":["184"],"URL":"https:\/\/doi.org\/10.1007\/s11390-011-0184-1","relation":{},"ISSN":["1000-9000","1860-4749"],"issn-type":[{"type":"print","value":"1000-9000"},{"type":"electronic","value":"1860-4749"}],"subject":[],"published":{"date-parts":[[2011,9]]}}}