{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T10:58:52Z","timestamp":1777546732920,"version":"3.51.4"},"reference-count":15,"publisher":"International Academy Publishing (IAP)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["JSW"],"DOI":"10.4304\/jsw.7.12.2695-2702","type":"journal-article","created":{"date-parts":[[2012,12,28]],"date-time":"2012-12-28T21:06:02Z","timestamp":1356728762000},"source":"Crossref","is-referenced-by-count":4,"title":["An Improved Implementation of Preconditioned Conjugate Gradient Method on GPU"],"prefix":"10.17706","volume":"7","author":[{"given":"Yechen","family":"Gui","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guijuan","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"7163","published-online":{"date-parts":[[2012,12,1]]},"reference":[{"key":"ref1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2005.42"},{"key":"ref2","doi-asserted-by":"publisher","DOI":"10.1137\/060675290"},{"key":"ref3","first-page":"11","article-title":"Implementing the conjugate gradient algorithm on multi-core systems","volume-title":"Proceedings of the International Symposium on System-on-Chip","author":"Wiggers","year":"2007","unstructured":"[7] W.A.Wiggers, V.Bakker, A.B.J.Kokkeler and G.J.M.Smit, \"Implementing the conjugate gradient algorithm on multi-core systems,\" In J.Nurmi,J.Takala and O.Vainio, editors, Proceedings of the International Symposium on System-on-Chip, Tampere, pages 11-14, Piscataway, NJ,November 2007,IEEE,ISBN 1-4244-1367-2"},{"key":"ref4","first-page":"438","volume-title":"International Conference","author":"Maringanti","year":"2009","unstructured":"[8] Maringanti.A.;Athavale.V.; Patkar.S.B.; \"Acceleration of Conjugate Gradient Method for cirruit simulation using CUDA\" In High Performance Computing (HiPC), 2009 International Conference, Kochi, pp438-444"},{"key":"ref5","doi-asserted-by":"publisher","DOI":"10.1145\/1362622.1362674"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"[1] J.D.Hall,N.A. Carr and J.C.Hart,\"Cache and Bandwidth aware matrix multiplication on the GPU,\",2003.UIUC Technical Report UIUCDCSR-2003-2328(2003)","DOI":"10.1145\/1362622.1362674"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"[3] Kendell A. Atkinson (1988), An introduction to numerical analysis (2nd ed.), Section 8.9, John Wiley and Sons.","DOI":"10.1145\/1362622.1362674"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"[5] Nvidia CUDA. Website,2009, http:\/\/www.nvidia.com\/cuda","DOI":"10.1145\/1362622.1362674"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"[6] J. Bolz, I. Farmer, E. Grinspun, and P. Schr\u00a8ooder, \"Sparse matrix solvers on the GPU: Conjugate gradients and multigrid,\" in SIGGRAPH '03:ACM SIGGRAPH 2003 Papers, 2003, pp. 917\u2013924.","DOI":"10.1145\/1362622.1362674"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"[9] A.Asgasri, J.E.Tate. Implementing the Chebyshev Polynomial Preconditioner for the iterative solutions of linear systems on massively parallel graphics processors http:\/\/www.ele.utoronto.ca\/zeb\/publications\/,2009","DOI":"10.1145\/1362622.1362674"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"[10] L. Buatois, G. Caumon, and B. Levy, \"Concurrent number cruncher: a GPU implementation of a general sparse linear solver,\" Int. J. Parallel Emerg. Distrib. Syst., 24(3):205\u2013223,2009. ISSN 1744-5760","DOI":"10.1145\/1362622.1362674"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"[11] Marco Ament, Gunter Knittel, Daniel Weiskopf, Wolfgang Strasser, \"A Parallel Preconditioned Conjugate Gradient Solver for the Poisson Problem on a Multi-GPU Platform,\" pdp, pp.583-592, 2010 18th Euromicro Conference on Parallel, Distributed and Network-based Processing, 2010","DOI":"10.1145\/1362622.1362674"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"[13] Nathan Bell \"Implementing Sparse Matrix-Vector Multiplication on Throughput-Oriented Processors\" in \"Proc. Supercomputing '09\", November 2009","DOI":"10.1145\/1362622.1362674"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"[14] Nvidia, \"CUDA Toolkit 4.0 CUBLAS Library\" NVIDIA Corporation, Santa Clara, April, 2011","DOI":"10.1145\/1362622.1362674"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"[15] Nvidia,\"CUDA Programming Guide 4.0\" NVIDIA Corporation, Santa Clara, April, 2011","DOI":"10.1145\/1362622.1362674"}],"container-title":["Journal of Software"],"original-title":[],"deposited":{"date-parts":[[2017,6,21]],"date-time":"2017-06-21T04:58:16Z","timestamp":1498021096000},"score":1,"resource":{"primary":{"URL":"http:\/\/ojs.academypublisher.com\/index.php\/jsw\/article\/view\/7119"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,12,1]]},"references-count":15,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2012,12,1]]}},"URL":"https:\/\/doi.org\/10.4304\/jsw.7.12.2695-2702","relation":{},"ISSN":["1796-217X"],"issn-type":[{"value":"1796-217X","type":"print"}],"subject":[],"published":{"date-parts":[[2012,12,1]]}}}