{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T22:46:32Z","timestamp":1777675592747,"version":"3.51.4"},"reference-count":28,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2020,1,13]],"date-time":"2020-01-13T00:00:00Z","timestamp":1578873600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61272062"],"award-info":[{"award-number":["61272062"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2020,5]]},"abstract":"<jats:p>With the new architecture and new programming paradigms such as task-based scheduling emerging in the parallel high performance computing area, it is of great importance to utilize these features to tune the monolithic computing codes. In this article, the classical conjugate gradient algorithms targeting at sparse linear system Ax = b in Krylov subspace are pipelining to execute interdependent tasks on Parallel Runtime Scheduling and Execution Controller (PaRSEC) runtime. Firstly, the sparse matrix A is split in rows to unfold more coarse-grained parallelism. Secondly, the partitioned sub-vectors are not assembled into one full vector in RAM to run sparse matrix\u2013vector product (SpMV) operations for eliminating the communication overhead. Moreover, in the SpMV computation, if all elements of one column in the split sub-matrix are zeros, the corresponding product operations of these elements may be removed by reorganizing sub-vectors. Finally, the latency of migrating sub-vector is partially overlapped by the duration of performing SpMV operations through the further splitting in columns of sparse matrix on GPUs. In experiments, a series of tests demonstrate that optimal speedup and higher pipelining efficiency has been achieved for the pipelined task scheduling on PaRSEC runtime. Fusing SpMV concurrency and dot product pipelining can achieve higher speedup and efficiency.<\/jats:p>","DOI":"10.1177\/1094342019899997","type":"journal-article","created":{"date-parts":[[2020,1,13]],"date-time":"2020-01-13T08:22:01Z","timestamp":1578903721000},"page":"306-315","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":1,"title":["Iteratively solving sparse linear system based on PaRSEC task scheduling"],"prefix":"10.1177","volume":"34","author":[{"given":"Tieqiang","family":"Mo","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Renfa","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information Science and Electronic Engineering, Hunan University, Changsha, Hunan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2020,1,13]]},"reference":[{"key":"bibr1-1094342019899997","first-page":"473","volume-title":"GPU Computing Gems","author":"Agullo E","year":"2010"},{"key":"bibr2-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.90"},{"key":"bibr3-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1137\/130915662"},{"key":"bibr4-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1145\/2898348"},{"key":"bibr5-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.1631"},{"key":"bibr6-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611971538"},{"key":"bibr7-1094342019899997","volume-title":"Distributed-Memory Task Execution and Dependence Tracking Within DAGuE and the DPLASMA Project","author":"Bosilca G","year":"2010"},{"key":"bibr8-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.299"},{"key":"bibr9-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2011.10.003"},{"key":"bibr10-1094342019899997","unstructured":"BSC Programming Models Group (2015) The Nanos++ parallel runtime. Available at https:\/\/pm.bsc.es\/nanox (accessed 7 February 2018)."},{"key":"bibr11-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1137\/110846427"},{"key":"bibr12-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1137\/120893057"},{"key":"bibr13-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-01970-8_90"},{"key":"bibr14-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1002\/nla.643"},{"key":"bibr15-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2017.04.005"},{"key":"bibr16-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898718881"},{"issue":"1","key":"bibr17-1094342019899997","first-page":"1","volume":"38","author":"Davis TA","year":"2011","journal-title":"ACM Transactions on Mathematical Software (TOMS)"},{"key":"bibr18-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2017.04.003"},{"key":"bibr19-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-14313-2_29"},{"key":"bibr20-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2013.06.001"},{"key":"bibr21-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1145\/3148226.3148233"},{"key":"bibr22-1094342019899997","volume-title":"Mini-symposium on \u201ctask-based scientific computing applications\u201d at SIAM CSE\u201915 conference","author":"Lacoste X","year":"2015"},{"key":"bibr23-1094342019899997","volume-title":"On the design of sparse hybrid linear solvers for modern parallel architectures","author":"Nakov S","year":"2015"},{"key":"bibr24-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898718003"},{"key":"bibr25-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2014.2311804"},{"key":"bibr26-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2002.1016651"},{"key":"bibr27-1094342019899997","volume-title":"Dynamic task execution on shared and distributed memory architectures","author":"YarKhan A","year":"2012"},{"key":"bibr28-1094342019899997","doi-asserted-by":"publisher","DOI":"10.1145\/3079079.3079091"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342019899997","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342019899997","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342019899997","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:15:56Z","timestamp":1777450556000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342019899997"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,1,13]]},"references-count":28,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2020,5]]}},"alternative-id":["10.1177\/1094342019899997"],"URL":"https:\/\/doi.org\/10.1177\/1094342019899997","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,1,13]]}}}