{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T23:00:27Z","timestamp":1777676427189,"version":"3.51.4"},"reference-count":39,"publisher":"SAGE Publications","issue":"2","license":[{"start":{"date-parts":[[2009,11,3]],"date-time":"2009-11-03T00:00:00Z","timestamp":1257206400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2010,5]]},"abstract":"<jats:p>Sparse matrix operations achieve only small fractions of peak CPU speeds because of the use of specialized, index-based matrix representations, which degrade cache utilization by imposing irregular memory accesses and increasing the number of overall accesses. Compounding the problem, the small number of floating-point operations in a single sparse iteration leads to low floating-point pipeline utilization. Operation stacking addresses these problems for large ensemble computations that solve multiple systems of linear equations with identical sparsity structure. By combining the data of multiple problems and solving them as one, operation stacking improves locality, reduces cache misses, and increases floating-point pipeline utilization. Operation stacking also requires less memory bandwidth because it involves fewer index array accesses. In this paper we present the Operation Stacking Framework (OSF), an object-oriented framework that provides runtime and code generation support for the development of stacked iterative solvers. OSF\u2019s runtime component provides an iteration engine that supports efficient ejection of converged problems from the stack. It separates the specific solver algorithm from the coding conventions and data representations that are necessary to implement stacking. Stacked solvers created with OSF can be used transparently without requiring significant changes to existing applications. Our results show that stacking can provide speedups up to 1.94\u00d7 with an average of 1.46\u00d7, even in scenarios in which the number of iterations required to converge varies widely within a stack of problems. Our evaluation shows that these improvements correlate with better cache utilization, improved floating-point utilization, and reduced memory accesses.<\/jats:p>","DOI":"10.1177\/1094342009347892","type":"journal-article","created":{"date-parts":[[2009,11,3]],"date-time":"2009-11-03T23:24:44Z","timestamp":1257290684000},"page":"194-212","source":"Crossref","is-referenced-by-count":0,"title":["Operation Stacking for Ensemble Computations With Variable Convergence"],"prefix":"10.1177","volume":"24","author":[{"given":"Mehmet","family":"Belgin","sequence":"first","affiliation":[{"name":"DEPARTMENT OF COMPUTER SCIENCE, VIRGINIA TECH, 2202 KRAFT DRIVE, BLACKSBURG, VA 24060, USA,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Godmar","family":"Back","sequence":"additional","affiliation":[{"name":"DEPARTMENT OF COMPUTER SCIENCE, VIRGINIA TECH, 2202 KRAFT DRIVE, BLACKSBURG, VA 24060, USA,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Calvin J.","family":"Ribbens","sequence":"additional","affiliation":[{"name":"DEPARTMENT OF COMPUTER SCIENCE, VIRGINIA TECH, 2202 KRAFT DRIVE, BLACKSBURG, VA 24060, USA,"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2009,11,3]]},"reference":[{"key":"atypb1","unstructured":"Allen, F.E. and Cocke, J. ( 1971). A catalogue of optimizing transformations. In Design and Optimization of Compilers. Englewood Cliffs, NJ: Prentice-Hall, pp. 1-30."},{"key":"atypb2","volume-title":"IEEE standard for information technology: Portable Operating System Interface (POSIX)","author":"American National Standards Institute","year":"1994"},{"key":"atypb3","volume-title":"Proceedings, Supercomputing \u201990: 12-16 November 1990, New York Hilton at Rockefeller Center","author":"Anderson, E."},{"key":"atypb4","volume-title":"Proceedings of SC99","author":"Anderson, W.K."},{"key":"atypb5","volume-title":"Structured Matrices in Operator Theory, Numerical Analysis, Control, Signal and Image Processing (Contemporary Mathematics, Vol. 280)","author":"Antoulas, A."},{"key":"atypb6","doi-asserted-by":"publisher","DOI":"10.1137\/040608088"},{"key":"atypb7","volume-title":"PETSc Web page","author":"Balay, S.","year":"2001"},{"key":"atypb8","volume-title":"Templates for the solution of linear systems","author":"Barrett, R.","year":"1994"},{"key":"atypb9","volume-title":"Proceedings of ICCSE2005, The International Conference on Computational Science and Engineering","author":"Belgin, M."},{"key":"atypb10","doi-asserted-by":"publisher","DOI":"10.1145\/1274971.1274986"},{"key":"atypb11","doi-asserted-by":"publisher","DOI":"10.1177\/109434200001400303"},{"key":"atypb12","volume-title":"The University of Florida sparse matrix collection","author":"Davis, T.","year":"2009"},{"key":"atypb13","volume-title":"Performance optimization of sparse kernels in TOPS","author":"Demmel, J.","year":"2006"},{"key":"atypb14","doi-asserted-by":"publisher","DOI":"10.1145\/77626.79170"},{"key":"atypb15","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898719611"},{"key":"atypb16","volume-title":"Proceedings of the Second Object Oriented Numerics Conference","author":"Dongarra, J."},{"key":"atypb17","doi-asserted-by":"publisher","DOI":"10.1126\/science.1115255"},{"key":"atypb18","volume-title":"Technical Report CS-TR-02-55, The University of Texas at Austin, Department of Computer Sciences","author":"Goto, K."},{"key":"atypb19","doi-asserted-by":"publisher","DOI":"10.1137\/060666123"},{"key":"atypb20","volume-title":"Computer Architecture, Fourth Edition: A Quantitative Approach","author":"Hennessy, J.L.","year":"2006"},{"key":"atypb21","volume-title":"ICCS \u201901: Proceedings of the International Conference on Computational Sciences-Part I","author":"Im, E.-J."},{"key":"atypb22","doi-asserted-by":"publisher","DOI":"10.1177\/1094342004041296"},{"key":"atypb23","volume-title":"PPSC","author":"Joshi, M.V."},{"key":"atypb24","volume-title":"CF \u201908: Proceedings of the 2008 Conference on Computing Frontiers","author":"Kourtis, K."},{"key":"atypb25","doi-asserted-by":"publisher","DOI":"10.1177\/1094342004038951"},{"key":"atypb26","doi-asserted-by":"publisher","DOI":"10.1002\/jcc.20081"},{"key":"atypb27","doi-asserted-by":"publisher","DOI":"10.1016\/0024-3795(80)90247-5"},{"key":"atypb28","volume-title":"Proceedings of Supercomputing\u201999 (CD-ROM)","author":"Pinar, A."},{"key":"atypb29","volume-title":"Technical Report, National Institute of Standards and Technology","author":"Remington, K.A."},{"key":"atypb30","unstructured":"Saad, Y. ( 1994). SPARSKIT: A basic tool kit for sparse matrix computations, Version 2. Technical Report, Computer Science Department, University of Minnesota, Minneapolis, MN 55455."},{"key":"atypb31","volume-title":"Iterative Methods for Sparse Linear Systems","author":"Saad, Y.","year":"1996"},{"key":"atypb32","doi-asserted-by":"publisher","DOI":"10.1016\/0024-3795(95)00093-3"},{"key":"atypb33","doi-asserted-by":"crossref","unstructured":"Simoncini, V. and Gallopoulos, E. ( 1996b). A hybrid block GMRES method for nonsymmetric systems with multiple right-hand sides, pp. 457-469.","DOI":"10.1016\/0377-0427(95)00198-0"},{"issue":"6","key":"atypb34","first-page":"711","volume":"41","author":"Toledo, S.","year":"1997","journal-title":"J. Res. Dev"},{"key":"atypb35","volume-title":"Official Aztec user\u2019s guide version 2.1. Technical Report SAND99-8801J","author":"Tuminaro, R.S.","year":"1999"},{"key":"atypb36","volume-title":"Proceedings of SciDAC 2005 (Journal of Physics: Confer-OPERATION STACKING 211 ence Series)","author":"Vuduc, R."},{"key":"atypb37","volume-title":"High Performance Computing and Communcations (Lecture Notes in Computer Science, Vol. 3726)","author":"Vuduc, R.W."},{"key":"atypb38","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-8191(00)00087-9"},{"key":"atypb39","volume-title":"ICS \u201906: Proceedings of the 20th Annual International Conference on Supercomputing","author":"Willcock, J."}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342009347892","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342009347892","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:18:52Z","timestamp":1777450732000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342009347892"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,11,3]]},"references-count":39,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2010,5]]}},"alternative-id":["10.1177\/1094342009347892"],"URL":"https:\/\/doi.org\/10.1177\/1094342009347892","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,11,3]]}}}