{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T23:14:45Z","timestamp":1776122085048,"version":"3.50.1"},"reference-count":28,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2020,9,30]],"date-time":"2020-09-30T00:00:00Z","timestamp":1601424000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001843","name":"Science and Engineering Research Board","doi-asserted-by":"crossref","award":["EMR\/2016\/008015"],"award-info":[{"award-number":["EMR\/2016\/008015"]}],"id":[{"id":"10.13039\/501100001843","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Program. Lang. Syst."],"published-print":{"date-parts":[[2020,9,30]]},"abstract":"<jats:p>Effective models for fusion of loop nests continue to remain a challenge in both general-purpose and domain-specific language (DSL) compilers. The difficulty often arises from the combinatorial explosion of grouping choices and their interaction with parallelism and locality. This article presents a new fusion algorithm for high-performance domain-specific compilers for image processing pipelines. The fusion algorithm is driven by dynamic programming and explores spaces of fusion possibilities not covered by previous approaches, and it is also driven by a cost function more concrete and precise in capturing optimization criteria than prior approaches. The fusion model is particularly tailored to the transformation and optimization sequence applied by PolyMage and Halide, two recent DSLs for image processing pipelines. Our model-driven technique when implemented in PolyMage provides significant improvements (up to 4.32\u00d7) over PolyMage\u2019s approach (which uses auto-tuning to aid its model) and over Halide\u2019s automatic approach (by up to 2.46\u00d7) on two state-of-the-art shared-memory multicore architectures.<\/jats:p>","DOI":"10.1145\/3404846","type":"journal-article","created":{"date-parts":[[2020,11,8]],"date-time":"2020-11-08T15:26:06Z","timestamp":1604849166000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["An Effective Fusion and Tile Size Model for PolyMage"],"prefix":"10.1145","volume":"42","author":[{"given":"Abhinav","family":"Jangda","sequence":"first","affiliation":[{"name":"Indian Institute of Science, Bengaluru, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Uday","family":"Bondhugula","sequence":"additional","affiliation":[{"name":"Indian Institute of Science, Bengaluru, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,11,8]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/HiPC.2013.6799131"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854317"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3168832"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178372.3179529"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/645670.665368"},{"key":"e_1_2_1_6_1","unstructured":"Google Inc. 2017. XLA (Accelerated Linear Algebra) for TensorFlow. Retrieved from https:\/\/www.tensorflow.org\/performance\/xla\/."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178487.3178507"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1012241830762"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the International Workshop on Languages and Compilers for Parallel Computing. 301--320","author":"Kennedy Ken","unstructured":"Ken Kennedy and Kathryn S. McKinley. 1993. Maximizing loop parallelism and improving data locality via loop fusion and distribution. In Proceedings of the International Workshop on Languages and Compilers for Parallel Computing. 301--320."},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the ACM SIGPLAN Conference on Programming Languages Design and Implementation (PLDI\u201907)","author":"Krishnamoorthy Sriram","unstructured":"Sriram Krishnamoorthy, Muthu Baskaran, Uday Bondhugula, J. Ramanujam, A. Rountev, and P. Sadayappan. 2007. Effective automatic parallelization of stencil computations. In Proceedings of the ACM SIGPLAN Conference on Programming Languages Design and Implementation (PLDI\u201907)."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/258492.258520"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2541228.2555292"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925952"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2694344.2694364"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.70"},{"key":"e_1_2_1_16_1","unstructured":"PolyMage project Apache 2.0 license 2017. PolyMage. Retrieved from https:\/\/bitbucket.org\/udayb\/polymage."},{"key":"e_1_2_1_17_1","unstructured":"PolyMagePage 2015. PolyMage: A DSL and compiler for automatic optimization of image processing pipelines. Retrieved from http:\/\/mcl.csa.iisc.ernet.in\/polymage.html."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1183401.1183437"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185528"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462176"},{"key":"e_1_2_1_21_1","volume-title":"Giles","author":"Reguly Istv\u00e1n Z.","year":"2017","unstructured":"Istv\u00e1n Z. Reguly, Gihan R. Mudalige, and Mike B. Giles. 2017. Loop tiling in large-scale stencil codes at run-time with OPS. CoRR abs\/1704.00693 (2017)."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/277830.277857"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-28652-0_6"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126968"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the ACM SIGPLAN Symposium on Programming Languages Design and Implementation. 30--44","author":"Wolf M.","unstructured":"M. Wolf and Monica S. Lam. 1991. A data locality optimizing algorithm. In Proceedings of the ACM SIGPLAN Symposium on Programming Languages Design and Implementation. 30--44."},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the 12th Workshop on Languages and Compilers for Parallel Computing. Springer-Verlag, 477--480","author":"Wonnacott David","year":"1999","unstructured":"David Wonnacott. 1999. Time skewing for parallel computers. In Proceedings of the 12th Workshop on Languages and Compilers for Parallel Computing. Springer-Verlag, 477--480."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342004038956"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2259016.2259044"}],"container-title":["ACM Transactions on Programming Languages and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3404846","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3404846","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:17:44Z","timestamp":1750191464000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3404846"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,30]]},"references-count":28,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2020,9,30]]}},"alternative-id":["10.1145\/3404846"],"URL":"https:\/\/doi.org\/10.1145\/3404846","relation":{},"ISSN":["0164-0925","1558-4593"],"issn-type":[{"value":"0164-0925","type":"print"},{"value":"1558-4593","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,9,30]]},"assertion":[{"value":"2018-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-11-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}