{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,30]],"date-time":"2025-10-30T06:59:04Z","timestamp":1761807544727,"version":"3.41.0"},"reference-count":24,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2010,2,1]],"date-time":"2010-02-01T00:00:00Z","timestamp":1264982400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000145","name":"Division of Information and Intelligent Systems","doi-asserted-by":"publisher","award":["CCF-0309461IIS-0513669"],"award-info":[{"award-number":["CCF-0309461IIS-0513669"]}],"id":[{"id":"10.13039\/100000145","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["CCF-0309461IIS-0513669"],"award-info":[{"award-number":["CCF-0309461IIS-0513669"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["60728206"],"award-info":[{"award-number":["60728206"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002920","name":"Research Grants Council, University Grants Committee, Hong Kong","doi-asserted-by":"publisher","award":["GRF CityU 123609GRF PolyU 5260\/07EHK CityU 7002473HK PolyU 1-ZV5S"],"award-info":[{"award-number":["GRF CityU 123609GRF PolyU 5260\/07EHK CityU 7002473HK PolyU 1-ZV5S"]}],"id":[{"id":"10.13039\/501100002920","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2010,2]]},"abstract":"<jats:p>\n            The widening gap between processor and memory performance is the main bottleneck for modern computer systems to achieve high processor utilization. To hide memory latency, a variety of techniques have been proposed\u2014from intermediate fast memories (caches) to various prefetching and memory management techniques. In this article, we propose a new loop scheduling with memory management technique,\n            <jats:italic>Iterational Retiming with Partitioning<\/jats:italic>\n            (IRP), that can completely hide memory latencies for applications with multidimensional loops on architectures like CELL processor. In IRP, the iteration space is first partitioned carefully. Then a two-part schedule, consisting of processor and memory parts, is produced such that the execution time of the memory part never exceeds the execution time of the processor part. These two parts are executed simultaneously and complete memory latency hiding is reached. In this article, we prove that such optimal two-part schedule can always be achieved given the right partition size and shape. Experiments on DSP benchmarks show that IRP consistently produces optimal solutions as well as significant improvement over previous techniques.\n          <\/jats:p>","DOI":"10.1145\/1698772.1698780","type":"journal-article","created":{"date-parts":[[2010,3,2]],"date-time":"2010-03-02T19:20:32Z","timestamp":1267557632000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["Iterational retiming with partitioning"],"prefix":"10.1145","volume":"9","author":[{"given":"Chun Jason","family":"Xue","sequence":"first","affiliation":[{"name":"City University of Hong Kong, Kowloon, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingtong","family":"Hu","sequence":"additional","affiliation":[{"name":"University of Texas, Dallas, Texas"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zili","family":"Shao","sequence":"additional","affiliation":[{"name":"Hong Kong Polytechnic University, Kowloon, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Edwin","family":"Sha","sequence":"additional","affiliation":[{"name":"University of Texas, Dallas, Texas"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2010,3,5]]},"reference":[{"doi-asserted-by":"publisher","key":"e_1_2_1_1_1","DOI":"10.1109\/71.466632"},{"doi-asserted-by":"publisher","key":"e_1_2_1_2_1","DOI":"10.1145\/277830.277925"},{"doi-asserted-by":"publisher","key":"e_1_2_1_3_1","DOI":"10.1109\/43.594829"},{"doi-asserted-by":"publisher","key":"e_1_2_1_4_1","DOI":"10.1109\/71.862210"},{"volume-title":"Proceedings of the International Symposium on System Synthesis.","author":"Chen F.","unstructured":"Chen , F. and Sha , E. H . -M. 1999. Loop scheduling and partitions for hiding memory latencies . In Proceedings of the International Symposium on System Synthesis. Chen, F. and Sha, E. H.-M. 1999. Loop scheduling and partitions for hiding memory latencies. In Proceedings of the International Symposium on System Synthesis.","key":"e_1_2_1_5_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_6_1","DOI":"10.1145\/191995.192030"},{"doi-asserted-by":"publisher","key":"e_1_2_1_7_1","DOI":"10.1109\/71.395402"},{"doi-asserted-by":"publisher","key":"e_1_2_1_8_1","DOI":"10.1145\/774572.774655"},{"key":"e_1_2_1_9_1","article-title":"Introduction to the cell multiprocessor. IBM","author":"Kahle J. A.","year":"2005","unstructured":"Kahle , J. A. , Day , M. N. , Hofstee , H. P. , Johns , C. R. , Maeurer , T. R. , and Shippy , D. 2005 . Introduction to the cell multiprocessor. IBM J. Resear. Dev. 49. Kahle, J. A., Day, M. N., Hofstee, H. P., Johns, C. R., Maeurer, T. R., and Shippy, D. 2005. Introduction to the cell multiprocessor. IBM J. Resear. Dev. 49.","journal-title":"J. Resear. Dev. 49."},{"doi-asserted-by":"publisher","key":"e_1_2_1_10_1","DOI":"10.1007\/BF01759032"},{"doi-asserted-by":"publisher","key":"e_1_2_1_11_1","DOI":"10.5555\/645533.656505"},{"volume-title":"Proceedings of the 28th Annual ACM\/IEEE International Symposium on Microarchitecure. 243--248","author":"Ozawa T.","unstructured":"Ozawa , T. , Kimura , Y. , and Nishizaki , S . 1995. Cache miss heuristics and preloading techniques for general-purpose programs . In Proceedings of the 28th Annual ACM\/IEEE International Symposium on Microarchitecure. 243--248 . Ozawa, T., Kimura, Y., and Nishizaki, S. 1995. Cache miss heuristics and preloading techniques for general-purpose programs. In Proceedings of the 28th Annual ACM\/IEEE International Symposium on Microarchitecure. 243--248.","key":"e_1_2_1_12_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_13_1","DOI":"10.1109\/92.736145"},{"doi-asserted-by":"publisher","key":"e_1_2_1_14_1","DOI":"10.1145\/237090.237151"},{"volume-title":"Proceedings of the 29th Annual ACM\/IEEE International Symposium on Microarchitecure. 214--225","author":"Pinter S. S.","unstructured":"Pinter , S. S. and Yoaz , A . 1996. Tango: A hardware-based data prefetching for superscalar processors . In Proceedings of the 29th Annual ACM\/IEEE International Symposium on Microarchitecure. 214--225 . Pinter, S. S. and Yoaz, A. 1996. Tango: A hardware-based data prefetching for superscalar processors. In Proceedings of the 29th Annual ACM\/IEEE International Symposium on Microarchitecure. 214--225.","key":"e_1_2_1_15_1"},{"volume-title":"Proceedings of the International Conference on Parallel Processing. 298--305","author":"Sheppstedt J.","unstructured":"Sheppstedt , J. and Dubois , M . 1997. Hybrid compiler-hardware prefetching for multiprocessors using low-overhead cache miss traps . In Proceedings of the International Conference on Parallel Processing. 298--305 . Sheppstedt, J. and Dubois, M. 1997. Hybrid compiler-hardware prefetching for multiprocessors using low-overhead cache miss traps. In Proceedings of the International Conference on Parallel Processing. 298--305.","key":"e_1_2_1_16_1"},{"volume-title":"Proceedings of the International Conference on Parallel Processing. 306--313","author":"Tcheun M. K.","unstructured":"Tcheun , M. K. , Yoon , H. , and Maeng , S. R . 1997. An adaptive sequential prefetching scheme in shared-memory multiprocessors . In Proceedings of the International Conference on Parallel Processing. 306--313 . Tcheun, M. K., Yoon, H., and Maeng, S. R. 1997. An adaptive sequential prefetching scheme in shared-memory multiprocessors. In Proceedings of the International Conference on Parallel Processing. 306--313.","key":"e_1_2_1_17_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_18_1","DOI":"10.1109\/71.689444"},{"doi-asserted-by":"publisher","key":"e_1_2_1_19_1","DOI":"10.1145\/859618.859663"},{"key":"e_1_2_1_20_1","first-page":"926","article-title":"Partitioning and scheduling dsp applications with maximal memory access hiding","volume":"9","author":"Wang Z.","year":"2002","unstructured":"Wang , Z. , Sha , E.-M. , and Wang , Y. 2002 . Partitioning and scheduling dsp applications with maximal memory access hiding . Eurasip J. Appl. Sing. Process. 9 , 926 -- 935 . Wang, Z., Sha, E.-M., and Wang, Y. 2002. Partitioning and scheduling dsp applications with maximal memory access hiding. Eurasip J. Appl. Sing. Process. 9, 926--935.","journal-title":"Eurasip J. Appl. Sing. Process."},{"doi-asserted-by":"publisher","key":"e_1_2_1_21_1","DOI":"10.1145\/500001.500042"},{"doi-asserted-by":"publisher","key":"e_1_2_1_22_1","DOI":"10.1145\/113445.113449"},{"doi-asserted-by":"publisher","key":"e_1_2_1_23_1","DOI":"10.1145\/1084834.1084910"},{"doi-asserted-by":"publisher","key":"e_1_2_1_24_1","DOI":"10.1145\/192724.192740"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1698772.1698780","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1698772.1698780","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:22:58Z","timestamp":1750278178000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1698772.1698780"}},"subtitle":["Loop scheduling with complete memory latency hiding"],"short-title":[],"issued":{"date-parts":[[2010,2]]},"references-count":24,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2010,2]]}},"alternative-id":["10.1145\/1698772.1698780"],"URL":"https:\/\/doi.org\/10.1145\/1698772.1698780","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2010,2]]},"assertion":[{"value":"2006-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2007-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2010-03-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}