{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T16:03:53Z","timestamp":1780675433399,"version":"3.54.1"},"reference-count":20,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2011,3,1]],"date-time":"2011-03-01T00:00:00Z","timestamp":1298937600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"GSRC, SRC Contract","award":["2009-TJ-1879"],"award-info":[{"award-number":["2009-TJ-1879"]}]},{"DOI":"10.13039\/100000144","name":"Division of Computer and Network Systems","doi-asserted-by":"publisher","award":["CNS-0725354"],"award-info":[{"award-number":["CNS-0725354"]}],"id":[{"id":"10.13039\/100000144","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2011,3]]},"abstract":"<jats:p>Memory bottleneck has become a limiting factor in satisfying the explosive demands on performance and cost in modern embedded system design. Selected computation kernels for acceleration are usually captured by nest loops, which are optimized by state-of-the-art techniques like loop tiling and loop pipelining. However, memory bandwidth bottlenecks prevent designs from reaching optimal throughput with respect to available parallelism. In this paper we present an automatic memory partitioning technique which can efficiently improve throughput and reduce energy consumption of pipelined loop kernels for given throughput constraints and platform requirements. Also, our proposed algorithm can handle general array access beyond affine array references.<\/jats:p>\n          <jats:p>Our partition scheme consists of two steps. The first step considers cycle accurate scheduling information to meet the hard constraints on memory bandwidth requirements specifically for synchronized hardware designs. An ILP formulation is proposed to solve the memory partitioning and scheduling problem optimally for small designs, followed by a heuristic algorithm which is more scalable and equally effective for solving large scale problems. Experimental results show an average 6\u00d7 throughput improvement on a set of real-world designs with moderate area increase (about 45% on average), given that less resource sharing opportunities exist with higher throughput in optimized designs. The second step further partitions the memory banks for reducing the dynamic power consumption of the final design. In contrast to previous approaches, our technique can statically compute memory access frequencies in polynomial time with little or no profiling. Experimental results show about 30% power reduction on the same set of benchmarks.<\/jats:p>","DOI":"10.1145\/1929943.1929947","type":"journal-article","created":{"date-parts":[[2011,4,6]],"date-time":"2011-04-06T16:08:07Z","timestamp":1302106087000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":64,"title":["Automatic memory partitioning and scheduling for throughput and power optimization"],"prefix":"10.1145","volume":"16","author":[{"given":"Jason","family":"Cong","sequence":"first","affiliation":[{"name":"University of California, Los Angeles, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei","family":"Jiang","sequence":"additional","affiliation":[{"name":"University of California, Los Angeles, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bin","family":"Liu","sequence":"additional","affiliation":[{"name":"University of California, Los Angeles, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Zou","sequence":"additional","affiliation":[{"name":"University of California, Los Angeles, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2011,4,7]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.466632"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/212094.212131"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/155090.155101"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1391962.1391969"},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Bomze I. M. Budinich M. Pardalos P. M. and Pelillo M. 1999. The maximum clique problem. In Handbook of Combinatorial Optimization Kluwer Academic Publishers 1--74.  Bomze I. M. Budinich M. Pardalos P. M. and Pelillo M. 1999. The maximum clique problem. In Handbook of Combinatorial Optimization Kluwer Academic Publishers 1--74.","DOI":"10.1007\/978-1-4757-3023-4_1"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the IEEE Systems on Chip Conference (SOCC).","author":"Cong J.","unstructured":"Cong , J. , Fan , Y. , Han , G. , Jiang , W. , and Zhang , Z . 2006. Platform-based behavior-level and system-level synthesis . In Proceedings of the IEEE Systems on Chip Conference (SOCC). Cong, J., Fan, Y., Han, G., Jiang, W., and Zhang, Z. 2006. Platform-based behavior-level and system-level synthesis. In Proceedings of the IEEE Systems on Chip Conference (SOCC)."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1687399.1687528"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1344671.1344683"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the International Conference on Computer-Aided Design (ICCAD).","author":"Gong W.","unstructured":"Gong , W. , Wang , G. , and Kastner , R . 2005. Storage assignment during high-level synthesis for configurable architectures . In Proceedings of the International Conference on Computer-Aided Design (ICCAD). Gong, W., Wang, G., and Kastner, R. 2005. Storage assignment during high-level synthesis for configurable architectures. In Proceedings of the International Conference on Computer-Aided Design (ICCAD)."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/266021.266171"},{"key":"e_1_2_1_11_1","unstructured":"Kennedy K. and Allen J. R. 2002. Optimizing Compilers for Modern Architectures: A Dependence-Based Approach. Morgan Kaufmann Publishers Inc. San Francisco CA.   Kennedy K. and Allen J. R. 2002. Optimizing Compilers for Modern Architectures: A Dependence-Based Approach. Morgan Kaufmann Publishers Inc. San Francisco CA."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/92.386219"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the International Symposium on Microarchitecture. 330--335","author":"Lee C.","unstructured":"Lee , C. , Potkonjak , M. , and Mangione-Smith , W. H . 1997. MediaBench: A Tool for Evaluating and Synthesizing Multimedia and Communicatons Systems . In Proceedings of the International Symposium on Microarchitecture. 330--335 . Lee, C., Potkonjak, M., and Mangione-Smith, W. H. 1997. MediaBench: A Tool for Evaluating and Synthesizing Multimedia and Communicatons Systems. In Proceedings of the International Symposium on Microarchitecture. 330--335."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/996566.996596"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/344166.344518"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.97903"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/192724.192731"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1049\/el:19991511"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00453-006-1231-0"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the Design, Automation and Test in Europe Conference (DATE).","author":"Wang Z.","unstructured":"Wang , Z. and Hu , X. S . 2004. Power aware variable partitioning and instruction scheduling for multiple memory banks . In Proceedings of the Design, Automation and Test in Europe Conference (DATE). Wang, Z. and Hu, X. S. 2004. Power aware variable partitioning and instruction scheduling for multiple memory banks. In Proceedings of the Design, Automation and Test in Europe Conference (DATE)."}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1929943.1929947","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1929943.1929947","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:26:32Z","timestamp":1750278392000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1929943.1929947"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,3]]},"references-count":20,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2011,3]]}},"alternative-id":["10.1145\/1929943.1929947"],"URL":"https:\/\/doi.org\/10.1145\/1929943.1929947","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,3]]},"assertion":[{"value":"2009-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2010-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-04-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}