{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,12,29]],"date-time":"2022-12-29T05:19:41Z","timestamp":1672291181667},"reference-count":28,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2007,9]]},"abstract":"<jats:p>\n            Modern embedded processors are designed to maximize execution efficiency\u2014the amount of performance achieved per unit of energy dissipated while meeting minimum performance levels. To increase this efficiency, we propose utilizing\n            <jats:italic>static strands<\/jats:italic>\n            , dependence chains without fan-out, which are exposed by a compiler pass. These dependent instructions are resequenced to be sequential and annotated to communicate their location to the hardware. Importantly, this modified application is binary compatible and functionally identical to the original, allowing transparent execution on a baseline processor. However, these static strands can be easily collapsed and optimized by simple processor modifications, significantly reducing the workload energy. Results show that over 30% of MediaBench and Spec2000int dynamic instructions can be collapsed, reducing issue logic energy by 20%, bypass energy 19%, and register file energy 14%. In addition, by increasing the effective capactity of pipeline resources by almost a third, average IPC can be improved up to 15%. This performance gain can then be traded in for a lower clock frequency to maintain a basline level of performance, further reducing energy.\n          <\/jats:p>","DOI":"10.1145\/1274858.1274862","type":"journal-article","created":{"date-parts":[[2007,9,26]],"date-time":"2007-09-26T17:18:32Z","timestamp":1190827112000},"page":"24","update-policy":"http:\/\/dx.doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Static strands"],"prefix":"10.1145","volume":"6","author":[{"given":"Peter G.","family":"Sassone","sequence":"first","affiliation":[{"name":"Intel Corporation, Austin, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"D. Scott","family":"Wills","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, Georgia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gabriel H.","family":"Loh","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, Georgia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2007,9]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Bik A. Girkar M. Grey P. and Tian X. 2001. Efficient exploitation of parallelism on Pentium III and Pentium 4 processor-based systems. In Intel Technology Journal.  Bik A. Girkar M. Grey P. and Tian X. 2001. Efficient exploitation of parallelism on Pentium III and Pentium 4 processor-based systems. In Intel Technology Journal."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2004.15"},{"key":"e_1_2_1_3_1","unstructured":"Brash D. 2002. The ARM architecture version 6 (ARMv6). White paper ARM.  Brash D. 2002. The ARM architecture version 6 (ARMv6). White paper ARM."},{"key":"e_1_2_1_4_1","volume-title":"Tech. Rep. 1342, Dept of Computer Science","author":"Burger D.","year":"1997","unstructured":"Burger , D. and Austin , T . 1997 . The Simplescalar tool set, version 2.0. Tech. Rep. 1342, Dept of Computer Science , University of Wisconsin-Madison. Burger, D. and Austin, T. 1997. The Simplescalar tool set, version 2.0. Tech. Rep. 1342, Dept of Computer Science, University of Wisconsin-Madison."},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the International Symposium on Microarchitecture.","author":"Butts A.","unstructured":"Butts , A. and Sohi , G . 2002. Characterizing and predicting value degree of use . In Proceedings of the International Symposium on Microarchitecture. Butts, A. and Sohi, G. 2002. Characterizing and predicting value degree of use. In Proceedings of the International Symposium on Microarchitecture."},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of IEEE Custom Integrated Circuits Conference.","author":"Cao Y.","unstructured":"Cao , Y. , Sato , T. , Sylvester , D. , Orshansky , M. , and Hu , C . 2000. New paradigm of predictive mosfet and interconnect modeling for early circuit design . In Proceedings of IEEE Custom Integrated Circuits Conference. Cao, Y., Sato, T., Sylvester, D., Orshansky, M., and Hu, C. 2000. New paradigm of predictive mosfet and interconnect modeling for early circuit design. In Proceedings of IEEE Custom Integrated Circuits Conference."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2004.5"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the International Symposium on Microarchitecture.","author":"Corbal J.","unstructured":"Corbal , J. , Valero , M. , and Espasa , R . 1999. Exploiting a new level of DLP in multimedia applications . In Proceedings of the International Symposium on Microarchitecture. Corbal, J., Valero, M., and Espasa, R. 1999. Exploiting a new level of DLP in multimedia applications. In Proceedings of the International Symposium on Microarchitecture."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the International Conference on Parallel Architectures and Compilation Techniques.","author":"Costa A.","unstructured":"Costa , A. , Franca , F. , and Filho , E . 2000. The dynamic trace memoization reuse technique . In Proceedings of the International Conference on Parallel Architectures and Compilation Techniques. Costa, A., Franca, F., and Filho, E. 2000. The dynamic trace memoization reuse technique. In Proceedings of the International Conference on Parallel Architectures and Compilation Techniques."},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the International Symposium on Computer Architecture.","author":"Ernst D.","unstructured":"Ernst , D. and Austin , T . 2002. Efficient dynamic scheduling through tag elimination . In Proceedings of the International Symposium on Computer Architecture. Ernst, D. and Austin, T. 2002. Efficient dynamic scheduling through tag elimination. In Proceedings of the International Symposium on Computer Architecture."},{"key":"e_1_2_1_11_1","first-page":"2","article-title":"The Intel Pentium M processor: Microarchitecture and performance","volume":"7","author":"Gochman S.","year":"2003","unstructured":"Gochman , S. , Ronen , R. , Anati , I. , Berkovits , A. , Kurts , T. , Naveh , A. , Saeed , A. , Sperber , Z. , and Valentine , R. 2003 . The Intel Pentium M processor: Microarchitecture and performance . Intel Technology Journal 7 , 2 (May). Gochman, S., Ronen, R., Anati, I., Berkovits, A., Kurts, T., Naveh, A., Saeed, A., Sperber, Z., and Valentine, R. 2003. The Intel Pentium M processor: Microarchitecture and performance. Intel Technology Journal 7, 2 (May).","journal-title":"Intel Technology Journal"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the International Symposium on High Performance Computer Architecture.","author":"Huang J.","unstructured":"Huang , J. and Lilja , D . 1999. Exploiting basic block value locality with block reuse . In Proceedings of the International Symposium on High Performance Computer Architecture. Huang, J. and Lilja, D. 1999. Exploiting basic block value locality with block reuse. In Proceedings of the International Symposium on High Performance Computer Architecture."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF01205185"},{"key":"e_1_2_1_14_1","unstructured":"IBM Corporation. PowerPC 750 RISC Microprocessor Technical Summary. http:\/\/www-3.ibm.com\/chips\/techlib\/techlib.nsf\/techdocs\/852569B20050FF778525699300470399\/$file\/750_ts.pdfwww.ibm.com.  IBM Corporation. PowerPC 750 RISC Microprocessor Technical Summary. http:\/\/www-3.ibm.com\/chips\/techlib\/techlib.nsf\/techdocs\/852569B20050FF778525699300470399\/$file\/750_ts.pdfwww.ibm.com."},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of the International Conference on Code Generation and Optimization.","author":"Kim H.","unstructured":"Kim , H. and Smith , J . 2003. Dynamic binary translation for accumulator-oriented architectures . In Proceedings of the International Conference on Code Generation and Optimization. Kim, H. and Smith, J. 2003. Dynamic binary translation for accumulator-oriented architectures. In Proceedings of the International Conference on Code Generation and Optimization."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/859618.859623"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the International Symposium on Microarchitecture.","author":"Kim I.","unstructured":"Kim , I. and Lipasti , M . 2003b. Macro-op scheduling: Relaxing scheduling loop constraints . In Proceedings of the International Symposium on Microarchitecture. Kim, I. and Lipasti, M. 2003b. Macro-op scheduling: Relaxing scheduling loop constraints. In Proceedings of the International Symposium on Microarchitecture."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the International Symposium on Microarchitecture.","author":"Lee C.","unstructured":"Lee , C. , Potkonjak , M. , and Mangione-Smith , W . 1997. Mediabench: A tool for evaluating multimedia and communications systems . In Proceedings of the International Symposium on Microarchitecture. Lee, C., Potkonjak, M., and Mangione-Smith, W. 1997. Mediabench: A tool for evaluating multimedia and communications systems. In Proceedings of the International Symposium on Microarchitecture."},{"key":"e_1_2_1_19_1","volume-title":"Tech. Rep. 04-28","author":"Mamidipaka M.","year":"2004","unstructured":"Mamidipaka , M. and Dutt , N . 2004 . eCACTI: An enhanced power estimation model for on-chip caches. Tech. Rep. 04-28 , Center for Embedded Computer Systems, University of California, Irvine . Mamidipaka, M. and Dutt, N. 2004. eCACTI: An enhanced power estimation model for on-chip caches. Tech. Rep. 04-28, Center for Embedded Computer Systems, University of California, Irvine."},{"key":"e_1_2_1_20_1","unstructured":"Marquez A. Theobald K. Tang X. and Gao G. 1997. A superstrand architecture. Technical Memo 14 University of Delaware Computer Architecture and Parallel Systems Laboratory.  Marquez A. Theobald K. Tang X. and Gao G. 1997. A superstrand architecture. Technical Memo 14 University of Delaware Computer Architecture and Parallel Systems Laboratory."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/264107.264201"},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the International Symposium on Microarchitecture.","author":"Park I.","unstructured":"Park , I. , Powell , M. , and Vijaykumar , T . 2002. Reducing register ports for higher speed and lower energy . In Proceedings of the International Symposium on Microarchitecture. Park, I., Powell, M., and Vijaykumar, T. 2002. Reducing register ports for higher speed and lower energy. In Proceedings of the International Symposium on Microarchitecture."},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the Computer Architecture and High Performance Computing.","author":"Pilla M.","unstructured":"Pilla , M. , Navaux , P. , Costa , A. , Franca , F. , Childers , B. , and Soffa , M . 2003. The limits of speculative trace reuse on deeply pipelined processors . In Proceedings of the Computer Architecture and High Performance Computing. Pilla, M., Navaux, P., Costa, A., Franca, F., Childers, B., and Soffa, M. 2003. The limits of speculative trace reuse on deeply pipelined processors. In Proceedings of the Computer Architecture and High Performance Computing."},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the International Symposium on Computer Architecture.","author":"Raasch S.","unstructured":"Raasch , S. , Binkert , N. , and Reinhardt , S . 2002. A scalable instruction queue design using dependence chains . In Proceedings of the International Symposium on Computer Architecture. Raasch, S., Binkert, N., and Reinhardt, S. 2002. A scalable instruction queue design using dependence chains. In Proceedings of the International Symposium on Computer Architecture."},{"key":"e_1_2_1_25_1","unstructured":"Renesas Technology. SH-4A Software Manual. http:\/\/documentation.renesas.com\/eng\/products\/mpumcu\/rej09b0003_sh4a.pdfwww.renesas.com.  Renesas Technology. SH-4A Software Manual. http:\/\/documentation.renesas.com\/eng\/products\/mpumcu\/rej09b0003_sh4a.pdfwww.renesas.com."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2004.16"},{"key":"e_1_2_1_27_1","unstructured":"UC Berkeley. Berkeley predictive technology model. http:\/\/www-device.eecs.berkeley.edu\/~ptmwww-device.eecs.berkeley.edu\/~ptm.  UC Berkeley. Berkeley predictive technology model. http:\/\/www-device.eecs.berkeley.edu\/~ptmwww-device.eecs.berkeley.edu\/~ptm."},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the International Symposium on Computer Architecture.","author":"Yehia S.","unstructured":"Yehia , S. and Temam , O . 2004. From sequences of dependent instructions to functions: A complexity-effective approach for improving performance without ILP or speculation . In Proceedings of the International Symposium on Computer Architecture. Yehia, S. and Temam, O. 2004. From sequences of dependent instructions to functions: A complexity-effective approach for improving performance without ILP or speculation. In Proceedings of the International Symposium on Computer Architecture."}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1274858.1274862","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T18:31:21Z","timestamp":1672252281000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1274858.1274862"}},"subtitle":["Safely exposing dependence chains for increasing embedded power efficiency"],"short-title":[],"issued":{"date-parts":[[2007,9]]},"references-count":28,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2007,9]]}},"alternative-id":["10.1145\/1274858.1274862"],"URL":"https:\/\/doi.org\/10.1145\/1274858.1274862","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2007,9]]},"assertion":[{"value":"2007-09-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}