{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:40:05Z","timestamp":1750192805392,"version":"3.41.0"},"reference-count":49,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2017]]},"DOI":"10.1587\/transinf.2016edp7414","type":"journal-article","created":{"date-parts":[[2017,3,31]],"date-time":"2017-03-31T22:24:35Z","timestamp":1490999075000},"page":"822-837","source":"Crossref","is-referenced-by-count":0,"title":["Skewed Multistaged Multibanked Register File for Area and Energy Efficiency"],"prefix":"10.1587","volume":"E100.D","author":[{"given":"Junji","family":"YAMADA","sequence":"first","affiliation":[{"name":"Graduate School of Information Science and Technology, The University of Tokyo"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ushio","family":"JIMBO","sequence":"additional","affiliation":[{"name":"Department of Informatics, School of Multidisciplinary Sciences, SOKENDAI (Graduate University for Advanced Studies)"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryota","family":"SHIOYA","sequence":"additional","affiliation":[{"name":"Graduate School of Engineering, Nagoya University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Masahiro","family":"GOSHIMA","sequence":"additional","affiliation":[{"name":"Information Systems Architecture Research Division, National Institute of Informatics"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuichi","family":"SAKAI","sequence":"additional","affiliation":[{"name":"Graduate School of Information Science and Technology, The University of Tokyo"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","unstructured":"[1] The Standard Performance Evaluation Corporation, SPEC CPU 2006. http:\/\/www.spec.org\/cpu2006\/"},{"key":"2","unstructured":"[2] J.L. Hennessy and D.A. Patterson, Computer Architecture \u2014 A Quantitative Approach, 5th ed., Morgan Kaufmann, 2011."},{"key":"3","doi-asserted-by":"crossref","unstructured":"[3] B. Sinharoy, R. Kalla, W.J. Starke, H.Q. Le, R. Cargnoni, J.A. Van Norstrand, B.J. Ronchetti, J. Stuecheli, J. Leenstra, G.L. Guthrie, D.Q. Nguyen, B. Blaner, C.F. Marino, E. Retter, and P. Williams, \u201cIBM POWER7 multicore server processor,\u201d IBM Journal of Research and Development, vol.55, no.3, pp.1:1-1:29, May 2011.","DOI":"10.1147\/JRD.2011.2127330"},{"key":"4","doi-asserted-by":"crossref","unstructured":"[4] B. Sinharoy, J.A. Van Norstrand, R.J. Eickemeyer, H.Q. Le, J. Leenstra, D.Q. Nguyen, B. Konigsburg, K. Ward, M.D. Brown, J.E. Moreira, D. Levitan, S. Tung, D. Hrusecky, J.W. Bishop, M. Gschwind, M. Boersma, M. Kroener, M. Kaltenbach, T. Karkhanis, and K.M. Fernsler, \u201cIBM POWER8 processor core microarchitecture,\u201d IBM Journal of Research and Development, vol.59, no.1, pp.2:1-2:21, Jan. 2015.","DOI":"10.1147\/JRD.2014.2376112"},{"key":"5","unstructured":"[5] K. Krewell, \u201cIntel&apos;s Haswell cuts core power,\u201d Microprocessor Report, Sept. 2012."},{"key":"6","doi-asserted-by":"crossref","unstructured":"[6] E. Fayneh, M. Yuffe, E. Knoll, M. Zelikson, M. Abozaed, Y. Talker, Z. Shmuely, and S.A. Rahme, \u201c4.1 14nm 6th-generation Core processor SoC with low power consumption and improved performance,\u201d IEEE International Solid-State Circuits Conference on (ISSCC 2016), Digest of Technical Papers, pp.72-73, Jan. 2016.","DOI":"10.1109\/ISSCC.2016.7417912"},{"key":"7","unstructured":"[7] N.H.E. Weste and D.M. Harris, CMOS VLSI Design: A Circuits and Systems Perspective, 4th ed., Addison Wesley, 2011."},{"key":"8","doi-asserted-by":"crossref","unstructured":"[8] S. Rixner, W.J. Dally, B. Khailany, P. Mattson, U.J. Kapasi, and J.D. Owens, \u201cRegister organization for media processing,\u201d Proc. International Symposium on High-Performance Computer Architecture (HPCA), pp.375-386, 2000.","DOI":"10.1109\/HPCA.2000.824366"},{"key":"9","unstructured":"[9] S. Thoziyoor, N. Muralimanohar, J. Ahn, and N. Jouppi, \u201cCACTI 5.1,\u201d Tech. Rep. HPL-2008-20, HP Laboratories, 2008."},{"key":"10","doi-asserted-by":"crossref","unstructured":"[10] T. Fischer, S. Arekapudi, E. Busta, C. Dietz, M. Golden, S. Hilker, A. Horiuchi, K.A. Hurd, D. Johnson, H. McIntyre, S. Naffziger, J. Vinh, J. White, and K. Wilcox, \u201cDesign solutions for the Bulldozer 32nm SOI 2-core processor module in an 8-core CPU,\u201d IEEE International Solid-State Circuits Conference on (ISSCC 2011), Digest of Technical Papers, pp.78-80, Feb. 2011.","DOI":"10.1109\/ISSCC.2011.5746227"},{"key":"11","doi-asserted-by":"crossref","unstructured":"[11] M. Golden, S. Arekapudi, and J. Vinh, \u201c40-entry unified out-of-order scheduler and integer execution unit for the AMD Bulldozer x86-64 core,\u201d IEEE International Solid-State Circuits Conference on (ISSCC 2011), Digest of Technical Papers, pp.80-82, Feb. 2011.","DOI":"10.1109\/ISSCC.2011.5746228"},{"key":"12","doi-asserted-by":"crossref","unstructured":"[12] H. McIntyre, S. Arekapudi, E. Busta, T. Fischer, M. Golden, A. Horiuchi, T. Meneghini, S. Naffziger, and J. Vinh, \u201cDesign of the two-core x86-64 AMD \u201cBulldozer\u201d module in 32nm SOI CMOS,\u201d IEEE J. Solid-State Circuits, vol.47, no.1, pp.164-176, Jan. 2012.","DOI":"10.1109\/JSSC.2011.2167823"},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] O. Ergin, D. Balkan, K. Ghose, and D. Ponomarev, \u201cRegister packing: Exploiting narrow-width operands for reducing register file pressure,\u201d Proc. 37th Annual International Symposium on Microarchitecture (MICRO), pp.304-315, 2004.","DOI":"10.1109\/MICRO.2004.29"},{"key":"14","doi-asserted-by":"crossref","unstructured":"[14] J. Abella, J. Carretero, P. Chaparro, and X. Vera, \u201cThe split register file,\u201d Design, Automation Test in Europe Conference Exhibition (DATE), pp.945-948, March 2010.","DOI":"10.1109\/DATE.2010.5456914"},{"key":"15","doi-asserted-by":"crossref","unstructured":"[15] G.S. Ditlow, R.K. Montoye, S.N. Storino, S.M. Dance, S. Ehrenreich, B.M. Fleischer, T.W. Fox, K.M. Holmes, J. Mihara, Y. Nakamura, S. Onishi, R. Shearer, D. Wendel, and L. Chang, \u201cA 4R2W register file for a 2.3GHz wire-speed POWER<sup>TM<\/sup> processor with double-pumped write operation,\u201d IEEE International Solid-State Circuits Conference on (ISSCC 2011), Digest of Technical Papers, pp.256-258, Feb. 2011.","DOI":"10.1109\/ISSCC.2011.5746308"},{"key":"16","doi-asserted-by":"crossref","unstructured":"[16] J.L. Shin, R. Golla, H. Li, S. Dash, Y. Choi, A. Smith, H. Sathianathan, M. Joshi, H. Park, M. Elgebaly, S. Turullols, S. Kim, R. Masleid, G.K. Konstadinidis, M.J. Doherty, G. Grohoski, and C. McAllister, \u201cThe next generation 64b SPARC core in a T4 SoC processor,\u201d IEEE J. Solid-State Circuits, vol.48, no.1, pp.82-90, Jan. 2013.","DOI":"10.1109\/JSSC.2012.2223036"},{"key":"17","doi-asserted-by":"crossref","unstructured":"[17] J. Feehrer, S. Jairath, P. Loewenstein, R. Sivaramakrishnan, D. Smentek, S. Turullols, and A. Vahidsafa, \u201cThe Oracle Sparc T5 16-core processor scales to eight sockets,\u201d IEEE Micro, vol.33, no.2, pp.48-57, March 2013.","DOI":"10.1109\/MM.2013.49"},{"key":"18","doi-asserted-by":"crossref","unstructured":"[18] J.M. Hart, H. Cho, Y. Ge, G. Gruber, D. Huang, C. Hwang, D. Jian, T. Johnson, G.K. Konstadinidis, V. Krishnaswamy, L. Kwong, R.P. Masleid, R. Mehta, U. Nawathe, A. Ramachandran, H. Sathianathan, Y. Sheng, J.L. Shin, S. Turullols, Z. Qin, and K.C. Yen, \u201cA 3.6GHz 16-core SPARC SoC processor in 28 nm,\u201d IEEE J. Solid-State Circuits, vol.49, no.1, pp.19-31, Jan. 2014.","DOI":"10.1109\/JSSC.2013.2284648"},{"key":"19","doi-asserted-by":"crossref","unstructured":"[19] J.A. Butts and G.S. Sohi, \u201cUse-based register caching with decoupled indexing,\u201d Proc. International Symposium on Computer Architecture (ISCA), pp.302-313, 2004.","DOI":"10.1109\/ISCA.2004.1310783"},{"key":"20","doi-asserted-by":"crossref","unstructured":"[20] R. Shioya, K. Horio, M. Goshima, and S. Sakai, \u201cRegister cache system not for latency reduction purpose,\u201d Proc. IEEE International Symposium on Microarchitecture (MICRO), pp.301-312, Dec. 2010.","DOI":"10.1109\/MICRO.2010.43"},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] M. Gebhart, D.R. Johnson, D. Tarjan, S.W. Keckler, W.J. Dally, E. Lindholm, and K. Skadron, \u201cEnergy-efficient mechanisms for managing thread context in throughput processors,\u201d Proc. 38th Annual International Symposium on Computer Architecture (ISCA), ISCA &apos;11, New York, NY, USA, pp.235-246, ACM, 2011.","DOI":"10.1145\/2000064.2000093"},{"key":"22","doi-asserted-by":"crossref","unstructured":"[22] M. Gebhart, D.R. Johnson, D. Tarjan, S.W. Keckler, W.J. Dally, E. Lindholm, and K. Skadron, \u201cA hierarchical thread scheduler and register file for energy-efficient throughput processors,\u201d ACM Trans. Comput. Syst., vol.30, no.2, pp.8:1-8:38, April 2012.","DOI":"10.1145\/2166879.2166882"},{"key":"23","doi-asserted-by":"crossref","unstructured":"[23] J.H. Tseng and K. Asanovic, \u201cBanked multiported register files for high-frequency superscalar microprocessors,\u201d Proc. 30th annual IEEE International Symposium on Computer Architecture (ISCA), pp.62-71, June 2003.","DOI":"10.1145\/859618.859627"},{"key":"24","doi-asserted-by":"crossref","unstructured":"[24] J.H. Tseng and K. Asanovic, \u201cA speculative control scheme for an energy-efficient banked register file,\u201d IEEE Transactions on Computers, vol.54, no.6, pp.741-751, June 2005.","DOI":"10.1109\/TC.2005.88"},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] T. Hironaka, M. Maeda, K. Tanigawa, T. Sueyoshi, K. Aoyama, T. Koide, H. Mattausch, and T. Saito, \u201cSuperscalar processor with multi-bank register file,\u201d Proc. Innovative Architecture for Future Generation High-Performance Processors and Systems (IWIA), pp.3-12, Jan. 2005.","DOI":"10.1109\/IWIA.2005.42"},{"key":"26","doi-asserted-by":"crossref","unstructured":"[26] N. Duong and R. Kumar, \u201cRegister multimapping: A technique for reducing register bank conflicts in processors with large register files,\u201d Proc. IEEE Symp. Application Specific Processors (SASP &apos;09), pp.50-53, 2009.","DOI":"10.1109\/SASP.2009.5226335"},{"key":"27","doi-asserted-by":"crossref","unstructured":"[27] I. Park, M.D. Powell, and T.N. Vijaykumar, \u201cReducing register ports for higher speed and lower energy,\u201d Proc. International Symposium on Microarchitecture (MICRO), pp.171-182, 2002.","DOI":"10.1109\/MICRO.2002.1176248"},{"key":"28","doi-asserted-by":"crossref","unstructured":"[28] K.C. Yeager, \u201cThe MIPS R10000 superscalar microprocessor,\u201d IEEE Micro, vol.16, no.2, pp.28-40, April 1996.","DOI":"10.1109\/40.491460"},{"key":"29","doi-asserted-by":"crossref","unstructured":"[29] J.R. Diamond, D.S. Fussell, and S.W. Keckler, \u201cArbitrary modulus indexing,\u201d Proc. 47th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO), pp.140-152, Dec. 2014.","DOI":"10.1109\/MICRO.2014.13"},{"key":"30","doi-asserted-by":"crossref","unstructured":"[30] J.H. Edmondson, P. Rubinfeld, R. Preston, and V. Rajagopalan, \u201cSuperscalar instruction execution in the 21164 Alpha microprocessor,\u201d IEEE Micro, vol.15, no.2, pp.33-43, April 1995.","DOI":"10.1109\/40.372349"},{"key":"31","unstructured":"[31] O. Brun and J.-M. Garcia, \u201cAnalytical solution of finite capacity M\/D\/1 queues,\u201d Journal of Applied Probability, vol.37, no.4, pp.1092-1098, Dec. 2000."},{"key":"32","unstructured":"[32] \u201cOnikiri 2.\u201d https:\/\/github.com\/onikiri\/onikiri2\/"},{"key":"33","doi-asserted-by":"crossref","unstructured":"[33] A. Perais, A. Seznec, P. Michaud, A. Sembranty, and E. Hagersten, \u201cCost-effective speculative scheduling in high performance processors,\u201d Proc. 42nd Annual International Symposium on Computer Architecture (ISCA), pp.247-259, 2015.","DOI":"10.1145\/2749469.2749470"},{"key":"34","doi-asserted-by":"crossref","unstructured":"[34] K. Bhanushali and W.R. Davis, \u201cFreePDK15: An open-source predictive process design kit for 15nm FinFET technology,\u201d Proc. International Symposium on Physical Design (ISPD), pp.165-170, 2015.","DOI":"10.1145\/2717764.2717782"},{"key":"35","doi-asserted-by":"crossref","unstructured":"[35] M. Martins, J.M. Matos, R.P. Ribas, A. Reis, G. Schlinker, L. Rech, and J. Michelsen, \u201cOpen cell library in 15nm FreePDK technology,\u201d Proc. International Symposium on Physical Design (ISPD), pp.171-178, 2015.","DOI":"10.1145\/2717764.2717783"},{"key":"36","unstructured":"[36] N. Muralimanohar, R. Balasubramonian, and N.P. Jouppi, \u201cCACTI 6.0: A tool to model large caches,\u201d Tech. Rep. HPL-2009-85, HP Laboratories, 2009."},{"key":"37","doi-asserted-by":"crossref","unstructured":"[37] S. Li, J.H. Ahn, R.D. Strong, J.B. Brockman, D.M. Tullsen, and N.P. Jouppi, \u201cThe McPAT framework for multicore and manycore architectures: Simultaneously modeling power, area, and timing,\u201d ACM Trans. Architecture and Code Optimization, vol.10, no.1, pp.5:1-5:29, April 2013.","DOI":"10.1145\/2445572.2445577"},{"key":"38","unstructured":"[38] S.Y. Wu, J.J. Liaw, C.Y. Lin, M.C. Chiang, C.K. Yang, J.Y. Cheng, M.H. Tsai, M.Y. Liu, P.H. Wu, C.H. Chang, L.C. Hu, C.I. Lin, H.F. Chen, S.Y. Chang, S.H. Wang, P.Y. Tong, Y.L. Hsieh, K.H. Pan, C.H. Hsieh, C.H. Chen, C.H. Yao, C.C. Chen, T.L. Lee, C.W. Chang, H.J. Lin, S.C. Chen, J.H. Shieh, S.M. Jang, K.S. Chen, Y. Ku, Y.C. See, and W.J. Lo, \u201cA highly manufacturable 28nm CMOS low power platform technology with fully functional 64Mb SRAM using dual\/tripe gate oxide process,\u201d Proc. Symposium on VLSI Technology 2009, pp.210-211, June 2009."},{"key":"39","doi-asserted-by":"crossref","unstructured":"[39] M. Yabuuchi, H. Fujiwara, Y. Tsukamoto, M. Tanaka, S. Tanaka, and K. Nii, \u201cA 28nm high density 1R\/1W 8T-SRAM macro with screening circuitry against read disturb failure,\u201d Proc. IEEE Custom Integrated Circuits Conference (CICC), pp.1-4, Sept. 2013.","DOI":"10.1109\/CICC.2013.6658451"},{"key":"40","doi-asserted-by":"crossref","unstructured":"[40] I. Kim and M.H. Lipasti, \u201cHalf-price architecture,\u201d Proc. International Symposium on Computer Architecture (ISCA), pp.28-38, 2003.","DOI":"10.1145\/859618.859623"},{"key":"41","doi-asserted-by":"crossref","unstructured":"[41] R. Sangireddy, \u201cReducing rename logic complexity for high-speed and low-power front-end architectures,\u201d IEEE Trans. Comput., vol.55, no.6, pp.672-685, June 2006.","DOI":"10.1109\/TC.2006.88"},{"key":"42","doi-asserted-by":"crossref","unstructured":"[42] G.S. Sohi, S.E. Breach, and T.N. Vijaykumar, \u201cMultiscalar processors,\u201d Proc. 22nd Annual International Symposium on Computer Architecture (ISCA), pp.414-425, 1995.","DOI":"10.1145\/223982.224451"},{"key":"43","doi-asserted-by":"crossref","unstructured":"[43] G.A. Kemp and M. Franklin, \u201cPEWs: a decentralized dynamic scheduler for ILP processing,\u201d Proc. International Conference on Parallel Processing, pp.239-246, Aug. 1996.","DOI":"10.1109\/ICPP.1996.537165"},{"key":"44","doi-asserted-by":"crossref","unstructured":"[44] S. Palacharla, N.P. Jouppi, and J.E. Smith, \u201cComplexity-effective superscalar processors,\u201d Proc. 24th Annual International Symposium on Computer Architecture (ISCA), pp.206-218, 1997.","DOI":"10.1145\/264107.264201"},{"key":"45","doi-asserted-by":"crossref","unstructured":"[45] K.I. Farkas, P. Chow, N.P. Jouppi, and Z. Vranesic, \u201cThe multicluster architecture: reducing cycle time through partitioning,\u201d Proc. 13th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO), pp.149-159, Dec. 1997.","DOI":"10.1109\/MICRO.1997.645806"},{"key":"46","doi-asserted-by":"crossref","unstructured":"[46] R. Kessler, \u201cThe Alpha 21264 microprocessor,\u201d IEEE Micro, vol.19, no.2, pp.24-36, March\/April 1999.","DOI":"10.1109\/40.755465"},{"key":"47","doi-asserted-by":"crossref","unstructured":"[47] P. Salverda and C. Zilles, \u201cA criticality analysis of clustering in superscalar processors,\u201d Proc. 38th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO), pp.12-66, Nov. 2005.","DOI":"10.1109\/MICRO.2005.6"},{"key":"48","doi-asserted-by":"crossref","unstructured":"[48] F. Tseng and Y.N. Patt, \u201cAchieving out-of-order performance with almost in-order complexity,\u201d Proc. 35th International Symposium on Computer Architecture (ISCA), pp.3-12, June 2008.","DOI":"10.1109\/ISCA.2008.23"},{"key":"49","doi-asserted-by":"crossref","unstructured":"[49] T. Stripf, R. Koenig, P. Rieder, and J. Becker, \u201cA compiler back-end for reconfigurable, mixed-ISA processors with clustered register files,\u201d Proc. IEEE 26th International Parallel and Distributed Processing Symposium Workshops &amp; PhD Forum (IPDPSW), pp.462-469, May 2012.","DOI":"10.1109\/IPDPSW.2012.60"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E100.D\/4\/E100.D_2016EDP7414\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:12:27Z","timestamp":1750191147000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E100.D\/4\/E100.D_2016EDP7414\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017]]},"references-count":49,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2017]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2016edp7414","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"type":"print","value":"0916-8532"},{"type":"electronic","value":"1745-1361"}],"subject":[],"published":{"date-parts":[[2017]]}}}