{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,26]],"date-time":"2025-09-26T13:26:42Z","timestamp":1758893202839},"reference-count":54,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2014,2,6]],"date-time":"2014-02-06T00:00:00Z","timestamp":1391644800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2014,5]]},"DOI":"10.1007\/s11227-013-1075-8","type":"journal-article","created":{"date-parts":[[2014,2,5]],"date-time":"2014-02-05T05:35:56Z","timestamp":1391578556000},"page":"948-977","source":"Crossref","is-referenced-by-count":1,"title":["Customized pipeline and instruction set architecture for embedded processing engines"],"prefix":"10.1007","volume":"68","author":[{"given":"Amir","family":"Yazdanbakhsh","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mostafa E.","family":"Salehi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sied Mehdi","family":"Fakhraie","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2014,2,6]]},"reference":[{"key":"1075_CR1","doi-asserted-by":"crossref","unstructured":"Swanson S, Putnam A, Mercaldi M, Michelson K, Petersen A, Schwerin A, Oskin M, Eggers SJ (2006) Area-performance trade-offs in tiled dataflow architectures. in: Proceedings of the 33rd international symposium on computer architecture (ISCA\u201906), pp. 314\u2013326","DOI":"10.1109\/ISCA.2006.10"},{"issue":"2","key":"1075_CR2","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1109\/MM.2010.41","volume":"30","author":"J Nickolls","year":"2010","unstructured":"Nickolls J, Dally WJ (2010) The GPU computing era. IEEE Micro 30(2):56\u201359","journal-title":"IEEE Micro"},{"key":"1075_CR3","doi-asserted-by":"crossref","unstructured":"Lee SJ (2010) A 345 mW heterogeneous many-core processor with an intelligent inference engine for robust object recognition. In: Porceedings of the IEEE international solid-state circuits conference, 2010, pp. 332\u2013334","DOI":"10.1109\/ISSCC.2010.5433905"},{"key":"1075_CR4","unstructured":"Bell S, et al (2008) TILE64 $$^{TM}$$ T M processor: a 64-core SoC with mesh interconnect. In: Porceedings ofthe IEEE international solid-state circuits conference, pp. 88\u201390"},{"key":"1075_CR5","unstructured":"Jotwani R, et al (2010) An x86\u201364 core implemented in 32 nm SOI CMOS. In: Porceedings of the IEEE international solid-state circuits conference, pp. 106\u2013107"},{"key":"1075_CR6","unstructured":"Howard J, et al (2010) A 48-Core IA-32 message-passing processor with DVFS in 45 nm CMOS. In: Poreedings of the IEEE international solid-state circuits conference, pp. 108\u2013110"},{"key":"1075_CR7","unstructured":"Shin JL, et al (2010) A 40 nm 16-Core 128-thread CMT SPARC SoC processor. In: Porceedings of the IEEE international solid-state circuits conference, pp. 98\u201399"},{"key":"1075_CR8","unstructured":"Johnson C, et al (2010) A wire-speed POWER $$^{TM}$$ T M processor: 2.3G Hz, 45 nm SOI with 16 cores and 64 threads. In: Porceedings of the IEEE international solid-state circuits conference, pp. 104\u2013106"},{"key":"1075_CR9","doi-asserted-by":"crossref","unstructured":"Azizi O, Mahesri A, Lee BC, Patel SJ, Horowitz M (2010) Energy-performance tradeoffs in processor architecture and circuit design: a marginal cost analysis. In: Proceedings of the 37th international symposium on computer architecture (ISCA\u201910), pp. 26\u201336","DOI":"10.1145\/1816038.1815967"},{"key":"1075_CR10","doi-asserted-by":"crossref","unstructured":"Kapre N, DeHon A (2009) Performance comparison of single-precision SPICE model-evaluation on FPGA, GPU, Cell, and multi-core processors. In: Proceedings of the international conference on field programmable logic and applications, pp. 65\u201372","DOI":"10.1109\/FPL.2009.5272548"},{"issue":"4","key":"1075_CR11","doi-asserted-by":"crossref","first-page":"1130","DOI":"10.1109\/JSSC.2009.2013772","volume":"44","author":"DN Truong","year":"2009","unstructured":"Truong DN et al (2009) A 167-processor computational platform in 65 nm CMOS. IEEE J Solid State Circuits 44(4):1130\u20131144","journal-title":"IEEE J Solid State Circuits"},{"issue":"7","key":"1075_CR12","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1109\/MC.2008.209","volume":"41","author":"MD Hill","year":"2008","unstructured":"Hill MD, Marty MR (2008) Amdahl\u2019s law in the multicore era. IEEE Comput 41(7):33\u201338","journal-title":"IEEE Comput"},{"key":"1075_CR13","unstructured":"Borkar S (2007) Thousand core chips\u2014a technology perspective. In: Proceedings of the design automation conference (DAC), pp. 746\u2013749"},{"key":"1075_CR14","doi-asserted-by":"crossref","unstructured":"Eyerman S, Eeckhout L (2010) Modeling critical sections in Amdahl\u2019s Law and its implications for multicore design. In: Proceedings of the 37th international symposium on computer, architecture (ISCA\u201910), pp. 362\u2013370","DOI":"10.1145\/1816038.1816011"},{"issue":"6","key":"1075_CR15","doi-asserted-by":"crossref","first-page":"1155","DOI":"10.1109\/TCAD.2008.923254","volume":"27","author":"S Park","year":"2008","unstructured":"Park S, Shrivastava A, Dutt N, Nicolau A, Paek Y, Earlie E (2008) Register file power reduction using bypass sensitive compiler. IEEE Trans Comput Aided Des Integr Circuits Syst 27(6):1155\u20131159","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"key":"1075_CR16","doi-asserted-by":"crossref","unstructured":"Nalluri R, Garg R, Panda PR (2007) Customization of register file banking architecture for low power. In: Proceedings of the 20th international conference on VLSI design (VLSID\u201907), pp. 239\u2013244","DOI":"10.1109\/VLSID.2007.58"},{"issue":"10","key":"1075_CR17","doi-asserted-by":"crossref","first-page":"1259","DOI":"10.1109\/TVLSI.2008.2001863","volume":"16","author":"P Bonzini","year":"2008","unstructured":"Bonzini P, Pozzi L (2008) Recurrence-aware instruction set selection for extensible embedded processors. IEEE Trans Very Large Scale Integr (VLSI) Syst 16(10):1259\u20131267","journal-title":"IEEE Trans Very Large Scale Integr (VLSI) Syst"},{"key":"1075_CR18","doi-asserted-by":"crossref","unstructured":"Atasu K, Pozzi L, Ienne P (2003) Automatic application-specific instruction-set extensions under microarchitectural constraints. In: Proceedings of the design automation conference (DAC), pp. 256\u2013261","DOI":"10.1145\/775832.775897"},{"key":"1075_CR19","doi-asserted-by":"crossref","unstructured":"Clark N, Zhong H, Mahlke S (2003) Processor acceleration through automated instruction set customization. In: Proceedings of the 36th Annu. IEEE\/ACM, MICRO, pp. 129\u2013140","DOI":"10.1109\/MICRO.2003.1253189"},{"key":"1075_CR20","doi-asserted-by":"crossref","unstructured":"Yu P, Mitra T (2004) Scalable custom instructions identification for instruction-set extensible processors. In: Proceedings of the CASES, pp. 69\u201378","DOI":"10.1145\/1023833.1023844"},{"key":"1075_CR21","doi-asserted-by":"crossref","first-page":"1209","DOI":"10.1109\/TCAD.2005.855950","volume":"25","author":"L Pozzi","year":"2006","unstructured":"Pozzi L, Atasu K, Ienne P (2006) Exact and approximate algorithms for the extension of embedded processor instruction sets. IEEE Trans Comput Aided Des Integr Circuits Syst 25:1209\u20131229","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"issue":"2","key":"1075_CR22","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1109\/TCAD.2006.883915","volume":"26","author":"X Chen","year":"2007","unstructured":"Chen X, Maskell DL, Sun Y (2007) Fast identification of custom instructions for extensible processors. IEEE Trans Comput Aided Des Integr Circuits Syst 26(2):359\u2013368","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"key":"1075_CR23","unstructured":"Zyuban VV, Kogge PM (1998) The energy complexity of register files. In: Proceedings of the international symposium on low power, electronic design, pp. 305\u2013310"},{"key":"1075_CR24","doi-asserted-by":"crossref","unstructured":"Leupers R, Karuri K, Kraemer S, Pandey M (2006) A design flow for configurable embedded processors based on optimized instruction set extension synthesis. In: Proceedings of the design, automation & test in Europe (DATE)","DOI":"10.1109\/DATE.2006.243972"},{"key":"1075_CR25","unstructured":"Altera Corp. Nios processor reference handbook"},{"key":"1075_CR26","unstructured":"Xilinx Inc., Microblaze soft processor core"},{"key":"1075_CR27","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1109\/40.848473","volume":"20","author":"RE Gonzalez","year":"2000","unstructured":"Gonzalez RE (2000) XTENSA: a configurable and extensible processor. IEEE Micro 20:60\u201370","journal-title":"IEEE Micro"},{"key":"1075_CR28","doi-asserted-by":"crossref","unstructured":"Karuri K, Chattopadhyay A, Hohenauer M, Leupers R, Ascheid G, Meyr H (2007) Increasing data-bandwidth to instruction-set extensions through register clustering. In: Proceedings of the international conference on computer aided design, pp. 166\u2013177","DOI":"10.1109\/ICCAD.2007.4397261"},{"key":"1075_CR29","unstructured":"Fischer JA, Faraboschi P, Young C (2005) Embedded computing: a VLIW approach to architecture. Elsevier Inc, Compiler and Tools, Amsterdam"},{"key":"1075_CR30","doi-asserted-by":"crossref","unstructured":"Kim NS, Mudge T (2003) Reducing register ports using delayed write-back queues and operand pre-fetch. In: Proceedings of the 17th annual international conference on Supercomputing, pp. 172\u2013182","DOI":"10.1145\/782814.782839"},{"key":"1075_CR31","doi-asserted-by":"crossref","unstructured":"Pozzi L, Ienne P (2005) Exploiting pipelining to relax register-file port constraints of instruction set extensions. In: Proceedings of the international conference on compilers, architectures and synthesis for embedded systems, pp. 2\u201310","DOI":"10.1145\/1086297.1086300"},{"key":"1075_CR32","doi-asserted-by":"crossref","unstructured":"Atasu K, Dimond R, Mencer O, Luk W, \u00d6zturan C, D\u00fcnda G (2007) Optimizing instruction-set extensible processors under data bandwidth constraints. In: Proceedings of the design automation and test in, Europe, Mar. 2007, pp. 588\u2013593","DOI":"10.1109\/DATE.2007.364657"},{"issue":"3","key":"1075_CR33","doi-asserted-by":"crossref","first-page":"528","DOI":"10.1109\/TCAD.2008.915536","volume":"27","author":"K Atasu","year":"2008","unstructured":"Atasu K, Ozturan C, Dundar G, Mencer O, Luk W (2008) CHIPS: custom hardware instruction processor synthesis. IEEE Trans Comput Aided Des Integr Circuits Syst 27(3):528\u2013541","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"issue":"3","key":"1075_CR34","doi-asserted-by":"crossref","first-page":"341","DOI":"10.1109\/TCAD.2010.2041849","volume":"29","author":"Ajay K Verma","year":"2010","unstructured":"Verma Ajay K, Brisk Philip, Ienne Paolo (2010) Fast, nearly optimal ISE identification with I\/O serialization through maximal clique enumeration. IEEE Trans Comput Aided Des Integr Circuits Syst 29(3):341\u2013354","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"key":"1075_CR35","doi-asserted-by":"crossref","unstructured":"Brisk P, Kaplan A, Sarrafzadeh M (2004) Area-efficient instruction set synthesis for reconfigurable system-on-chip designs. In: Proceedings of the design automation conference (DAC), pp. 395\u2013400","DOI":"10.1145\/996566.996679"},{"issue":"7","key":"1075_CR36","doi-asserted-by":"crossref","first-page":"969","DOI":"10.1109\/TCAD.2005.850844","volume":"24","author":"N Moreano","year":"2005","unstructured":"Moreano N, Borin E, de Souza C, Araujo G (2005) Efficient datapath merging for partially reconfigurable architectures. IEEE Trans Comput Aided Des Integr Circuits Syst 24(7):969\u2013980","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"key":"1075_CR37","doi-asserted-by":"crossref","unstructured":"Dinh Q, Chen D, Wong MDF (2008) Efficient ASIP design for configurable processors with fine-grained resource sharing. In: Proceedings of the ACM\/SIGDA 16th international symposium on FPGA, pp. 99\u2013106","DOI":"10.1145\/1344671.1344687"},{"issue":"12","key":"1075_CR38","doi-asserted-by":"crossref","first-page":"1788","DOI":"10.1109\/TCAD.2009.2026355","volume":"28","author":"M Zuluaga","year":"2009","unstructured":"Zuluaga M, Topham N (2009) Design-space exploration of resource-sharing solutions for custom instruction set extensions. IEEE Trans Comput Aided Des Integr Circuits Syst 28(12):1788\u20131801","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"key":"1075_CR39","unstructured":"Hennessy JL, Patterson DA (2005) Computer organization and design: the hardware\/software interface, the Morgan Kaufmann Series in computer architecture and design, 3rd edn. Elsevier Inc., Amsterdam"},{"key":"1075_CR40","unstructured":"Powell PMD, Vijaykumar TN (2002) Reducing register ports for higher speed and lower energy. In: Proceedings of the 35th annual IEEE\/ACM international symposium on microarchitecture, pp. 171\u2013182"},{"key":"1075_CR41","doi-asserted-by":"crossref","unstructured":"Cong J, et al (2005) Instruction set extension with shadow registers for configurable processors. In: Proceedings of the FPGA, pp. 99\u2013106","DOI":"10.1145\/1046192.1046206"},{"key":"1075_CR42","unstructured":"Liu H, Jayaseelan R, Mitra T (2006) Exploiting forwarding to improve data bandwidth of instruction-set extensions. In: Proceedings of the design automation conference (DAC), pp. 43\u201348"},{"key":"1075_CR43","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1016\/j.sysarc.2006.10.006","volume":"53","author":"X Chen","year":"2007","unstructured":"Chen X, Maskell DL (2007) Supporting multiple-input, multiple-output custom functions in configurable processors. J Syst Architect 53:263\u2013271","journal-title":"J Syst Architect"},{"key":"1075_CR44","doi-asserted-by":"crossref","first-page":"373","DOI":"10.1016\/j.sysarc.2009.07.001","volume":"55","author":"ME Salehi","year":"2009","unstructured":"Salehi ME, Fakhraie SM (2009) Quantitative analysis of packet-processing applications regarding architectural guidelines for network-processing-engine development. J Syst Architect 55:373\u2013386","journal-title":"J Syst Architect"},{"key":"1075_CR45","doi-asserted-by":"crossref","first-page":"112","DOI":"10.1016\/j.sysarc.2012.02.004","volume":"58","author":"ME Salehi","year":"2012","unstructured":"Salehi ME, Fakhraie SM, Yazdanbakhsh A (2012) Instruction set architectural guidelines for embedded packet-processing engines. J Syst Architect 58:112\u2013125","journal-title":"J Syst Architect"},{"key":"1075_CR46","unstructured":"The GNU operating system, available online: http:\/\/www.gnu.org"},{"key":"1075_CR47","doi-asserted-by":"crossref","unstructured":"Yazdanbakhsh A, Salehi ME, Fakhraie SM (2010) Architecture-aware graph-covering algorithm for custom instruction selection. In: Proceedings of the international conference on future information technology (FutureTech), pp. 1\u20136","DOI":"10.1109\/FUTURETECH.2010.5482719"},{"key":"1075_CR48","doi-asserted-by":"crossref","unstructured":"Yazdanbakhsh A, Salehi ME, Fakhraie SM (2010) Locality considerations in exploring custom instruction selection algorithms. In: Proceedings of the ASQED","DOI":"10.1109\/ASQED.2010.5548232"},{"key":"1075_CR49","doi-asserted-by":"crossref","unstructured":"Yazdanbakhsh A, Kamal M, Salehi ME, Noori H, Fakhraie SM (2010) Energy-aware design space exploration of registerfile for extensible processors. In: Proceedings of the SAMOS","DOI":"10.1109\/ICSAMOS.2010.5642055"},{"key":"1075_CR50","doi-asserted-by":"crossref","first-page":"313","DOI":"10.1016\/S0166-218X(02)00205-6","volume":"126","author":"S Sakai","year":"2003","unstructured":"Sakai S, Togasaki M, Yamazaki K (2003) A note on greedy algorithms for the maximum weighted independent set problem. Discret Appl Math 126:313\u2013322","journal-title":"Discret Appl Math"},{"issue":"10\u2014-12","key":"1075_CR51","doi-asserted-by":"crossref","first-page":"421","DOI":"10.1016\/j.sysarc.2009.09.001","volume":"55","author":"R Ramaswamy","year":"2009","unstructured":"Ramaswamy R, Weng N, Wolf T (2009) Analysis of network processing workloads. J Syst Architect 55(10\u2014-12):421\u2013433","journal-title":"J Syst Architect"},{"key":"1075_CR52","doi-asserted-by":"crossref","unstructured":"Biswas P, Atasu K, Choudhary V, Pozzi L, Dutt N, Ienne P (2004) Introduction of local memory elements in instruction set extensions. In: Proceedings of the 41st design automation conference, June 2004, pp. 729\u2013734","DOI":"10.1145\/996566.996765"},{"key":"1075_CR53","doi-asserted-by":"crossref","unstructured":"She D, He Y, Corporaal H (2012) Energy efficient special instruction support in an embedded processor with compact ISA. In: proceedings of the CASES, pp. 131\u2013140","DOI":"10.1145\/2380403.2380430"},{"key":"1075_CR54","doi-asserted-by":"crossref","unstructured":"Wu D, Ahn J, Lee I, Choi K (2012) Resource-shared custom instruction generation under performance\/area constraints. International symposium on system on chip (SoC), pp. 1\u20136","DOI":"10.1109\/ISSoC.2012.6376353"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-013-1075-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11227-013-1075-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-013-1075-8","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,8,7]],"date-time":"2019-08-07T05:57:18Z","timestamp":1565157438000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11227-013-1075-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,2,6]]},"references-count":54,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2014,5]]}},"alternative-id":["1075"],"URL":"https:\/\/doi.org\/10.1007\/s11227-013-1075-8","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,2,6]]}}}