{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T14:54:43Z","timestamp":1777128883041,"version":"3.51.4"},"reference-count":27,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2015,1,29]],"date-time":"2015-01-29T00:00:00Z","timestamp":1422489600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2015,1,29]],"date-time":"2015-01-29T00:00:00Z","timestamp":1422489600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Hum. Cent. Comput. Inf. Sci."],"published-print":{"date-parts":[[2015,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The power-performance trade-off is one of the major considerations in micro-architecture design. Pipelined architecture has brought a radical change in the design to capitalize on the parallel operation of various functional blocks involved in the instruction execution process, which is widely used in all modern processors. Pipeline introduces the instruction level parallelism (ILP) because of the potential overlap of instructions, and it does have drawbacks in the form of hazards, which is a result of data dependencies and resource conflicts. To overcome these hazards, stalls were introduced, which are basically delayed execution of instructions to diffuse the problematic situation. Out-of-order (OOO) execution is a ramification of the stall approach since it executes the instruction in an order governed by the availability of the input data rather than by their original order in the program. This paper presents a new algorithm called Left-Right (LR) for reducing stalls in pipelined processors. This algorithm is built by combining the traditional in-order and the out-of-order (OOO) instruction execution, resulting in the best of both approaches. As instruction input, we take the Tomasulo\u2019s algorithm for scheduling out-of-order and the in-order instruction execution and we compare the proposed algorithm\u2019s efficiency against both in terms of power-performance gain. Experimental simulations are conducted using Sim-Panalyzer, an instruction level simulator, showing that our proposed algorithm optimizes the power-performance with an effective increase of 30% in terms of energy consumption benefits compared to the Tomasulo\u2019s algorithm and 3% compared to the in-order algorithm.<\/jats:p>","DOI":"10.1186\/s13673-014-0016-8","type":"journal-article","created":{"date-parts":[[2015,1,28]],"date-time":"2015-01-28T12:24:48Z","timestamp":1422447888000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":26,"title":["An optimizing pipeline stall reduction algorithm for power and performance on multi-core CPUs"],"prefix":"10.1186","volume":"5","author":[{"given":"Vijayalakshmi","family":"Saravanan","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kothari Dwarkadas","family":"Pralhaddas","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dwarkadas Pralhaddas","family":"Kothari","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Isaac","family":"Woungang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2015,1,29]]},"reference":[{"key":"16_CR1","unstructured":"Kogge PM (1981) The Architecture of pipelined computers. McGraw-Hill advanced computer science series, Hemisphere, Washington, New York, Paris, Includes index."},{"key":"16_CR2","unstructured":"Johnson WM (1989) Super-scalar processor design. Technical report."},{"issue":"2","key":"16_CR3","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1145\/545214.545217","volume":"30","author":"A Hartstein","year":"2002","unstructured":"Hartstein A, Puzak TR (2002) The optimum pipeline depth for a microprocessor. SIGARCH Comput Archit News 30(2): 7\u201313.","journal-title":"SIGARCH Comput Archit News"},{"issue":"4","key":"16_CR4","doi-asserted-by":"publisher","first-page":"673","DOI":"10.1145\/1109118.1109124","volume":"10","author":"S Shamshiri","year":"2005","unstructured":"Shamshiri S, Esmaeilzadeh H, Navabi Z (2005) Instruction-level test methodology for cpu core self-testing. ACM Trans Des Autom Electron Syst 10(4): 673\u2013689.","journal-title":"ACM Trans Des Autom Electron Syst"},{"key":"16_CR5","unstructured":"Patterson DA, Hennessy JL (2006) In praise of computer architecture: a quantitative approach. Number 704. Morgan Kaufmann."},{"key":"16_CR6","doi-asserted-by":"publisher","first-page":"237","DOI":"10.1109\/MICRO.2010.45","volume-title":"Proceedings of the 2010 43rd Annual IEEE\/ACM International Symposium on Microarchitecture","author":"M Steffen","year":"2010","unstructured":"Steffen M, Zambreno J (2010) Improving simt efficiency of global rendering algorithms with architectural support for dynamic micro-kernels In: Proceedings of the 2010 43rd Annual IEEE\/ACM International Symposium on Microarchitecture, 237\u2013248.. MICRO \u201943, IEEE Computer Society, Washington, DC, USA."},{"key":"16_CR7","first-page":"399","volume-title":"Proceedings of the 2012 20th Euromicro International Conference on Parallel, Distributed and Network-based Processing PDP \u201912","author":"S Frey","year":"2012","unstructured":"Frey S, Reina G, Ertl T (2012) Simt microscheduling: Reducing thread stalling in divergent iterative algorithms In: Proceedings of the 2012 20th Euromicro International Conference on Parallel, Distributed and Network-based Processing PDP \u201912, 399\u2013406.. IEEE Computer Society, Washington, DC, USA."},{"key":"16_CR8","doi-asserted-by":"crossref","unstructured":"Han TD, Abdelrahman TS (2011) Reducing branch divergence in gpu programs In: Proceedings of the Fourth Workshop on General Purpose Processing on Graphics Processing Units, 3:1\u20133:8, GPGPU-4, ACM, New York, NY, USA.","DOI":"10.1145\/1964179.1964184"},{"key":"16_CR9","unstructured":"Lawrence R (1998) A survey of cache coherence mechanisms in shared memory multiprocessors."},{"issue":"12","key":"16_CR10","first-page":"47","volume":"12","author":"MK Chaudhary","year":"2011","unstructured":"Chaudhary MK, Kumar M, Rai M, Dwivedi RK (2011) Article: A Modified Algorithm for Buffer Cache Management. Int J Comput Appl 12(12): 47\u201349.","journal-title":"Int J Comput Appl"},{"key":"16_CR11","unstructured":"Bennett JE, Flynn MJ (1996) Reducing Cache Miss Rates Using Prediction Caches. Technical report."},{"key":"16_CR12","unstructured":"Schnberg S, Mehnert F, Hamann C-J, Hamann Clj, Reuther L, Hrtig H (1998) Performance and Bus Transfer Influences In: In First Workshop on PC-Based Syatem Performance and Analysis."},{"key":"16_CR13","doi-asserted-by":"crossref","unstructured":"Bahar RI, Albera G, Manne S (1998) Power and performance tradeoffs using various caching strategies In: Proceedings of the 1998 international symposium on Low power electronics and design, 64\u201369, ISLPED \u201998, ACM, New York, NY, USA.","DOI":"10.1145\/280756.295115"},{"key":"16_CR14","doi-asserted-by":"crossref","unstructured":"Jeon HS, Noh SH (1998) A database disk buffer management algorithm based on prefetching In: Proceedings of the seventh international conference on Information and knowledge management, 167\u2013174, CIKM \u201998, ACM, New York, NY, USA.","DOI":"10.1145\/288627.288654"},{"key":"16_CR15","first-page":"439","volume-title":"Proceedings of the 20th International Conference on Very Large Data Bases","author":"T Johnson","year":"1994","unstructured":"Johnson T, Shasha D (1994) 2Q: A Low Overhead High Performance Buffer Management Replacement Algorithm In: Proceedings of the 20th International Conference on Very Large Data Bases, 439\u2013450.. VLDB \u201994, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA."},{"issue":"4","key":"16_CR16","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1109\/92.645068","volume":"5","author":"RS Bajwa","year":"1997","unstructured":"Bajwa RS, Hiraki M, Kojima H, Gorny DJ, Nitta K, Shridhar A, Seki K, Sasaki K (1997) Instruction buffering to reduce power in processors for signal processing. IEEE Trans. Very Large Scale Integr Syst 5(4): 417\u2013424.","journal-title":"IEEE Trans. Very Large Scale Integr Syst"},{"issue":"5","key":"16_CR17","doi-asserted-by":"publisher","first-page":"38","DOI":"10.1145\/780822.781137","volume":"38","author":"C-H Hsu","year":"2003","unstructured":"Hsu C-H, Kremer U (2003) The design, implementation, and evaluation of a compiler algorithm for cpu energy reduction. SIGPLAN Not 38(5): 38\u201348.","journal-title":"SIGPLAN Not"},{"key":"16_CR18","unstructured":"Marculescu D (2000) On the Use of Microarchitecture-Driven Dynamic Voltage Scaling."},{"key":"16_CR19","first-page":"132","volume-title":"Proceedings of the 25th annual international symposium on Computer architecture","author":"S Manne","year":"1998","unstructured":"Manne S, Klauser A, Grunwald D (1998) Pipeline gating: speculation control for energy reduction In: Proceedings of the 25th annual international symposium on Computer architecture, 132\u2013141.. ISCA \u201998, IEEE Computer Society, Washington, DC, USA."},{"issue":"1","key":"16_CR20","doi-asserted-by":"publisher","first-page":"24","DOI":"10.1145\/1044111.1044114","volume":"10","author":"S-J Ruan","year":"2005","unstructured":"Ruan S-J, Tsai K-L, Naroska E, Lai F (2005) Bipartitioning and encoding in low-power pipelined circuits. ACM Trans Des Autom Electron Syst 10(1): 24\u201332.","journal-title":"ACM Trans Des Autom Electron Syst"},{"key":"16_CR21","unstructured":"Lei H, Duchamp D (1997) An Analytical Approach to File Prefetching In: In Proceedings of the USENIX 1997 Annual Technical Conference, 275\u2013288."},{"key":"16_CR22","doi-asserted-by":"crossref","unstructured":"Li Y, Henkel Jrg, Jrghenkel Y (1998) A Framework for Estimating and Minimizing Energy Dissipation of Embedded HW\/SW Systems.","DOI":"10.1145\/277044.277097"},{"key":"16_CR23","doi-asserted-by":"publisher","first-page":"434","DOI":"10.1145\/1186822.1073211","volume-title":"ACM SIGGRAPH 2005 Papers","author":"S Woop","year":"2005","unstructured":"Woop S, Schmittler J, Slusallek P (2005) Rpu: a programmable ray processing unit for realtime ray tracing In: ACM SIGGRAPH 2005 Papers, 434\u2013444.. SIGGRAPH \u201905. ACM, New York, NY, USA."},{"key":"16_CR24","unstructured":"Johnson M, William M (1989) Super-Scalar Processor Design. Technical report."},{"key":"16_CR25","unstructured":"Whitham J (2013) Simple scalar\/ARM VirtualBox Appliance. Website. http:\/\/www.jwhitham.org\/simplescalar."},{"issue":"2","key":"16_CR26","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1145\/342001.339657","volume":"28","author":"D Brooks","year":"2000","unstructured":"Brooks D, Tiwari V, Martonosi M (2000) Wattch: a framework for architectural-level power analysis and optimizations. SIGARCH Comput Archit News 28(2): 83\u201394.","journal-title":"SIGARCH Comput Archit News"},{"issue":"4","key":"16_CR27","doi-asserted-by":"publisher","first-page":"8","DOI":"10.1109\/MM.2006.73","volume":"26","author":"AR Alameldeen","year":"2006","unstructured":"Alameldeen AR, Wood DA (2006) Ipc considered harmful for multiprocessor workloads. IEEE Micro 26(4): 8\u201317.","journal-title":"IEEE Micro"}],"updated-by":[{"DOI":"10.1186\/s13673-015-0026-1","type":"erratum","label":"Erratum","source":"publisher","updated":{"date-parts":[[2015,4,7]],"date-time":"2015-04-07T00:00:00Z","timestamp":1428364800000}}],"container-title":["Human-centric Computing and Information Sciences"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13673-014-0016-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13673-014-0016-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13673-014-0016-8","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13673-014-0016-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,7,30]],"date-time":"2021-07-30T07:10:15Z","timestamp":1627629015000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1186\/s13673-014-0016-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,1,29]]},"references-count":27,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2015,12]]}},"alternative-id":["16"],"URL":"https:\/\/doi.org\/10.1186\/s13673-014-0016-8","relation":{},"ISSN":["2192-1962"],"issn-type":[{"value":"2192-1962","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,1,29]]},"assertion":[{"value":"11 February 2014","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 August 2014","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 January 2015","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"2"}}