{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:54:51Z","timestamp":1750308891013,"version":"3.41.0"},"reference-count":30,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2004,6,1]],"date-time":"2004-06-01T00:00:00Z","timestamp":1086048000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2004,6]]},"abstract":"<jats:p>Fetch engine performance is a key topic in superscalar processors, since it limits the instruction-level parallelism that can be exploited by the execution core. In the search of high performance, the fetch engine has evolved toward more efficient designs, but its complexity has also increased.In this paper, we present the stream fetch engine, a novel architecture based on the execution of long streams of sequential instructions, taking maximum advantage of code layout optimizations. We describe our design in detail, showing that it achieves high fetch performance, while requiring less complexity than other state-of-the-art fetch architectures.<\/jats:p>","DOI":"10.1145\/1011528.1011532","type":"journal-article","created":{"date-parts":[[2005,8,1]],"date-time":"2005-08-01T17:31:42Z","timestamp":1122917502000},"page":"220-245","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["A low-complexity fetch architecture for high-performance superscalar processors"],"prefix":"10.1145","volume":"1","author":[{"given":"Oliverio J.","family":"Santana","sequence":"first","affiliation":[{"name":"Universitat Polit\u00e9cnica de Catalunya, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alex","family":"Ramirez","sequence":"additional","affiliation":[{"name":"Universitat Polit\u00e9cnica de Catalunya, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Josep L.","family":"Larriba-Pey","sequence":"additional","affiliation":[{"name":"Universitat Polit\u00e9cnica de Catalunya, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mateo","family":"Valero","sequence":"additional","affiliation":[{"name":"Universitat Polit\u00e9cnica de Catalunya, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2004,6]]},"reference":[{"volume-title":"Proceedings of the 27th International Symposium on Computer Architecture. 10","author":"Agarwal V.","key":"e_1_2_1_1_1","unstructured":"Agarwal , V. , Hrishikesh , M. S. , Keckler , S. W. , and Burger , D . 2000. Clock rate versus IPC: The end of the road for conventional microarchitectures . In Proceedings of the 27th International Symposium on Computer Architecture. 10 .1145\/339647.339691 Agarwal, V., Hrishikesh, M. S., Keckler, S. W., and Burger, D. 2000. Clock rate versus IPC: The end of the road for conventional microarchitectures. In Proceedings of the 27th International Symposium on Computer Architecture. 10.1145\/339647.339691"},{"volume-title":"Proceedings of the 26th International Symposium on Computer Architecture. 10","author":"Black B.","key":"e_1_2_1_2_1","unstructured":"Black , B. , Rychlik , B. , and Shen , J. P . 1999. The block-based trace cache . In Proceedings of the 26th International Symposium on Computer Architecture. 10 .1145\/300979.300996 Black, B., Rychlik, B., and Shen, J. P. 1999. The block-based trace cache. In Proceedings of the 26th International Symposium on Computer Architecture. 10.1145\/300979.300996"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the 21st International Symposium on Computer Architecture. 10","author":"Calder B.","year":"1919","unstructured":"Calder , B. and Grunwald , D . 1994. Fast & accurate instruction fetch and branch prediction . In Proceedings of the 21st International Symposium on Computer Architecture. 10 .1145\/ 1919 95.192011 Calder, B. and Grunwald, D. 1994. Fast & accurate instruction fetch and branch prediction. In Proceedings of the 21st International Symposium on Computer Architecture. 10.1145\/191995.192011"},{"volume-title":"Proceedings of the 22nd International Symposium on Computer Architecture. 10","author":"Conte T.","key":"e_1_2_1_4_1","unstructured":"Conte , T. , Menezes , K. , Mills , P. , and Patell , B . 1995. Optimization of instruction fetch mechanism for high issue rates . In Proceedings of the 22nd International Symposium on Computer Architecture. 10 .1145\/223982.224444 Conte, T., Menezes, K., Mills, P., and Patell, B. 1995. Optimization of instruction fetch mechanism for high issue rates. In Proceedings of the 22nd International Symposium on Computer Architecture. 10.1145\/223982.224444"},{"volume-title":"Proceedings of the 31st International Symposium on Microarchitecture.","author":"Driesen K.","key":"e_1_2_1_5_1","unstructured":"Driesen , K. and H\u00f6lzle , U . 1998. The cascaded predictor: economical and adaptive branch target prediction . In Proceedings of the 31st International Symposium on Microarchitecture. Driesen, K. and H\u00f6lzle, U. 1998. The cascaded predictor: economical and adaptive branch target prediction. In Proceedings of the 31st International Symposium on Microarchitecture."},{"volume-title":"Proceedings of the 30th International Symposium on Microarchitecture.","author":"Friendly D. H.","key":"e_1_2_1_6_1","unstructured":"Friendly , D. H. , Patel , S. J. , and Patt , Y. N . 1997. Alternative fetch and issue techniques from the trace cache mechanism . In Proceedings of the 30th International Symposium on Microarchitecture. Friendly, D. H., Patel, S. J., and Patt, Y. N. 1997. Alternative fetch and issue techniques from the trace cache mechanism. In Proceedings of the 30th International Symposium on Microarchitecture."},{"key":"e_1_2_1_7_1","unstructured":"Hinton G. Sager D. Upton M. Boggs D. Caerman D. Kyker A. and Roussel P. 2001. The microarchitecture of the Pentium 4 processor. Intel Technology Journal.  Hinton G. Sager D. Upton M. Boggs D. Caerman D. Kyker A. and Roussel P. 2001. The microarchitecture of the Pentium 4 processor. Intel Technology Journal."},{"volume-title":"Proceedings of the 29th International Symposium on Computer Architecture.","author":"Hrishikesh M. S.","key":"e_1_2_1_8_1","unstructured":"Hrishikesh , M. S. , Jouppi , N. P. , Farkas , K. I. , Burger , D. , Keckler , S. W. , and Shivakumar , P . 2002. The optimal useful logic depth per pipeline stage is 6--8 FO4 . In Proceedings of the 29th International Symposium on Computer Architecture. Hrishikesh, M. S., Jouppi, N. P., Farkas, K. I., Burger, D., Keckler, S. W., and Shivakumar, P. 2002. The optimal useful logic depth per pipeline stage is 6--8 FO4. In Proceedings of the 29th International Symposium on Computer Architecture."},{"volume-title":"Proceedings of the 30th International Symposium on Microarchitecture.","author":"Jacobson Q.","key":"e_1_2_1_9_1","unstructured":"Jacobson , Q. , Rotenberg , E. , and Smith , J. E . 1997. Path-based next trace prediction . In Proceedings of the 30th International Symposium on Microarchitecture. Jacobson, Q., Rotenberg, E., and Smith, J. E. 1997. Path-based next trace prediction. In Proceedings of the 30th International Symposium on Microarchitecture."},{"volume-title":"Proceedings of the 33rd International Symposium on Microarchitecture. 10","author":"Jimenez D. A.","key":"e_1_2_1_11_1","unstructured":"Jimenez , D. A. , Keckler , S. W. , and Lin , C . 2000. The impact of delay on the design of branch predictors . In Proceedings of the 33rd International Symposium on Microarchitecture. 10 .1145\/360128.360137 Jimenez, D. A., Keckler, S. W., and Lin, C. 2000. The impact of delay on the design of branch predictors. In Proceedings of the 33rd International Symposium on Microarchitecture. 10.1145\/360128.360137"},{"volume-title":"Proceedings of the 7th International Conference on High Performance Computer Architecture.","author":"Jimenez D. A.","key":"e_1_2_1_12_1","unstructured":"Jimenez , D. A. and Lin , C . 2001. Dynamic branch prediction with perceptrons . In Proceedings of the 7th International Conference on High Performance Computer Architecture. Jimenez, D. A. and Lin, C. 2001. Dynamic branch prediction with perceptrons. In Proceedings of the 7th International Conference on High Performance Computer Architecture."},{"volume-title":"Proceedings of the 18th International Symposium on Computer Architecture. 10","author":"Kaeli D.","key":"e_1_2_1_13_1","unstructured":"Kaeli , D. and Emma , P . 1991. Branch history table prediction of moving target branches due to subroutine returns . In Proceedings of the 18th International Symposium on Computer Architecture. 10 .1145\/115952.115957 Kaeli, D. and Emma, P. 1991. Branch history table prediction of moving target branches due to subroutine returns. In Proceedings of the 18th International Symposium on Computer Architecture. 10.1145\/115952.115957"},{"key":"e_1_2_1_14_1","unstructured":"Nair R. 2001. Method and apparatus for prefetching superblocks in a computer processing system. U.S. Patent Number 6 304 962 B1.  Nair R. 2001. Method and apparatus for prefetching superblocks in a computer processing system. U.S. Patent Number 6 304 962 B1."},{"volume-title":"Proceedings of the 33rd International Symposium on Microarchitecture. 10","author":"Patel S. J.","key":"e_1_2_1_15_1","unstructured":"Patel , S. J. , Tung , T. , Bose , S. , and Crum , M. M . 2000. Increasing the size of atomic instruction blocks using control flow assertions . In Proceedings of the 33rd International Symposium on Microarchitecture. 10 .1145\/360128.360160 Patel, S. J., Tung, T., Bose, S., and Crum, M. M. 2000. Increasing the size of atomic instruction blocks using control flow assertions. In Proceedings of the 33rd International Symposium on Microarchitecture. 10.1145\/360128.360160"},{"key":"e_1_2_1_16_1","first-page":"381","article-title":"Dynamic flow instruction cache memory organized around trace segments independent of virtual address line","volume":"5","author":"Peleg A.","year":"1995","unstructured":"Peleg , A. and Weiser , U. 1995 . Dynamic flow instruction cache memory organized around trace segments independent of virtual address line . U.S. Patent Number 5 , 381 ,533. Peleg, A. and Weiser, U. 1995. Dynamic flow instruction cache memory organized around trace segments independent of virtual address line. U.S. Patent Number 5,381,533.","journal-title":"U.S. Patent Number"},{"volume-title":"Proceedings of the 28th International Symposium on Computer Architecture. 10","author":"Ramirez A.","key":"e_1_2_1_17_1","unstructured":"Ramirez , A. , Barroso , L. , Gharachorloo , K. , Cohn , R. , Larriba-Pey , J. L. , Lawney , G. , and Valero , M . 2001. Code layout optimizations for transaction processing workloads . In Proceedings of the 28th International Symposium on Computer Architecture. 10 .1145\/379240.379260 Ramirez, A., Barroso, L., Gharachorloo, K., Cohn, R., Larriba-Pey, J. L., Lawney, G., and Valero, M. 2001. Code layout optimizations for transaction processing workloads. In Proceedings of the 28th International Symposium on Computer Architecture. 10.1145\/379240.379260"},{"volume-title":"Proceedings of the 13th International Conference on Supercomputing. 10","author":"Ramirez A.","key":"e_1_2_1_18_1","unstructured":"Ramirez , A. , Larriba-Pey , J. L. , Navarro , C. , Torrellas , J. , and Valero , M . 1999. Software trace cache . In Proceedings of the 13th International Conference on Supercomputing. 10 .1145\/305138.305178 Ramirez, A., Larriba-Pey, J. L., Navarro, C., Torrellas, J., and Valero, M. 1999. Software trace cache. In Proceedings of the 13th International Conference on Supercomputing. 10.1145\/305138.305178"},{"volume-title":"Proceedings of the 6th International Conference on High Performance Computer Architecture.","author":"Ramirez A.","key":"e_1_2_1_19_1","unstructured":"Ramirez , A. , Larriba-Pey , J. L. , and Valero , M . 2000. Trace cache redundancy: Red & blue traces . In Proceedings of the 6th International Conference on High Performance Computer Architecture. Ramirez, A., Larriba-Pey, J. L., and Valero, M. 2000. Trace cache redundancy: Red & blue traces. In Proceedings of the 6th International Conference on High Performance Computer Architecture."},{"volume-title":"Proceedings of the 35th International Symposium on Microarchitecture.","author":"Ramirez A.","key":"e_1_2_1_20_1","unstructured":"Ramirez , A. , Santana , O. J. , Larriba-Pey , J. L. , and Valero , M . 2002. Fetching instruction streams . In Proceedings of the 35th International Symposium on Microarchitecture. Ramirez, A., Santana, O. J., Larriba-Pey, J. L., and Valero, M. 2002. Fetching instruction streams. In Proceedings of the 35th International Symposium on Microarchitecture."},{"volume-title":"Proceedings of the 26th International Symposium on Computer Architecture. 10","author":"Reinman G.","key":"e_1_2_1_21_1","unstructured":"Reinman , G. , Austin , T. , and Calder , B . 1999. A scalable front-end architecture for fast instruction delivery . In Proceedings of the 26th International Symposium on Computer Architecture. 10 .1145\/300979.300999 Reinman, G., Austin, T., and Calder, B. 1999. A scalable front-end architecture for fast instruction delivery. In Proceedings of the 26th International Symposium on Computer Architecture. 10.1145\/300979.300999"},{"volume-title":"Proceedings of the 10th International Conference on Parallel Architectures and Compilation Techniques.","author":"Rosner R.","key":"e_1_2_1_22_1","unstructured":"Rosner , R. , Mendelson , A. , and Ronen , R . 2001. Filtering techniques to improve trace cache efficiency . In Proceedings of the 10th International Conference on Parallel Architectures and Compilation Techniques. Rosner, R., Mendelson, A., and Ronen, R. 2001. Filtering techniques to improve trace cache efficiency. In Proceedings of the 10th International Conference on Parallel Architectures and Compilation Techniques."},{"volume-title":"Proceedings of the 29th International Symposium on Microarchitecture.","author":"Rotenberg E.","key":"e_1_2_1_23_1","unstructured":"Rotenberg , E. , Benett , S. , and Smith , J. E . 1996. Trace cache: a low latency approach to high bandwidth instruction fetching . In Proceedings of the 29th International Symposium on Microarchitecture. Rotenberg, E., Benett, S., and Smith, J. E. 1996. Trace cache: a low latency approach to high bandwidth instruction fetching. In Proceedings of the 29th International Symposium on Microarchitecture."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/12.752652"},{"volume-title":"Proceedings of the 4th International Symposium on High Performance Computing.","author":"Santana O. J.","key":"e_1_2_1_25_1","unstructured":"Santana , O. J. , Falcon , A. , Fernandez , E. , Medina , P. , Ramirez , A. , and Valero , M . 2002. A comprehensive analysis of indirect branch prediction . In Proceedings of the 4th International Symposium on High Performance Computing. Santana, O. J., Falcon, A., Fernandez, E., Medina, P., Ramirez, A., and Valero, M. 2002. A comprehensive analysis of indirect branch prediction. In Proceedings of the 4th International Symposium on High Performance Computing."},{"key":"e_1_2_1_26_1","unstructured":"Santana O. J. Falcon A. Ramirez A. Larriba-Pey J. L. and Valero M. 2002. Differences Between the Next Stream Predictor and the Apparatus For Prefetching Superblocks Described in U.S. Patent 6 304 962 B1. Tech. rep. DAC-UPC-2002-18 Universitat Polit\u00e8cnica de Catalunya.  Santana O. J. Falcon A. Ramirez A. Larriba-Pey J. L. and Valero M. 2002. Differences Between the Next Stream Predictor and the Apparatus For Prefetching Superblocks Described in U.S. Patent 6 304 962 B1. Tech. rep. DAC-UPC-2002-18 Universitat Polit\u00e8cnica de Catalunya."},{"volume-title":"Proceedings of the 29th International Symposium on Computer Architecture.","author":"Seznec A.","key":"e_1_2_1_27_1","unstructured":"Seznec , A. , Felix , S. , Krishnan , V. , and Sazeides , Y . 2002. Design tradeoffs for the alpha EV8 conditional branch predictor . In Proceedings of the 29th International Symposium on Computer Architecture. Seznec, A., Felix, S., Krishnan, V., and Sazeides, Y. 2002. Design tradeoffs for the alpha EV8 conditional branch predictor. In Proceedings of the 29th International Symposium on Computer Architecture."},{"volume-title":"Proceedings of the 10th International Conference on Parallel Architectures and Compilation Techniques.","author":"Sherwood T.","key":"e_1_2_1_28_1","unstructured":"Sherwood , T. , Perelman , E. , and Calder , B . 2001. Basic block distribution analysis to find periodic behavior and simulation points in applications . In Proceedings of the 10th International Conference on Parallel Architectures and Compilation Techniques. Sherwood, T., Perelman, E., and Calder, B. 2001. Basic block distribution analysis to find periodic behavior and simulation points in applications. In Proceedings of the 10th International Conference on Parallel Architectures and Compilation Techniques."},{"key":"e_1_2_1_29_1","volume-title":"Research Report 2001\/2, Western Research Laboratory.","author":"Shivakumar P.","year":"2001","unstructured":"Shivakumar , P. and Jouppi , N. P . 2001 . Cacti 3.0: An integrated cache timing, power and area model. Research Report 2001\/2, Western Research Laboratory. Shivakumar, P. and Jouppi, N. P. 2001. Cacti 3.0: An integrated cache timing, power and area model. Research Report 2001\/2, Western Research Laboratory."},{"volume-title":"Proceedings of the 7th International Conference on Supercomputing. 10","author":"Yeh T.-Y.","key":"e_1_2_1_30_1","unstructured":"Yeh , T.-Y. , Marr , D. T. , and Patt , Y. N . 1993. Increasing the instruction fetch rate via multiple branch prediction and a branch address cache . In Proceedings of the 7th International Conference on Supercomputing. 10 .1145\/165939.165956 Yeh, T.-Y., Marr, D. T., and Patt, Y. N. 1993. Increasing the instruction fetch rate via multiple branch prediction and a branch address cache. In Proceedings of the 7th International Conference on Supercomputing. 10.1145\/165939.165956"},{"volume-title":"Proceedings of the 25th International Symposium on Microarchitecture.","author":"Yeh T.-Y.","key":"e_1_2_1_31_1","unstructured":"Yeh , T.-Y. and Patt , Y. N . 1992. A comprehensive instruction fetch mechanism for a processor supporting speculative execution . In Proceedings of the 25th International Symposium on Microarchitecture. Yeh, T.-Y. and Patt, Y. N. 1992. A comprehensive instruction fetch mechanism for a processor supporting speculative execution. In Proceedings of the 25th International Symposium on Microarchitecture."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1011528.1011532","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1011528.1011532","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:25:42Z","timestamp":1750281942000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1011528.1011532"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,6]]},"references-count":30,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2004,6]]}},"alternative-id":["10.1145\/1011528.1011532"],"URL":"https:\/\/doi.org\/10.1145\/1011528.1011532","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2004,6]]},"assertion":[{"value":"2004-06-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}