{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:35:10Z","timestamp":1750307710367,"version":"3.41.0"},"reference-count":21,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2009,2,1]],"date-time":"2009-02-01T00:00:00Z","timestamp":1233446400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100006602","name":"Air Force Research Laboratory","doi-asserted-by":"publisher","award":["___amp___num;F30602-01-C-0171"],"award-info":[{"award-number":["___amp___num;F30602-01-C-0171"]}],"id":[{"id":"10.13039\/100006602","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["___amp___num;CCR-0000988"],"award-info":[{"award-number":["___amp___num;CCR-0000988"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000185","name":"Defense Advanced Research Projects Agency","doi-asserted-by":"publisher","award":["___amp___num;NBCH1050022","___amp___num;F30602-01-C-0171"],"award-info":[{"award-number":["___amp___num;NBCH1050022","___amp___num;F30602-01-C-0171"]}],"id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Comput. Syst."],"published-print":{"date-parts":[[2009,2]]},"abstract":"<jats:p>The key to high performance in Simultaneous MultiThreaded (SMT) processors lies in optimizing the distribution of shared resources to active threads. Existing resource distribution techniques optimize performance only indirectly. They infer potential performance bottlenecks by observing indicators, like instruction occupancy or cache miss counts, and take actions to try to alleviate them. While the corrective actions are designed to improve performance, their actual performance impact is not known since end performance is never monitored. Consequently, potential performance gains are lost whenever the corrective actions do not effectively address the actual bottlenecks occurring in the pipeline.<\/jats:p>\n          <jats:p>\n            We propose a different approach to SMT resource distribution that optimizes end performance directly. Our approach observes the impact that resource distribution decisions have on performance at runtime, and feeds this information back to the resource distribution mechanisms to improve future decisions. By evaluating many different resource distributions, our approach tries to\n            <jats:italic>learn<\/jats:italic>\n            the best distribution over time. Because we perform learning online, learning time is crucial. We develop a\n            <jats:italic>hill-climbing algorithm<\/jats:italic>\n            that quickly learns the best distribution of resources by following the performance gradient within the resource distribution space. We also develop several ideal learning algorithms to enable deeper insights through limit studies.\n          <\/jats:p>\n          <jats:p>This article conducts an in-depth investigation of hill-climbing SMT resource distribution using a comprehensive suite of 63 multiprogrammed workloads. Our results show hill-climbing outperforms ICOUNT, FLUSH, and DCRA (three existing SMT techniques) by 11.4%, 11.5%, and 2.8%, respectively, under the weighted IPC metric. A limit study conducted using our ideal learning algorithms shows our approach can potentially outperform the same techniques by 19.2%, 18.0%, and 7.6%, respectively, thus demonstrating additional room exists for further improvement. Using our ideal algorithms, we also identify three bottlenecks that limit online learning speed: local maxima, phased behavior, and interepoch jitter. We define metrics to quantify these learning bottlenecks, and characterize the extent to which they occur in our workloads. Finally, we conduct a sensitivity study, and investigate several extensions to improve our hill-climbing technique.<\/jats:p>","DOI":"10.1145\/1482619.1482620","type":"journal-article","created":{"date-parts":[[2009,2,10]],"date-time":"2009-02-10T16:42:19Z","timestamp":1234284139000},"page":"1-47","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Hill-climbing SMT processor resource distribution"],"prefix":"10.1145","volume":"27","author":[{"given":"Seungryul","family":"Choi","sequence":"first","affiliation":[{"name":"Google"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Donald","family":"Yeung","sequence":"additional","affiliation":[{"name":"University of Maryland, College Park, MD"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2009,2,13]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/268806.268810"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2004.17"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2006.25"},{"key":"e_1_2_1_4_1","first-page":"1","article-title":"Optimizing SMT processors for high single-thread performance","volume":"5","author":"Dorai G. K.","year":"2003","unstructured":"Dorai , G. K. , Yeung , D. , and Choi , S. 2003 . Optimizing SMT processors for high single-thread performance . J. Instruction-Level Parallel. 5 , 1 -- 35 . Dorai, G. K., Yeung, D., and Choi, S. 2003. Optimizing SMT processors for high single-thread performance. J. Instruction-Level Parallel. 5, 1--35.","journal-title":"J. Instruction-Level Parallel."},{"volume-title":"Proceedings of the 9th International Conference on High Performance Computer Architecture. IEEE Computer Society, 31--40","author":"El-Moursy A.","key":"e_1_2_1_5_1","unstructured":"El-Moursy , A. and Albonesi , D. H . 2003. Front-End policies for improved issue efficiency in SMT processors . In Proceedings of the 9th International Conference on High Performance Computer Architecture. IEEE Computer Society, 31--40 . El-Moursy, A. and Albonesi, D. H. 2003. Front-End policies for improved issue efficiency in SMT processors. In Proceedings of the 9th International Conference on High Performance Computer Architecture. IEEE Computer Society, 31--40."},{"volume-title":"Proceedings of the 13th Symposium on Computer Architecture and High Performance Computing.","author":"Goncalves R.","key":"e_1_2_1_6_1","unstructured":"Goncalves , R. , Ayguade , E. , Valero , M. , and Navau , P. O. A. 2001. Performance evaluation of decoding and dispatching stages in simultaneous multithreaded architectures . In Proceedings of the 13th Symposium on Computer Architecture and High Performance Computing. Goncalves, R., Ayguade, E., Valero, M., and Navau, P. O. A. 2001. Performance evaluation of decoding and dispatching stages in simultaneous multithreaded architectures. In Proceedings of the 13th Symposium on Computer Architecture and High Performance Computing."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2004.1289290"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1006209.1006254"},{"volume-title":"Proceedings of the International Parallel and Distributed Processing Symposium. IEEE Computer Society.","author":"Luo K.","key":"e_1_2_1_9_1","unstructured":"Luo , K. , Franklin , M. , Mukherjee , S. S. , and Seznec , A . 2001. Boosting SMT performance by speculation control . In Proceedings of the International Parallel and Distributed Processing Symposium. IEEE Computer Society. Luo, K., Franklin, M., Mukherjee, S. S., and Seznec, A. 2001. Boosting SMT performance by speculation control. In Proceedings of the International Parallel and Distributed Processing Symposium. IEEE Computer Society."},{"volume-title":"Proceedings of the International Symposium on Performance Analysis of Systems and Software. IEEE Computer Society, 164--171","author":"Luo K.","key":"e_1_2_1_10_1","unstructured":"Luo , K. , Gummaraju , J. , and Franklin , M . 2001. Balancing throughput and fairness in SMT processors . In Proceedings of the International Symposium on Performance Analysis of Systems and Software. IEEE Computer Society, 164--171 . Luo, K., Gummaraju, J., and Franklin, M. 2001. Balancing throughput and fairness in SMT processors. In Proceedings of the International Symposium on Performance Analysis of Systems and Software. IEEE Computer Society, 164--171."},{"volume-title":"Proceedings of EuroPar'99","author":"Madon D.","key":"e_1_2_1_11_1","unstructured":"Madon , D. , Sanchez , E. , and Monnier , S . 1999. A study of a simultaneous multithreaded processor implementation . In Proceedings of EuroPar'99 . Springer, 716--726. Madon, D., Sanchez, E., and Monnier, S. 1999. A study of a simultaneous multithreaded processor implementation. In Proceedings of EuroPar'99. Springer, 716--726."},{"key":"e_1_2_1_12_1","first-page":"4","article-title":"Hyper-Threading technology architecture and microarchitecture","volume":"6","author":"Marr D. T.","year":"2002","unstructured":"Marr , D. T. , Binns , F. , Hill , D. , Hinton , G. , Koufaty , D. , Miller , J. A. , and Upton , M. 2002 . Hyper-Threading technology architecture and microarchitecture . Intel Technol. J. 6 , 1, 4 -- 15 . Marr, D. T., Binns, F., Hill, D., Hinton, G., Koufaty, D., Miller, J. A., and Upton, M. 2002. Hyper-Threading technology architecture and microarchitecture. Intel Technol. J. 6, 1, 4--15.","journal-title":"Intel Technol. J."},{"key":"e_1_2_1_13_1","unstructured":"Pentium4. 2002. Intel Pentium 4 processor. http:\/\/www.intel.com\/design\/Pentium4\/index.htm.  Pentium4. 2002. Intel Pentium 4 processor. http:\/\/www.intel.com\/design\/Pentium4\/index.htm."},{"volume-title":"Proceedings of the Multithreaded Execution, Architecture, and Compilation Workshop.","author":"Raasch S. E.","key":"e_1_2_1_14_1","unstructured":"Raasch , S. E. and Reinhardt , S. K . 1999. Applications of thread prioritization in SMT processors . In Proceedings of the Multithreaded Execution, Architecture, and Compilation Workshop. Raasch, S. E. and Reinhardt, S. K. 1999. Applications of thread prioritization in SMT processors. In Proceedings of the Multithreaded Execution, Architecture, and Compilation Workshop."},{"volume-title":"Proceedings of the 12th International Conference on Parallel Architectures and Compilation Techniques. IEEE Computer Society, 15--25","author":"Raasch S. E.","key":"e_1_2_1_15_1","unstructured":"Raasch , S. E. and Reinhardt , S. K . 2003. The impact of resource partitioning on SMT processors . In Proceedings of the 12th International Conference on Parallel Architectures and Compilation Techniques. IEEE Computer Society, 15--25 . Raasch, S. E. and Reinhardt, S. K. 2003. The impact of resource partitioning on SMT processors. In Proceedings of the 12th International Conference on Parallel Architectures and Compilation Techniques. IEEE Computer Society, 15--25."},{"volume-title":"Proceedings of the 10th International Conference on Parallel Architectures and Compilation Techniques. IEEE Computer Society, 3--14","author":"Sherwood T.","key":"e_1_2_1_16_1","unstructured":"Sherwood , T. , Perelman , E. , and Calder , B . 2001. Basic block distribution analysis to find periodic behavior and simulation points in applications . In Proceedings of the 10th International Conference on Parallel Architectures and Compilation Techniques. IEEE Computer Society, 3--14 . Sherwood, T., Perelman, E., and Calder, B. 2001. Basic block distribution analysis to find periodic behavior and simulation points in applications. In Proceedings of the 10th International Conference on Parallel Architectures and Compilation Techniques. IEEE Computer Society, 3--14."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/605397.605403"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/859618.859657"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/511334.511343"},{"volume-title":"Proceedings of the 34th Annual ACM\/IEEE International Symposium on Microarchitecture. IEEE Computer Society, 318--327","author":"Tullsen D. M.","key":"e_1_2_1_20_1","unstructured":"Tullsen , D. M. and Brown , J. A . 2001. Handling long-latency loads in a simultaneous multithreading processor . In Proceedings of the 34th Annual ACM\/IEEE International Symposium on Microarchitecture. IEEE Computer Society, 318--327 . Tullsen, D. M. and Brown, J. A. 2001. Handling long-latency loads in a simultaneous multithreading processor. In Proceedings of the 34th Annual ACM\/IEEE International Symposium on Microarchitecture. IEEE Computer Society, 318--327."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/232973.232993"}],"container-title":["ACM Transactions on Computer Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1482619.1482620","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1482619.1482620","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T13:30:10Z","timestamp":1750253410000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1482619.1482620"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,2]]},"references-count":21,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2009,2]]}},"alternative-id":["10.1145\/1482619.1482620"],"URL":"https:\/\/doi.org\/10.1145\/1482619.1482620","relation":{},"ISSN":["0734-2071","1557-7333"],"issn-type":[{"type":"print","value":"0734-2071"},{"type":"electronic","value":"1557-7333"}],"subject":[],"published":{"date-parts":[[2009,2]]},"assertion":[{"value":"2007-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2008-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-02-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}