{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T15:26:14Z","timestamp":1759332374159,"version":"3.41.0"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2012,7,1]],"date-time":"2012-07-01T00:00:00Z","timestamp":1341100800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2012,7]]},"abstract":"<jats:p>In the field of high-performance computing, systems harboring reconfigurable devices, such as field-programmable gate arrays (FPGAs), are gaining more widespread interest. Such systems range from supercomputers with tightly coupled reconfigurable hardware to clusters with reconfigurable devices at each node. The use of these architectures for scientific computing provides an alternative for computationally demanding problems and has advantages in metrics, such as operating cost\/performance and power\/performance. However, performance optimization of these systems can be challenging even with knowledge of the system\u2019s characteristics. Our analytic performance model includes parameters representing the reconfigurable hardware, application load imbalance across the nodes, background user load, basic message-passing communication, and processor heterogeneity. In this article, we provide an overview of the analytical model and demonstrate its application for optimization and scheduling of high-performance reconfigurable computing (HPRC) resources. We examine cost functions for minimum runtime and other optimization problems commonly found in shared computing resources. Finally, we discuss additional scheduling issues and other potential applications of the model.<\/jats:p>","DOI":"10.1145\/2220336.2220348","type":"journal-article","created":{"date-parts":[[2012,7,31]],"date-time":"2012-07-31T13:42:45Z","timestamp":1343742165000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Optimization of Shared High-Performance Reconfigurable Computing Resources"],"prefix":"10.1145","volume":"11","author":[{"given":"Melissa C.","family":"Smith","sequence":"first","affiliation":[{"name":"Clemson University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gregory D.","family":"Peterson","sequence":"additional","affiliation":[{"name":"University of Tennessee"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,7]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Alpha Data. 2012. http:\/\/www.alpha-data.com. Alpha Data. 2012. http:\/\/www.alpha-data.com."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/0743-7315(92)90015-F"},{"volume-title":"Proceedings of the 9th SIAM Conference on Parallel Processing for Scientific Computing.","author":"Basney J.","key":"e_1_2_1_3_1","unstructured":"Basney , J. , Raman , B. , and Livny , M . 1999. High throughput Monte Carlo . In Proceedings of the 9th SIAM Conference on Parallel Processing for Scientific Computing. Basney, J., Raman, B., and Livny, M. 1999. High throughput Monte Carlo. In Proceedings of the 9th SIAM Conference on Parallel Processing for Scientific Computing."},{"key":"e_1_2_1_4_1","unstructured":"BLAS. 2012. Basic Linear Algebra Subprograms. http:\/\/www.netlib.org\/blas\/. BLAS . 2012. Basic Linear Algebra Subprograms. http:\/\/www.netlib.org\/blas\/."},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the 3rd Annual Conference on Genetic Programming.","author":"Cantu-Paz E.","year":"1998","unstructured":"Cantu-Paz , E. 1998 . Designing efficient master-slave parallel genetic algorithms . In Proceedings of the 3rd Annual Conference on Genetic Programming. Cantu-Paz, E. 1998. Designing efficient master-slave parallel genetic algorithms. In Proceedings of the 3rd Annual Conference on Genetic Programming."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/32.4634"},{"key":"e_1_2_1_7_1","unstructured":"Celoxica. 2012. http:\/\/www.celoxica.com. Celoxica. 2012. http:\/\/www.celoxica.com."},{"key":"e_1_2_1_8_1","unstructured":"ClearSpeed. 2012. http:\/\/www.clearspeed.com. ClearSpeed. 2012. http:\/\/www.clearspeed.com."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/508352.508353"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/800157.805047"},{"key":"e_1_2_1_11_1","unstructured":"Cray. 2012. http:\/\/www.cray.com. Cray. 2012. http:\/\/www.cray.com."},{"volume-title":"Reconfigurable Processing Unit","key":"e_1_2_1_12_1","unstructured":"DRC. 2012. Reconfigurable Processing Unit , DRC Computer Corporation . http:\/\/www.drccomp.com. DRC. 2012. Reconfigurable Processing Unit, DRC Computer Corporation. http:\/\/www.drccomp.com."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2008.65"},{"key":"e_1_2_1_14_1","unstructured":"El-Rewini H. and Lewis T. G. 1994. Task Scheduling in Parallel Distributed Systems. Prentice Hall Upper Saddle River NJ. El-Rewini H. and Lewis T. G. 1994. Task Scheduling in Parallel Distributed Systems. Prentice Hall Upper Saddle River NJ."},{"key":"e_1_2_1_15_1","unstructured":"Garey M. R. and Johnson D. S. 1979. Computers and Intractability: A Guide to the Theory of NP-Completeness. W.H. Freeman and Company New York. Garey M. R. and Johnson D. S. 1979. Computers and Intractability: A Guide to the Theory of NP-Completeness. W.H. Freeman and Company New York."},{"volume-title":"Proceedings of the 10th International Parallel Processing Symposium (IPPS\u201996)","author":"Govindan V.","key":"e_1_2_1_16_1","unstructured":"Govindan , V. and Franklin , M. A . 1996. Application load imbalance on parallel processors . In Proceedings of the 10th International Parallel Processing Symposium (IPPS\u201996) . Govindan, V. and Franklin, M. A. 1996. Application load imbalance on parallel processors. In Proceedings of the 10th International Parallel Processing Symposium (IPPS\u201996)."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1462586.1462591"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.265940"},{"volume-title":"Introduction to Computer System Performance Evaluation","author":"Kant K.","key":"e_1_2_1_19_1","unstructured":"Kant , K. 1992. Introduction to Computer System Performance Evaluation . McGraw-Hill, Inc. , New York . Kant, K. 1992. Introduction to Computer System Performance Evaluation. McGraw-Hill, Inc., New York."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.220.4598.671"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2008.01.008"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/344588.344618"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/1058426.1058879"},{"key":"e_1_2_1_24_1","unstructured":"Maxwell at EPCC. 2012. http:\/\/www.epcc.ed.ac.uk\/facilities\/maxwell. Maxwell at EPCC. 2012. http:\/\/www.epcc.ed.ac.uk\/facilities\/maxwell."},{"key":"e_1_2_1_25_1","unstructured":"Nallatech. 2012. http:\/\/www.nallatech.com. Nallatech . 2012. http:\/\/www.nallatech.com."},{"key":"e_1_2_1_26_1","unstructured":"NIST. 2005. Guideline for implementing cryptography in the federal government. Tech. rep. NIST SP800-21. http:\/\/csrc.nist.gov\/publications. NIST. 2005. Guideline for implementing cryptography in the federal government. Tech. rep. NIST SP800-21. http:\/\/csrc.nist.gov\/publications."},{"key":"e_1_2_1_27_1","unstructured":"Novo-G. 2012. http:\/\/www.chrec.org\/facilities.html. Novo-G. 2012. http:\/\/www.chrec.org\/facilities.html."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2007.59"},{"key":"e_1_2_1_30_1","unstructured":"SGI RASC. 2012. http:\/\/www.sgi.com. SGI RASC. 2012. http:\/\/www.sgi.com."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.peva.2004.10.004"},{"key":"e_1_2_1_33_1","unstructured":"SRC MAPstation. 2012. http:\/\/www.srccomp.com. SRC MAPstation. 2012. http:\/\/www.srccomp.com."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.993206"},{"key":"e_1_2_1_35_1","unstructured":"XtremeData XD1000 Development System. 2012. http:\/\/www.xtremedata.com. XtremeData XD1000 Development System. 2012. http:\/\/www.xtremedata.com."}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2220336.2220348","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2220336.2220348","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T09:20:57Z","timestamp":1750238457000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2220336.2220348"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,7]]},"references-count":33,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2012,7]]}},"alternative-id":["10.1145\/2220336.2220348"],"URL":"https:\/\/doi.org\/10.1145\/2220336.2220348","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2012,7]]},"assertion":[{"value":"2009-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-07-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}