{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:28:28Z","timestamp":1750307308211,"version":"3.41.0"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2011,8,1]],"date-time":"2011-08-01T00:00:00Z","timestamp":1312156800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000149","name":"Division of Engineering Education and Centers","doi-asserted-by":"publisher","award":["EEC-0642422"],"award-info":[{"award-number":["EEC-0642422"]}],"id":[{"id":"10.13039\/100000149","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2011,8]]},"abstract":"<jats:p>Reconfigurable Computing (RC) systems based on FPGAs are becoming an increasingly attractive solution to building parallel systems of the future. Applications targeting such systems have demonstrated superior performance and reduced energy consumption versus their traditional counterparts based on microprocessors. However, most of such work has been limited to small system sizes. Unlike traditional HPC systems, lack of integrated, system-wide, parallel-programming models and languages presents a significant design challenge for creating applications targeting scalable, reconfigurable HPC systems. In this article, we extend the traditional Partitioned Global Address Space (PGAS) model to provide a multilevel integration of memory, which simplifies development of parallel applications for such systems and improves developer productivity. The new multilevel-PGAS programming model captures the unique characteristics of reconfigurable HPC systems, such as the existence of multiple levels of memory hierarchy and heterogeneous computation resources. Based on this model, we extend and adapt the SHMEM communication library to become what we call SHMEM+, the first known SHMEM library enabling coordination between FPGAs and CPUs in a reconfigurable, heterogeneous HPC system. Applications designed with SHMEM+ yield improved developer productivity compared to current methods of multidevice RC design and exhibit a high degree of portability. In addition, our design of SHMEM+ library itself is portable and provides peak communication bandwidth comparable to vendor-proprietary versions of SHMEM. Application case studies are presented to illustrate the advantages of SHMEM+.<\/jats:p>","DOI":"10.1145\/2000832.2000838","type":"journal-article","created":{"date-parts":[[2011,8,30]],"date-time":"2011-08-30T13:30:18Z","timestamp":1314711018000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["SHMEM+"],"prefix":"10.1145","volume":"4","author":[{"given":"Vikas","family":"Aggarwal","sequence":"first","affiliation":[{"name":"University of Florida, Gainesville, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alan D.","family":"George","sequence":"additional","affiliation":[{"name":"University of Florida, Gainesville, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Changil","family":"Yoon","sequence":"additional","affiliation":[{"name":"University of Florida, Gainesville, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kishore","family":"Yalamanchili","sequence":"additional","affiliation":[{"name":"University of Florida, Gainesville, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Herman","family":"Lam","sequence":"additional","affiliation":[{"name":"University of Florida, Gainesville, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2011,8,22]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1646461.1646464"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1646461.1646467"},{"key":"e_1_2_1_3_1","volume-title":"Workshop on asynchrony in the PGAS programming model. http:\/\/research.ihost.com\/apgas09\/.","author":"APGAS.","year":"2009","unstructured":"APGAS. 2009 . Workshop on asynchrony in the PGAS programming model. http:\/\/research.ihost.com\/apgas09\/. APGAS. 2009. Workshop on asynchrony in the PGAS programming model. http:\/\/research.ihost.com\/apgas09\/."},{"key":"e_1_2_1_4_1","unstructured":"Bonachea D. and Jeong J. Spring 2002. GASNet: A portable high-performance communication layer for global address-space languages. CS258 Parallel Computer Architecture Project.  Bonachea D. and Jeong J. Spring 2002. GASNet: A portable high-performance communication layer for global address-space languages. CS258 Parallel Computer Architecture Project."},{"volume-title":"The Fast Fourier Transform and its Application","author":"Brigham E. O.","key":"e_1_2_1_5_1","unstructured":"Brigham , E. O. 1988. The Fast Fourier Transform and its Application . Prentice Hall . Brigham, E. O. 1988. The Fast Fourier Transform and its Application. Prentice Hall."},{"key":"e_1_2_1_6_1","unstructured":"Carlson W. W. Draper J. M. Culler D. E. Yelick K. Brooks E. and Warren K. 1999. Introduction to UPC and language specification. Tech. rep. University of California-Berkeley Berkeley CA.  Carlson W. W. Draper J. M. Culler D. E. Yelick K. Brooks E. and Warren K. 1999. Introduction to UPC and language specification. Tech. rep. University of California-Berkeley Berkeley CA."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1090\/S0025-5718-1965-0178586-1"},{"key":"e_1_2_1_8_1","unstructured":"Cray T3ETM Fortran Optimization Guide - 004-2518-002. 2011. SHMEM. http:\/\/docs.cray.com\/books\/004-2518-002\/html-004-2518-002\/z826920364dep.html.  Cray T3ETM Fortran Optimization Guide - 004-2518-002. 2011. SHMEM. http:\/\/docs.cray.com\/books\/004-2518-002\/html-004-2518-002\/z826920364dep.html."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/648138.746808"},{"volume-title":"Proceedings of the Reconfigurable Systems Summer Institute.","author":"El-Ghazawi T.","key":"e_1_2_1_10_1","unstructured":"El-Ghazawi , T. , Serres , O. , Bahra , S. , Huang , M. , and El-Araby , E . 2008. Parallel programming of high-performance reconfigurable computing systems with Unified Parallel C . In Proceedings of the Reconfigurable Systems Summer Institute. El-Ghazawi, T., Serres, O., Bahra, S., Huang, M., and El-Araby, E. 2008. Parallel programming of high-performance reconfigurable computing systems with Unified Parallel C. In Proceedings of the Reconfigurable Systems Summer Institute."},{"key":"e_1_2_1_11_1","unstructured":"El-Ghazawi T. A. Carlson W. W. and Draper J. M. 2001. UPC language specifications v1.0. http:\/\/upc.gwu.edu\/docs\/upc_spec_1.1.1.pdf.  El-Ghazawi T. A. Carlson W. W. and Draper J. M. 2001. UPC language specifications v1.0. http:\/\/upc.gwu.edu\/docs\/upc_spec_1.1.1.pdf."},{"volume-title":"Proceedings of the Workshop on Asynchrony in the PGAS Programming Model.","author":"Farreras M.","key":"e_1_2_1_12_1","unstructured":"Farreras , M. , Marjanovic , V. , Ayguade , E. , and Labarta , J . 1997. Gaining asynchrony by using hybrid UPC\/SMPSs . In Proceedings of the Workshop on Asynchrony in the PGAS Programming Model. Farreras, M., Marjanovic, V., Ayguade, E., and Labarta, J. 1997. Gaining asynchrony by using hybrid UPC\/SMPSs. In Proceedings of the Workshop on Asynchrony in the PGAS Programming Model."},{"key":"e_1_2_1_13_1","unstructured":"Gonzales R. and Woods R. E. 2002. Digital Image Processing. Addison-Wesley.  Gonzales R. and Woods R. E. 2002. Digital Image Processing. Addison-Wesley."},{"key":"e_1_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Guyon I. Gunn S. Nikravesh M. and Zadeh L. 2006. Feature Extraction Foundations and Applications. Springer.   Guyon I. Gunn S. Nikravesh M. and Zadeh L. 2006. Feature Extraction Foundations and Applications. Springer.","DOI":"10.1007\/978-3-540-35488-8"},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 762--768","author":"Huang J.","key":"e_1_2_1_15_1","unstructured":"Huang , J. , Kumar , S. , Mitra , M. , Zhu , W.-J. , and Zabih , R . 1997. Image indexing using color correlograms . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 762--768 . Huang, J., Kumar, S., Mitra, M., Zhu, W.-J., and Zabih, R. 1997. Image indexing using color correlograms. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 762--768."},{"key":"e_1_2_1_16_1","unstructured":"MPI. 2011. MPI standard. http:\/\/www.mcs.anl.gov\/research\/projects\/mpi\/.  MPI. 2011. MPI standard. http:\/\/www.mcs.anl.gov\/research\/projects\/mpi\/."},{"key":"e_1_2_1_17_1","volume-title":"MVAPICH: MPI over InfiniBand and iWARP","author":"Network-Based Computing","year":"2011","unstructured":"Network-Based Computing Laboratory. 2011 . MVAPICH: MPI over InfiniBand and iWARP . http:\/\/mvapich.cse.ohio-state.edu. Network-Based Computing Laboratory. 2011. MVAPICH: MPI over InfiniBand and iWARP. http:\/\/mvapich.cse.ohio-state.edu."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5161076"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/289918.289920"},{"key":"e_1_2_1_20_1","unstructured":"OpenMP. 2011. The OpenMP API specification for parallel programming. http:\/\/openmp.org\/wp\/.  OpenMP. 2011. The OpenMP API specification for parallel programming. http:\/\/openmp.org\/wp\/."},{"volume-title":"Proceedings of the 3rd International Workshop on High-Performance Reconfigurable Computing Technology and Applications (HPRCTA'08)","author":"Saldana M.","key":"e_1_2_1_21_1","unstructured":"Saldana , M. , Patel , A. , Madill , C. , Nunes , D. , Danyao , W. , Styles , H. , Putnam , A. , Wittig , R. , and Chow , P . 2008. MPI as an abstraction for software-hardware interaction for HPRCs . In Proceedings of the 3rd International Workshop on High-Performance Reconfigurable Computing Technology and Applications (HPRCTA'08) . ACM, New York. Saldana, M., Patel, A., Madill, C., Nunes, D., Danyao, W., Styles, H., Putnam, A., Wittig, R., and Chow, P. 2008. MPI as an abstraction for software-hardware interaction for HPRCs. In Proceedings of the 3rd International Workshop on High-Performance Reconfigurable Computing Technology and Applications (HPRCTA'08). ACM, New York."},{"key":"e_1_2_1_22_1","unstructured":"SGI. 2011. Introduction to the SHMEM programming model. http:\/\/docs.sgi.com\/library\/tpl\/cgi-bin\/getdoc.cgi?coll=linux&db=man&fname=\/usr\/share\/catman\/man3\/intro_shmem.3.html&srch=intro_shmem.  SGI. 2011. Introduction to the SHMEM programming model. http:\/\/docs.sgi.com\/library\/tpl\/cgi-bin\/getdoc.cgi?coll=linux&db=man&fname=\/usr\/share\/catman\/man3\/intro_shmem.3.html&srch=intro_shmem."},{"volume-title":"Proceedings of the International Conference on Engineering of Reconfigurable Systems and Algorithms.","author":"Shih K.","key":"e_1_2_1_23_1","unstructured":"Shih , K. , Balachandran , A. , Nagarajan , K. , Holland , B. , Slatton , C. , and George , A . 2008. Fast real-time LIDAR processing on FPGAs . In Proceedings of the International Conference on Engineering of Reconfigurable Systems and Algorithms. Shih, K., Balachandran, A., Nagarajan, K., Holland, B., Slatton, C., and George, A. 2008. Fast real-time LIDAR processing on FPGAs. In Proceedings of the International Conference on Engineering of Reconfigurable Systems and Algorithms."},{"key":"e_1_2_1_24_1","doi-asserted-by":"crossref","unstructured":"Shirazi N. Athanas P. M. and Abbott A. L. 1995. Implementation of a 2-D fast Fourier transform on an FPGA-based custom computing machine. In Field Programmable Logic and Application. Springer Berlin 282--292.   Shirazi N. Athanas P. M. and Abbott A. L. 1995. Implementation of a 2-D fast Fourier transform on an FPGA-based custom computing machine. In Field Programmable Logic and Application. Springer Berlin 282--292.","DOI":"10.1007\/3-540-60294-1_122"},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Skarpathiotis C. and Dimond K. 2004. A hardware implementation of a content based image retrieval algorithm. In Field Programmable Logic and Application. Springer Berlin 1165--1167.  Skarpathiotis C. and Dimond K. 2004. A hardware implementation of a content based image retrieval algorithm. In Field Programmable Logic and Application. Springer Berlin 1165--1167.","DOI":"10.1007\/978-3-540-30117-2_157"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the Many-Core and Reconfigurable Supercomputing Conference (MRSC).","author":"Storaasli O.","year":"2008","unstructured":"Storaasli , O. 2008 . Accelerating senome sequencing 100-1000X with FPGAs . In Proceedings of the Many-Core and Reconfigurable Supercomputing Conference (MRSC). Storaasli, O. 2008. Accelerating senome sequencing 100-1000X with FPGAs. In Proceedings of the Many-Core and Reconfigurable Supercomputing Conference (MRSC)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/1058426.1058880"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1049\/ip-vis:20041114"},{"volume-title":"Proceedings of the ACM Workshop on Java for High-Performance Network Computing. ACM Press","author":"Yelick K.","key":"e_1_2_1_29_1","unstructured":"Yelick , K. , Semenzato , L. , Pike , G. , Miyamoto , C. , Liblit , B. , Krishnamurthy , A. , Hilfinger , P. , Graham , S. , Gay , D. , Colella , P. , and Aiken , A . 1998. Titanium: A high-performance Java dialect . In Proceedings of the ACM Workshop on Java for High-Performance Network Computing. ACM Press , New York. Yelick, K., Semenzato, L., Pike, G., Miyamoto, C., Liblit, B., Krishnamurthy, A., Hilfinger, P., Graham, S., Gay, D., Colella, P., and Aiken, A. 1998. Titanium: A high-performance Java dialect. In Proceedings of the ACM Workshop on Java for High-Performance Network Computing. ACM Press, New York."}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2000832.2000838","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2000832.2000838","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T11:00:03Z","timestamp":1750244403000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2000832.2000838"}},"subtitle":["A multilevel-PGAS programming model for reconfigurable supercomputing"],"short-title":[],"issued":{"date-parts":[[2011,8]]},"references-count":29,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2011,8]]}},"alternative-id":["10.1145\/2000832.2000838"],"URL":"https:\/\/doi.org\/10.1145\/2000832.2000838","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"type":"print","value":"1936-7406"},{"type":"electronic","value":"1936-7414"}],"subject":[],"published":{"date-parts":[[2011,8]]},"assertion":[{"value":"2010-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2010-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-08-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}