{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T04:35:22Z","timestamp":1783571722613,"version":"3.55.0"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2015,1,21]],"date-time":"2015-01-21T00:00:00Z","timestamp":1421798400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Science Foundation","award":["CCF-0916652 and CCF-1055094 (CAREER)"],"award-info":[{"award-number":["CCF-0916652 and CCF-1055094 (CAREER)"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2015,1,21]]},"abstract":"<jats:p>Recent industry trends show a drastic rise in the use of hand-held embedded devices, from everyday applications to medical (e.g., monitoring devices) and critical defense applications (e.g., sensor nodes). The two key requirements in the design of such devices are their processing capabilities and battery life. There is therefore an urgency to build high-performance and power-efficient embedded devices, inspiring researchers to develop novel system designs for the same. The use of a coprocessor (application-specific hardware) to offload power-hungry computations is gaining favor among system designers to suit their power budgets. We propose the use of CGRAs (Coarse-Grained Reconfigurable Arrays) as a power-efficient coprocessor. Though CGRAs have been widely used for streaming applications, the extensive compiler support required limits its applicability and use as a general purpose coprocessor. In addition, a CGRA structure can efficiently execute only one statically scheduled kernel at a time, which is a serious limitation when used as an accelerator to a multithreaded or multitasking processor. In this work, we envision a multithreaded CGRA where multiple schedules (or kernels) can be executed simultaneously on the CGRA (as a coprocessor). We propose a comprehensive software scheme that transforms the traditionally single-threaded CGRA into a multithreaded coprocessor to be used as a power-efficient accelerator for multithreaded embedded processors. Our software scheme includes (1) a compiler framework that integrates with existing CGRA mapping techniques to prepare kernels for execution on the multithreaded CGRA and (2) a runtime mechanism that dynamically schedules multiple kernels (offloaded from the processor) to execute simultaneously on the CGRA coprocessor. Our multithreaded CGRA coprocessor implementation thus makes it possible to achieve improved power-efficient computing in modern multithreaded embedded systems.<\/jats:p>","DOI":"10.1145\/2638558","type":"journal-article","created":{"date-parts":[[2015,1,28]],"date-time":"2015-01-28T14:05:51Z","timestamp":1422453951000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["A Software Scheme for Multithreading on CGRAs"],"prefix":"10.1145","volume":"14","author":[{"given":"Jared","family":"Pager","sequence":"first","affiliation":[{"name":"Compiler Microarchitecture Lab, Arizona State University, Arizona"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Reiley","family":"Jeyapaul","sequence":"additional","affiliation":[{"name":"Compiler Microarchitecture Lab, Arizona State University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Aviral","family":"Shrivastava","sequence":"additional","affiliation":[{"name":"Compiler Microarchitecture Lab, Arizona State University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2015,1,21]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"ARM-A9. 2009. ARM-A9 Datasheet. Retrieved from http:\/\/www.arm.com\/files\/pdf\/ARMCortexA-9Processors.pdf.  ARM-A9. 2009. ARM-A9 Datasheet. Retrieved from http:\/\/www.arm.com\/files\/pdf\/ARMCortexA-9Processors.pdf."},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","unstructured":"F. Bouwens M. Berekovic A. Kanstein and G. Gaydadjiev. 2007. Architectural exploration of the ADRES coarse-grained reconfigurable array. In ARC\u201907. 1--13. http:\/\/dl.acm.org\/citation.cfm&quest;id=1764631.1764633.   F. Bouwens M. Berekovic A. Kanstein and G. Gaydadjiev. 2007. Architectural exploration of the ADRES coarse-grained reconfigurable array. In ARC\u201907. 1--13. http:\/\/dl.acm.org\/citation.cfm&quest;id=1764631.1764633.","DOI":"10.1007\/978-3-540-71431-6_1"},{"key":"e_1_2_1_3_1","volume-title":"Tesla S2050 GPU Computing System.","year":"2010"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2008.07.002"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2005.8"},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"C. Ebeling D. C. Cronquist P. Franklin J. Secosky and S. G. Berg. 1997. Mapping applications to the RaPiD configurable architecture. In FCCM\u201997. IEEE Computer Society 106--115. DOI: http:\/\/dx.doi.org\/10.1109\/FPGA.1997.624610   C. Ebeling D. C. Cronquist P. Franklin J. Secosky and S. G. Berg. 1997. Mapping applications to the RaPiD configurable architecture. In FCCM\u201997. IEEE Computer Society 106--115. DOI: http:\/\/dx.doi.org\/10.1109\/FPGA.1997.624610","DOI":"10.1109\/FPGA.1997.624610"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1508128.1508158"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2228360.2228600"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463209.2488756"},{"key":"e_1_2_1_10_1","volume-title":"DATE\u201901","author":"Hartenstein R."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/224818.224959"},{"key":"e_1_2_1_12_1","doi-asserted-by":"crossref","unstructured":"A. Hatanaka and N. Bagherzadeh. 2007. A modulo scheduling algorithm for a coarse-grain reconfigurable array template. In IPDPS\u201907. 1--8. DOI: http:\/\/dx.doi.org\/10.1109\/IPDPS.2007.370371  A. Hatanaka and N. Bagherzadeh. 2007. A modulo scheduling algorithm for a coarse-grain reconfigurable array template. In IPDPS\u201907. 1--8. DOI: http:\/\/dx.doi.org\/10.1109\/IPDPS.2007.370371","DOI":"10.1109\/IPDPS.2007.370371"},{"key":"e_1_2_1_13_1","unstructured":"Intel-N550. 2010. Intel N550 Datasheet. Retrieved from http:\/\/ark.intel.com\/products\/50154\/Intel-Atom- Processor-N550-(1M-Cache-1_50-GHz).  Intel-N550. 2010. Intel N550 Datasheet. Retrieved from http:\/\/ark.intel.com\/products\/50154\/Intel-Atom- Processor-N550-(1M-Cache-1_50-GHz)."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/DATE.2005.260"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2009.2025280"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1155\/2009\/518659"},{"key":"e_1_2_1_17_1","volume-title":"DRESC: A retargetable compiler for coarse-grained reconfigurable architectures. In FTP\u201902. 166--173. DOI: http:\/\/dx.doi.org\/10.1109\/FPT.2002.1188678","author":"Mei B.","year":"2002"},{"key":"e_1_2_1_18_1","doi-asserted-by":"crossref","unstructured":"B. Mei S. Vernalde D. Verkest H. De Man and R. Lauwereins. 2003. Exploiting loop-level parallelism on coarse-grained reconfigurable architectures using modulo scheduling. In DATE\u201903. IEEE Computer Society. 296--301. DOI: http:\/\/dx.doi.org\/10.1109\/DATE.2003.1253623   B. Mei S. Vernalde D. Verkest H. De Man and R. Lauwereins. 2003. Exploiting loop-level parallelism on coarse-grained reconfigurable architectures using modulo scheduling. In DATE\u201903. IEEE Computer Society. 296--301. DOI: http:\/\/dx.doi.org\/10.1109\/DATE.2003.1253623","DOI":"10.1109\/DATE.2003.1253623"},{"key":"e_1_2_1_19_1","volume-title":"International Conference on Field Programmable Logic and Applications, 2005","author":"Mei B.","year":"2005"},{"key":"e_1_2_1_20_1","doi-asserted-by":"crossref","unstructured":"B. Mei M. Berekovic and J.-Y. Mignolet. 2007. ADRES & DRESC: Architecture and compiler for coarse-grain reconfigurable processors. In Fine- and Coarse-Grain Reconfigurable Computing S. Vassiliadis and D. Soudris (Eds.). Springer Netherlands 255--297. DOI: http:\/\/dx.doi.org\/10.1007\/978-1-4020-6505-76  B. Mei M. Berekovic and J.-Y. Mignolet. 2007. ADRES & DRESC: Architecture and compiler for coarse-grain reconfigurable processors. In Fine- and Coarse-Grain Reconfigurable Computing S. Vassiliadis and D. Soudris (Eds.). Springer Netherlands 255--297. DOI: http:\/\/dx.doi.org\/10.1007\/978-1-4020-6505-76","DOI":"10.1007\/978-1-4020-6505-7_6"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1454115.1454140"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669160"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1176760.1176778"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1629395.1629433"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1629395.1629433"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/192724.192731"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2011.77"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/12.859540"},{"key":"e_1_2_1_29_1","volume-title":"SPKM: A novel graph drawing based algorithm for application mapping onto coarse-grained reconfigurable architectures. In DAC\u201908. 776--782. DOI: http:\/\/dx.doi.org\/10.1109\/ASPDAC.2008.4484056","author":"Yoon J. W.","year":"2008"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2638558","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2638558","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T08:10:33Z","timestamp":1750234233000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2638558"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,1,21]]},"references-count":29,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2015,1,21]]}},"alternative-id":["10.1145\/2638558"],"URL":"https:\/\/doi.org\/10.1145\/2638558","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,1,21]]},"assertion":[{"value":"2011-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-01-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}