{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:07:40Z","timestamp":1750306060612,"version":"3.41.0"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2017,6,28]],"date-time":"2017-06-28T00:00:00Z","timestamp":1498608000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2017,6,30]]},"abstract":"<jats:p>In today\u2019s computers, heterogeneous processing is used to meet performance targets at manageable power. In adopting increased compute specialization, however, the relative amount of time spent on communication increases. System and software optimizations for communication often come at the costs of increased complexity and reduced portability. The Decoupled Supply-Compute (DeSC) approach offers a way to attack communication latency bottlenecks automatically, while maintaining good portability and low complexity. Our work expands prior Decoupled Access Execute techniques with hardware\/software specialization. For a range of workloads, DeSC offers roughly 2 \u00d7 speedup, and additional specialized compression optimizations reduce traffic between decoupled units by 40%.<\/jats:p>","DOI":"10.1145\/3075620","type":"journal-article","created":{"date-parts":[[2017,6,30]],"date-time":"2017-06-30T12:36:19Z","timestamp":1498826179000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Decoupling Data Supply from Computation for Latency-Tolerant Communication in Heterogeneous Architectures"],"prefix":"10.1145","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2669-6849","authenticated-orcid":false,"given":"Tae Jun","family":"Ham","sequence":"first","affiliation":[{"name":"Princeton University, NJ, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Juan L.","family":"Arag\u00f3n","sequence":"additional","affiliation":[{"name":"University of Murcia, Murcia, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Margaret","family":"Martonosi","sequence":"additional","affiliation":[{"name":"Princeton University, NJ, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,6,28]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/2337159.2337169"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1122971.1122990"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the IEEE International Symposium on Circuits and Systems","volume":"4","author":"Benini L.","unstructured":"L. Benini , D. Bruni , B. Ricco , A. Macii , and E. Macii . 2002. An adaptive data compression scheme for memory traffic minimization in processor-based systems . In Proceedings of the IEEE International Symposium on Circuits and Systems , Vol. 4 . L. Benini, D. Bruni, B. Ricco, A. Macii, and E. Macii. 2002. An adaptive data compression scheme for memory traffic minimization in processor-based systems. In Proceedings of the IEEE International Symposium on Circuits and Systems, Vol. 4."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/165939.165952"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1950413.1950423"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063454"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2629677"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555814"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"volume-title":"Proceedings of the 49th International Symposium on Microarchitecture (MICRO\u201916)","author":"Chen T.","key":"e_1_2_1_11_1","unstructured":"T. Chen and G. E. Suh . 2016. Efficient data supply for hardware accelerators with prefetching and access\/execute decoupling . In Proceedings of the 49th International Symposium on Microarchitecture (MICRO\u201916) . T. Chen and G. E. Suh. 2016. Efficient data supply for hardware accelerators with prefetching and access\/execute decoupling. In Proceedings of the 49th International Symposium on Microarchitecture (MICRO\u201916)."},{"volume-title":"Proceedings of the 34th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO\u201901)","author":"Collins Jamison D.","key":"e_1_2_1_12_1","unstructured":"Jamison D. Collins , Dean M. Tullsen , Hong Wang , and John P. Shen . 2001. Dynamic speculative precomputation . In Proceedings of the 34th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO\u201901) . Jamison D. Collins, Dean M. Tullsen, Hong Wang, and John P. Shen. 2001. Dynamic speculative precomputation. In Proceedings of the 34th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO\u201901)."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/379240.379248"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2228360.2228512"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2744769.2744794"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2000064.2000079"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1044823.1044825"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1735688.1735702"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/WWC.2003.1249063"},{"volume-title":"Proceedings of the 2nd IEEE Symposium on High-Performance Computer Architecture (HPCA\u201996)","author":"Espasa R.","key":"e_1_2_1_20_1","unstructured":"R. Espasa and M. Valero . 1996. Decoupled vector architectures . In Proceedings of the 2nd IEEE Symposium on High-Performance Computer Architecture (HPCA\u201996) . R. Espasa and M. Valero. 1996. Decoupled vector architectures. In Proceedings of the 2nd IEEE Symposium on High-Performance Computer Architecture (HPCA\u201996)."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2008.4771800"},{"volume-title":"Proceedings of the 12th Annual International Symposium on Computer Architecture (ISCA\u201985)","author":"Goodman J. R.","key":"e_1_2_1_23_1","unstructured":"J. R. Goodman , Jian-tu Hsieh, Koujuch Liou , Andrew R. Pleszkun , P. B. Schechter , and Honesty C. Young . 1985. PIPE: A VLSI decoupled architecture . In Proceedings of the 12th Annual International Symposium on Computer Architecture (ISCA\u201985) . J. R. Goodman, Jian-tu Hsieh, Koujuch Liou, Andrew R. Pleszkun, P. B. Schechter, and Honesty C. Young. 1985. PIPE: A VLSI decoupled architecture. In Proceedings of the 12th Annual International Symposium on Computer Architecture (ISCA\u201985)."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/2014698.2014884"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2830772.2830800"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750390"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2259016.2259038"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1993498.1993516"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2581122.2544161"},{"volume-title":"Proceedings of the 1st IEEE Symposium on High-Performance Computer Architecture (HPCA\u201995)","author":"John L. K.","key":"e_1_2_1_30_1","unstructured":"L. K. John , V. Reddy , P. T. Hulina , and L. D. Coraor . 1995. Program balance and its impact on high performance RISC architectures . In Proceedings of the 1st IEEE Symposium on High-Performance Computer Architecture (HPCA\u201995) . L. K. John, V. Reddy, P. T. Hulina, and L. D. Coraor. 1995. Program balance and its impact on high performance RISC architectures. In Proceedings of the 1st IEEE Symposium on High-Performance Computer Architecture (HPCA\u201995)."},{"volume-title":"Proceedings of the 39th Annual International Symposium on Computer Architecture (ISCA\u201912)","author":"Kambadur Melanie","key":"e_1_2_1_31_1","unstructured":"Melanie Kambadur , Kui Tang , and Martha A. Kim . 2012. Harmony: Collection and analysis of parallel block vectors . In Proceedings of the 39th Annual International Symposium on Computer Architecture (ISCA\u201912) . Melanie Kambadur, Kui Tang, and Martha A. Kim. 2012. Harmony: Collection and analysis of parallel block vectors. In Proceedings of the 39th Annual International Symposium on Computer Architecture (ISCA\u201912)."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/605432.605415"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2005.35"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2464996.2465012"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2005.379"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2593069.2593105"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2005.18"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/379240.379250"},{"key":"e_1_2_1_39_1","volume-title":"Proceedings of the 23rd Annual Hawaii International Conference on System Sciences","author":"Mangione-Smith W.","year":"1990","unstructured":"W. Mangione-Smith , S. G. Abraham , and E. S. Davidson . 1990. The effects of memory latency and fine-grain parallelism on Astronautics ZS-1 performance . In Proceedings of the 23rd Annual Hawaii International Conference on System Sciences , 1990 , Vol. i. W. Mangione-Smith, S. G. Abraham, and E. S. Davidson. 1990. The effects of memory latency and fine-grain parallelism on Astronautics ZS-1 performance. In Proceedings of the 23rd Annual Hawaii International Conference on System Sciences, 1990, Vol. i."},{"volume-title":"Proceedings of the 9th International Symposium on High-Performance Computer Architecture (HPCA\u201903)","author":"Mutlu Onur","key":"e_1_2_1_40_1","unstructured":"Onur Mutlu , Jared Stark , Chris Wilkerson , and Yale N. Patt . 2003. Runahead execution: An alternative to very large instruction windows for out-of-order processors . In Proceedings of the 9th International Symposium on High-Performance Computer Architecture (HPCA\u201903) . Onur Mutlu, Jared Stark, Chris Wilkerson, and Yale N. Patt. 2003. Runahead execution: An alternative to very large instruction windows for out-of-order processors. In Proceedings of the 9th International Symposium on High-Performance Computer Architecture (HPCA\u201903)."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.5555\/1299042.1299107"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.5555\/2665671.2665678"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/1356058.1356074"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.5555\/1025127.1026007"},{"key":"e_1_2_1_46_1","unstructured":"Justin Rattner. 2002. Making the Right Hand Turn to Power Efficient Computing. (2002). Retrieved from http:\/\/www.microarch.org\/micro35\/keynote\/JRattner.pdf.  Justin Rattner. 2002. Making the Right Hand Turn to Power Efficient Computing. (2002). Retrieved from http:\/\/www.microarch.org\/micro35\/keynote\/JRattner.pdf."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370864"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.5555\/2665671.2665689"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.5555\/602770.602793"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/357401.357403"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.1986.1676820"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/1024393.1024407"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2009.05.002"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2008.28"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/224170.224301"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.5555\/800078.802557"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/1013948.1013953"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2007.346187"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2005.18"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3075620","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3075620","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:03:42Z","timestamp":1750215822000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3075620"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,6,28]]},"references-count":56,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2017,6,30]]}},"alternative-id":["10.1145\/3075620"],"URL":"https:\/\/doi.org\/10.1145\/3075620","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2017,6,28]]},"assertion":[{"value":"2017-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-03-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-06-28","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}