{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:45:54Z","timestamp":1787017554263,"version":"build-2736575974"},"reference-count":90,"publisher":"Association for Computing Machinery (ACM)","issue":"PLDI","funder":[{"name":"NSF","award":["2403144"],"award-info":[{"award-number":["2403144"]}]},{"name":"U.S. Department of Energy Computational Science Graduate Fellowship","award":["DESC0022158"],"award-info":[{"award-number":["DESC0022158"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Program. Lang."],"published-print":{"date-parts":[[2025,6,10]]},"abstract":"<jats:p>\n                    Spatial dataflow architectures (SDAs) are a promising and versatile accelerator platform. They are software-programmable and achieve near-ASIC performance and energy efficiency, beating CPUs by orders of magnitude. Unfortunately, many SDAs struggle to efficiently implement irregular computations because they suffer from an\n                    <jats:italic toggle=\"yes\">abstraction inversion:<\/jats:italic>\n                    they fail to capture coarse-grain dataflow semantics in the application \u2014 namely asynchronous communication, pipelining, and queueing \u2014 that are naturally supported by the dataflow execution model and existing SDA hardware.\n                  <\/jats:p>\n                  <jats:p>\n                    <jats:italic toggle=\"yes\">\n                      <jats:sc>Ripple<\/jats:sc>\n                    <\/jats:italic>\n                    is a language and architecture that corrects the abstraction inversion by preserving dataflow semantics down the stack.\n                    <jats:italic toggle=\"yes\">\n                      <jats:sc>Ripple<\/jats:sc>\n                    <\/jats:italic>\n                    provides\n                    <jats:italic toggle=\"yes\">asynchronous iterators<\/jats:italic>\n                    , shared-memory\n                    <jats:monospace>atomics<\/jats:monospace>\n                    , and a familiar task-parallel interface to concisely express the asynchronous pipeline parallelism enabled by an SDA.\n                    <jats:italic toggle=\"yes\">\n                      <jats:sc>Ripple<\/jats:sc>\n                    <\/jats:italic>\n                    efficiently implements deadlock-free, asynchronous task communication by exposing hardware token queues in its ISA. Across nine important workloads, compared to a recent ordered-dataflow SDA,\n                    <jats:italic toggle=\"yes\">\n                      <jats:sc>Ripple<\/jats:sc>\n                    <\/jats:italic>\n                    shrinks programs by 1.9\u00d7, improves performance by 3\u00d7, increases IPC by 58%, and reduces dynamic instructions by 44%.\n                  <\/jats:p>","DOI":"10.1145\/3729256","type":"journal-article","created":{"date-parts":[[2025,6,13]],"date-time":"2025-06-13T16:02:27Z","timestamp":1749830547000},"page":"249-276","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Ripple: Asynchronous Programming for Spatial Dataflow Architectures"],"prefix":"10.1145","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0656-4726","authenticated-orcid":false,"given":"Souradip","family":"Ghosh","sequence":"first","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9142-1028","authenticated-orcid":false,"given":"Yufei","family":"Shi","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4130-1099","authenticated-orcid":false,"given":"Brandon","family":"Lucia","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6301-714X","authenticated-orcid":false,"given":"Nathan","family":"Beckmann","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,13]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/325164.325100"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.5555\/1177220"},{"key":"e_1_3_2_4_2","volume-title":"The (preliminary) Id report: an asynchronous programming language and computing machine (revised)","author":"Arvind Kim P Gostelow","year":"1978","unstructured":"Arvind, Kim P Gostelow, and Wil Plouffe. 1978. The (preliminary) Id report: an asynchronous programming language and computing machine (revised). Technical Report. University of California, Irvine (UCI). https:\/\/escholarship.org\/uc\/item\/0rr7573w"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/69558.69562"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/12.48862"},{"key":"e_1_3_2_7_2","unstructured":"David A. Bader Henning Meyerhenke Peter Sanders and Dorothea Wagner (Eds.). 2013. Graph Partitioning and Graph Clustering 10th DIMACS Implementation Challenge Workshop Georgia Institute of Technology Atlanta GA USA February 13-14 2012. Proceedings. Contemporary Mathematics Vol. 588. American Mathematical Society. http:\/\/dblp.uni-trier.de\/db\/conf\/dimacs\/dimacs2012.html"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/3539845.3539913"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507772"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.5555\/645420.652538"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.5555\/17299"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3078597.3078616"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/209937.209958"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1155\/2010\/521797"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.5555\/1023556"},{"key":"e_1_3_2_16_2","volume-title":"Pegasus: An Efficient Intermediate Representation","author":"Budiu Mihai","year":"2002","unstructured":"Mihai Budiu and Seth Copen Goldstein. 2002. Pegasus: An Efficient Intermediate Representation. Technical Report CMU-CS-02-107. Carnegie Mellon University."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/2093157.2093165"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1177\/1094342007078442"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/1094811.1094852"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2013.78"},{"key":"e_1_3_2_21_2","unstructured":"Gabor Csardi and Tamas Nepusz. 2006. The igraph software package for complex network research. InterJournal Complex Systems (2006) 1695. https:\/\/igraph.org"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.5555\/90523.90636"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/115372.115320"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00053"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358276"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/62297.62346"},{"key":"e_1_3_2_27_2","first-page":"5","article-title":"Deadlock-Free Message Routing in Multiprocessor Interconnection Networks","volume":"36","author":"Dally W. J.","year":"1987","unstructured":"W. J. Dally and C. L. Seitz. 1987. Deadlock-Free Message Routing in Multiprocessor Interconnection Networks. IEEE Trans. Computers 36, 5 (1987).","journal-title":"IEEE Trans. Computers"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2019.00010"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/642089.642111"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.5555\/2851099"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.5486\/PMD.1959.6.3-4.12"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/360363.360369"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/2641638.2641652"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00084"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00046"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/2.839324"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2012.51"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/1941553.1941557"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1006\/jpdc.1993.1065"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3104255"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","unstructured":"Rohan Juneja Pranav Dangi Thilini Kaushalya Bandara Zhaoying Li Tulika Mitra and Li shiuan Peh. 2025. Nexus Machine: An Active Message Inspired Reconfigurable Architecture for Irregular Workloads. arXiv:2502.12380 [cs.AR] https:\/\/arxiv.org\/abs\/2502.12380","DOI":"10.1145\/3725843.3756091"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3061639.3062262"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.5555\/502981"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037749"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3192366.3192379"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.21105\/joss.01244"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/1504176.1504181"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/1250734.1250759"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.5555\/977395.977673"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/PROC.1987.13876"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3448128"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.5555\/3023549.3023589"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/1060745.1060829"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","unstructured":"Mahim Mishra Timothy J. Callahan Tiberiu Chelcea Girish Venkataramani Seth C. Goldstein and Mihai Budiu. 2006. Tartan: evaluating spatial computation for whole program execution. (2006) 163\u2013174. doi:10.1145\/1168857.1168878","DOI":"10.1145\/1168857.1168878"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480048"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10071026"},{"key":"e_1_3_2_58_2","unstructured":"R.S. Nikhil. 1991. ID Reference Manual. Memo 284-2."},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3243176.3243212"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080255"},{"key":"e_1_3_2_61_2","article-title":"Exploring the potential of heterogeneous von neumann\/dataflow execution models","volume":"43","author":"Nowatzki Tony","year":"2015","unstructured":"Tony Nowatzki, Vinay Gangadhar, and Karthikeyan Sankaralingam. 2015. Exploring the potential of heterogeneous von neumann\/dataflow execution models. In ACM SIGARCH Computer Architecture News, Vol. 43.","journal-title":"ACM SIGARCH Computer Architecture News"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462163"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485935"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/1152154.1152184"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/1993498.1993501"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO61859.2024.00100"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080256"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1006\/jpdc.2001.1746"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480047"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA57654.2024.00016"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/859618.859667"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/HCS52781.2021.9567306"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3614283"},{"key":"e_1_3_2_74_2","volume-title":"MPI: The Complete Reference","author":"Snir Mark","year":"1996","unstructured":"Mark Snir, Steve Otto, Steven Huss-Lederman, David Walker, and Jack Dongarra. 1996. MPI: The Complete Reference. The MIT Press."},{"key":"e_1_3_2_75_2","unstructured":"Svend Haugaard Sorensen. 2013. Linux call graph (version 3.7.10). https:\/\/sparse.tamu.edu\/Sorensen\/Linux_call_graph"},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/2740908.2744708"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.5555\/956417.956546"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA53966.2022.00030"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2011.87"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.5555\/647478.727935"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/1065944.1065975"},{"key":"e_1_3_2_82_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00042"},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00066"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00039"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.5555\/2665671.2665703"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2018.00013"},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.1109\/2.612254"},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00063"},{"key":"e_1_3_2_89_2","doi-asserted-by":"crossref","unstructured":"Joyce Whang Andrew Lenharth Inderjit S. Dhillon and Keshav Pingali. 2015. Scalable Data-driven PageRank: Algorithms System Issues and Lessons Learned. In International European Conference on Parallel and Distributed Computing (Euro-Par).","DOI":"10.1007\/978-3-662-48096-0_34"},{"key":"e_1_3_2_90_2","unstructured":"Tomofumi Yuki and Louis-Noel Pouchet. 2016. PolyBench 4.2.1: The polyhedral benchmark suite."},{"key":"e_1_3_2_91_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00085"}],"container-title":["Proceedings of the ACM on Programming Languages"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729256","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T10:05:32Z","timestamp":1784196332000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729256"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,10]]},"references-count":90,"journal-issue":{"issue":"PLDI","published-print":{"date-parts":[[2025,6,10]]}},"alternative-id":["10.1145\/3729256"],"URL":"https:\/\/doi.org\/10.1145\/3729256","relation":{},"ISSN":["2475-1421"],"issn-type":[{"value":"2475-1421","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,10]]},"assertion":[{"value":"2024-11-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}