{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T10:08:06Z","timestamp":1781258886154,"version":"3.54.1"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"4s","license":[{"start":{"date-parts":[[2014,4,1]],"date-time":"2014-04-01T00:00:00Z","timestamp":1396310400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Language and Algorithms for Heterogeneous Graph Streams"},{"DOI":"10.13039\/100004682","name":"Oracle","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100004682","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100004311","name":"Advanced Micro Devices","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100004311","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100002418","name":"Intel Corporation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100002418","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003816","name":"Huawei Technologies","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100003816","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000185","name":"Defense Advanced Research Projects Agency","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100005492","name":"Stanford University","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100005492","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000781","name":"European Research Council","doi-asserted-by":"publisher","award":["587327"],"award-info":[{"award-number":["587327"]}],"id":[{"id":"10.13039\/501100000781","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007065","name":"Nvidia","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100007065","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2014,7]]},"abstract":"<jats:p>Developing high-performance software is a difficult task that requires the use of low-level, architecture-specific programming models (e.g., OpenMP for CMPs, CUDA for GPUs, MPI for clusters). It is typically not possible to write a single application that can run efficiently in different environments, leading to multiple versions and increased complexity. Domain-Specific Languages (DSLs) are a promising avenue to enable programmers to use high-level abstractions and still achieve good performance on a variety of hardware. This is possible because DSLs have higher-level semantics and restrictions than general-purpose languages, so DSL compilers can perform higher-level optimization and translation. However, the cost of developing performance-oriented DSLs is a substantial roadblock to their development and adoption. In this article, we present an overview of the Delite compiler framework and the DSLs that have been developed with it. Delite simplifies the process of DSL development by providing common components, like parallel patterns, optimizations, and code generators, that can be reused in DSL implementations. Delite DSLs are embedded in Scala, a general-purpose programming language, but use metaprogramming to construct an Intermediate Representation (IR) of user programs and compile to multiple languages (including C++, CUDA, and OpenCL). DSL programs are automatically parallelized and different parts of the application can run simultaneously on CPUs and GPUs. We present Delite DSLs for machine learning, data querying, graph analysis, and scientific computing and show that they all achieve performance competitive to or exceeding C++ code.<\/jats:p>","DOI":"10.1145\/2584665","type":"journal-article","created":{"date-parts":[[2014,4,29]],"date-time":"2014-04-29T12:32:32Z","timestamp":1398774752000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":143,"title":["Delite"],"prefix":"10.1145","volume":"13","author":[{"given":"Arvind K.","family":"Sujeeth","sequence":"first","affiliation":[{"name":"Stanford University, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kevin J.","family":"Brown","sequence":"additional","affiliation":[{"name":"Stanford University, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hyoukjoong","family":"Lee","sequence":"additional","affiliation":[{"name":"Stanford University, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tiark","family":"Rompf","sequence":"additional","affiliation":[{"name":"Oracle Labs and EPFL"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hassan","family":"Chafi","sequence":"additional","affiliation":[{"name":"Oracle Labs and Stanford University, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Martin","family":"Odersky","sequence":"additional","affiliation":[{"name":"EPFL"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kunle","family":"Olukotun","sequence":"additional","affiliation":[{"name":"Stanford University, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2014,4]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Apache. 2014. Hadoop. http:\/\/hadoop.apache.org\/.  Apache. 2014. Hadoop. http:\/\/hadoop.apache.org\/."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1869459.1869469"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/2050135.2050143"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.scico.2007.11.003"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.15"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1941553.1941562"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1932682.1869527"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1941553.1941561"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1806596.1806638"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2254064.2254079"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the 6th Conference on Symposium on Opearting Systems Design and Implementation (OSDI'04)","author":"Dean Jeffrey","year":"2004","unstructured":"Jeffrey Dean and Sanjay Ghemawat . 2004 . MapReduce: Simplified data processing on large clusters . In Proceedings of the 6th Conference on Symposium on Opearting Systems Design and Implementation (OSDI'04) . 137--150. Jeffrey Dean and Sanjay Ghemawat. 2004. MapReduce: Simplified data processing on large clusters. In Proceedings of the 6th Conference on Symposium on Opearting Systems Design and Implementation (OSDI'04). 137--150."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063396"},{"key":"e_1_2_1_13_1","series-title":"Lecture Notes in Computer Science","volume-title":"Semantics, Applications, and Implementation of Program Generation, Walid Taha, Ed.","author":"Elliott Conal","unstructured":"Conal Elliott , Sigbj\u00f3rn Finne , and Oege de Moor . 2000. Compiling embedded languages . In Semantics, Applications, and Implementation of Program Generation, Walid Taha, Ed. , Lecture Notes in Computer Science , vol. 1924 , Springer , 9--26. Conal Elliott, Sigbj\u00f3rn Finne, and Oege de Moor. 2000. Compiling embedded languages. In Semantics, Applications, and Implementation of Program Generation, Walid Taha, Ed., Lecture Notes in Computer Science, vol. 1924, Springer, 9--26."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2048147.2048199"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2150976.2151013"},{"key":"e_1_2_1_16_1","unstructured":"Intel. 2014. Cilk plus. http:\/\/software.intel.com\/en-us\/articles\/intel-cilk-plus\/.  Intel. 2014. Cilk plus. http:\/\/software.intel.com\/en-us\/articles\/intel-cilk-plus\/."},{"key":"e_1_2_1_17_1","unstructured":"Intel. 2010. Intel array building blocks. http:\/\/software.intel.com\/en-us\/articles\/intel-array-building-blocks.  Intel. 2010. Intel array building blocks. http:\/\/software.intel.com\/en-us\/articles\/intel-array-building-blocks."},{"key":"e_1_2_1_18_1","unstructured":"Intel. 2013. Intel math kernel library. http:\/\/software.intel.com\/en-us\/intel-mkl.  Intel. 2013. Intel math kernel library. http:\/\/software.intel.com\/en-us\/intel-mkl."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1272996.1273005"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559962"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1869459.1869497"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2004.840447"},{"key":"e_1_2_1_23_1","unstructured":"The Khronos Group. 2014. OpenCL 1.0. http:\/\/www.khronos.org\/opencl\/.  The Khronos Group. 2014. OpenCL 1.0. http:\/\/www.khronos.org\/opencl\/."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/977395.977673"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/355841.355847"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/331960.331977"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1863523.1863533"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807184"},{"key":"e_1_2_1_29_1","unstructured":"MathWorks. 2014. Matlab. http:\/\/www.mathworks.com\/products\/matlab\/.  MathWorks. 2014. Matlab. http:\/\/www.mathworks.com\/products\/matlab\/."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1142473.1142552"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2103746.2103769"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2047862.2047883"},{"key":"e_1_2_1_33_1","unstructured":"Lawrence Page Sergey Brin Rajeev Motwani and Terry Winograd. 1999. The pagerank citation ranking: Bringing order to the web. Tech. rep. 1999-66. Stanford Info Lab. http:\/\/ilpubs.stanford.edu:8090\/422\/.  Lawrence Page Sergey Brin Rajeev Motwani and Terry Winograd. 1999. The pagerank citation ranking: Bringing order to the web. Tech. rep. 1999-66. Stanford Info Lab. http:\/\/ilpubs.stanford.edu:8090\/422\/."},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the USENIX Java Virtual Machine Research and Technology Symposium. 1--12","author":"Paleczny Michael","year":"2001","unstructured":"Michael Paleczny , Christopher Vick , and Cliff Click . 2001 . The java hotspot(tm) server compiler . In Proceedings of the USENIX Java Virtual Machine Research and Technology Symposium. 1--12 . Michael Paleczny, Christopher Vick, and Cliff Click. 2001. The java hotspot(tm) server compiler. In Proceedings of the USENIX Java Virtual Machine Research and Technology Symposium. 1--12."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2004.840306"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/1868294.1868314"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2429069.2429128"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.4204\/EPTCS.66.5"},{"key":"e_1_2_1_39_1","volume-title":"Armadillo: An open source c&plus;&plus","author":"Sanderson Conrad","year":"2006","unstructured":"Conrad Sanderson . 2006 . Armadillo: An open source c&plus;&plus ; linear algebra library for fast prototyping and computationally intensive experiments. Tech. rep., NICTA. http:\/\/arma.sourceforge.net\/armadillo_nicta_2010.pdf. Conrad Sanderson. 2006. Armadillo: An open source c&plus;&plus; linear algebra library for fast prototyping and computationally intensive experiments. Tech. rep., NICTA. http:\/\/arma.sourceforge.net\/armadillo_nicta_2010.pdf."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517208.2517220"},{"key":"e_1_2_1_41_1","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML'11)","author":"Sujeeth Arvind K.","unstructured":"Arvind K. Sujeeth , Hyoukjoong Lee , Kevin J. Brown , Tiark Rompf , Michael Wu , Anand R. Atreya , Martin Odersky, and Kunle Olukotun. 2011. OptiML: An implicitly parallel domain-specific language for machine learning . In Proceedings of the 28th International Conference on Machine Learning (ICML'11) . Arvind K. Sujeeth, Hyoukjoong Lee, Kevin J. Brown, Tiark Rompf, Michael Wu, Anand R. Atreya, Martin Odersky, and Kunle Olukotun. 2011. OptiML: An implicitly parallel domain-specific language for machine learning. In Proceedings of the 28th International Conference on Machine Learning (ICML'11)."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-39038-8_3"},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation (NSDI'11)","author":"Zaharia Matei","year":"2011","unstructured":"Matei Zaharia , Mosharaf Chowdhury , Tathagata Das , Ankur Dave , Justin Ma , Murphy McCauley , Michael Franklin , Scott Shenker , and Ion Stoica . 2011 . Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing . In Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation (NSDI'11) . Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauley, Michael Franklin, Scott Shenker, and Ion Stoica. 2011. Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. In Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation (NSDI'11)."}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2584665","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2584665","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T07:01:43Z","timestamp":1750230103000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2584665"}},"subtitle":["A Compiler Architecture for Performance-Oriented Embedded Domain-Specific Languages"],"short-title":[],"issued":{"date-parts":[[2014,4]]},"references-count":43,"journal-issue":{"issue":"4s","published-print":{"date-parts":[[2014,7]]}},"alternative-id":["10.1145\/2584665"],"URL":"https:\/\/doi.org\/10.1145\/2584665","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,4]]},"assertion":[{"value":"2013-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-04-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}