{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T10:08:06Z","timestamp":1781258886231,"version":"3.54.1"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2017,8,16]],"date-time":"2017-08-16T00:00:00Z","timestamp":1502841600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"C-FAR"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2017,9,30]]},"abstract":"<jats:p>\n            Specialized image processing accelerators are necessary to deliver the performance and energy efficiency required by important applications in computer vision, computational photography, and augmented reality. But creating, \u201cprogramming,\u201d and integrating this hardware into a hardware\/software system is difficult. We address this problem by extending the image processing language\n            <jats:italic>Halide<\/jats:italic>\n            so users can specify which portions of their applications should become hardware accelerators, and then we provide a compiler that uses this code to automatically create the accelerator along with the \u201cglue\u201d code needed for the user\u2019s application to access this hardware. Starting with Halide not only provides a very high-level functional description of the hardware but also allows our compiler to generate a complete software application, which accesses the hardware for acceleration when appropriate. Our system also provides high-level semantics to explore different mappings of applications to a heterogeneous system, including the flexibility of being able to change the throughput rate of the generated hardware.\n          <\/jats:p>\n          <jats:p>We demonstrate our approach by mapping applications to a commercial Xilinx Zynq system. Using its FPGA with two low-power ARM cores, our design achieves up to 6\u00d7 higher performance and 38\u00d7 lower energy compared to the quad-core ARM CPU on an NVIDIA Tegra K1, and 3.5\u00d7 higher performance with 12\u00d7 lower energy compared to the K1\u2019s 192-core GPU.<\/jats:p>","DOI":"10.1145\/3107953","type":"journal-article","created":{"date-parts":[[2017,8,24]],"date-time":"2017-08-24T11:49:04Z","timestamp":1503575344000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":84,"title":["Programming Heterogeneous Systems from an Image Processing DSL"],"prefix":"10.1145","volume":"14","author":[{"given":"Jing","family":"Pu","sequence":"first","affiliation":[{"name":"Stanford University, Stanford, California"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Steven","family":"Bell","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, California"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xuan","family":"Yang","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, California"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jeff","family":"Setter","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, California"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stephen","family":"Richardson","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, California"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jonathan","family":"Ragan-Kelley","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, California"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mark","family":"Horowitz","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, California"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2017,8,16]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778766"},{"key":"e_1_2_1_2_1","unstructured":"Altera. 2016. Intel FPGA SDK for OpenCL. Retrieved from https:\/\/www.altera.com\/products\/design-software\/embedded-software-developers\/opencl\/overview.html.  Altera. 2016. Intel FPGA SDK for OpenCL. Retrieved from https:\/\/www.altera.com\/products\/design-software\/embedded-software-developers\/opencl\/overview.html."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.347995"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2228360.2228411"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2514740"},{"key":"e_1_2_1_7_1","unstructured":"J. M. P. Cardoso and P. C. Diniz. 2011. Compilation Techniques for Reconfigurable Architectures. Vol. 81. Springer Science & Business Media.   J. M. P. Cardoso and P. C. Diniz. 2011. Compilation Techniques for Reconfigurable Architectures. Vol. 81. Springer Science & Business Media."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2842615"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2948976"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2967938.2967969"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2011.2110592"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 2012 22nd International Conference on Field Programmable Logic and Applications (FPL\u201912)","author":"Czajkowski T. S.","unstructured":"T. S. Czajkowski , U. Aydonat , D. Denisenko , J. Freeman , M. Kinsner , D. Neto , J. Wong , P. Yiannacouras , and D. P. Singh . 2012. From OpenCLto high-performance hardware on FPGAs . In Proceedings of the 2012 22nd International Conference on Field Programmable Logic and Applications (FPL\u201912) . IEEE, 531--534. T. S. Czajkowski, U. Aydonat, D. Denisenko, J. Freeman, M. Kinsner, D. Neto, J. Wong, P. Yiannacouras, and D. P. Singh. 2012. From OpenCLto high-performance hardware on FPGAs. In Proceedings of the 2012 22nd International Conference on Field Programmable Logic and Applications (FPL\u201912). IEEE, 531--534."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/518909.791826"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819226"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.982918"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201916)","author":"Gao M.","unstructured":"M. Gao and C. Kozyrakis . 2016. HRL: Efficient and flexible reconfigurable logic for near-data processing . In Proceedings of the 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201916) . IEEE, 126--137. M. Gao and C. Kozyrakis. 2016. HRL: Efficient and flexible reconfigurable logic for near-data processing. In Proceedings of the 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201916). IEEE, 126--137."},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 2014 24th International Conference on Field Programmable Logic and Applications (FPL\u201914)","author":"George N.","unstructured":"N. George , H.-J. Lee , D. Novo , T. Rompf , K. J. Brown , A. K. Sujeeth , M. Odersky , K. Olukotun , and P. Ienne . 2014. Hardware system synthesis from domain-specific languages . In Proceedings of the 2014 24th International Conference on Field Programmable Logic and Applications (FPL\u201914) . IEEE, 1--8. N. George, H.-J. Lee, D. Novo, T. Rompf, K. J. Brown, A. K. Sujeeth, M. Odersky, K. Olukotun, and P. Ienne. 2014. Hardware system synthesis from domain-specific languages. In Proceedings of the 2014 24th International Conference on Field Programmable Logic and Applications (FPL\u201914). IEEE, 1--8."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 2011 IEEE 17th International Symposium on High Performance Computer Architecture (HPCA\u201911)","author":"Govindaraju V.","unstructured":"V. Govindaraju , C.-H. Ho , and K. Sankaralingam . 2011. Dynamically specialized datapaths for energy efficient computing . In Proceedings of the 2011 IEEE 17th International Symposium on High Performance Computer Architecture (HPCA\u201911) . IEEE, 503--514. V. Govindaraju, C.-H. Ho, and K. Sankaralingam. 2011. Dynamically specialized datapaths for energy efficient computing. In Proceedings of the 2011 IEEE 17th International Symposium on High Performance Computer Architecture (HPCA\u201911). IEEE, 503--514."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/DATE.2005.234"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1369396.1369402"},{"key":"e_1_2_1_21_1","volume-title":"Reconfigurable Computing: The Theory and Practice of FPGA-Based Computation.","author":"Hauck S.","year":"2010","unstructured":"S. Hauck and A. DeHon . 2010 . Reconfigurable Computing: The Theory and Practice of FPGA-Based Computation. Vol. 1 . Morgan Kaufmann , Burlington, MA . S. Hauck and A. DeHon. 2010. Reconfigurable Computing: The Theory and Practice of FPGA-Based Computation. Vol. 1. Morgan Kaufmann, Burlington, MA."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2601097.2601174"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925892"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/4.102668"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.20"},{"key":"e_1_2_1_27_1","doi-asserted-by":"crossref","unstructured":"Y. LeCun Y. Bengio and G. Hinton. 2015. Deep learning. Nature 521 7553 (2015) 436--444.  Y. LeCun Y. Bengio and G. Hinton. 2015. Deep learning. Nature 521 7553 (2015) 436--444.","DOI":"10.1038\/nature14539"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/MDT.2009.83"},{"key":"e_1_2_1_29_1","unstructured":"Maxeler Acceleration Technology. 2011. MaxCompiler White Paper. Retrieved from https:\/\/www.maxeler.com\/media\/documents\/MaxelerWhitePaperMaxCompiler.pdf.  Maxeler Acceleration Technology. 2011. MaxCompiler White Paper. Retrieved from https:\/\/www.maxeler.com\/media\/documents\/MaxelerWhitePaperMaxCompiler.pdf."},{"key":"e_1_2_1_30_1","volume-title":"ADRES: An Architecture with Tightly Coupled VLIW Processor and Coarse-Grained Reconfigurable Matrix","author":"Mei B.","year":"2003","unstructured":"B. Mei , S. Vernalde , D. Verkest , H. De Man , and R. Lauwereins . 2003 . ADRES: An Architecture with Tightly Coupled VLIW Processor and Coarse-Grained Reconfigurable Matrix . Springer , Berlin , 61--70. B. Mei, S. Vernalde, D. Verkest, H. De Man, and R. Lauwereins. 2003. ADRES: An Architecture with Tightly Coupled VLIW Processor and Coarse-Grained Reconfigurable Matrix. Springer, Berlin, 61--70."},{"key":"e_1_2_1_31_1","unstructured":"R. Membarth and O. Reiche. 2016. Fork of HIPAccgenerating code for Vivado HLS. https:\/\/github.com\/hipacc\/hipacc-vivado. (2016).  R. Membarth and O. Reiche. 2016. Fork of HIPAccgenerating code for Vivado HLS. https:\/\/github.com\/hipacc\/hipacc-vivado. (2016)."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2015.2394802"},{"key":"e_1_2_1_33_1","unstructured":"Mentor Graphics. 2016. Catapult High-Level Synthesis. Retrieved from https:\/\/www.mentor.com\/hls-lp\/catapult-high-level-synthesis\/. (2016).  Mentor Graphics. 2016. Catapult High-Level Synthesis. Retrieved from https:\/\/www.mentor.com\/hls-lp\/catapult-high-level-synthesis\/. (2016)."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2159542.2159547"},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915)","author":"Moreau T.","unstructured":"T. Moreau , M. Wyse , J. Nelson , A. Sampson , H. Esmaeilzadeh , L. Ceze , and M. Oskin . 2015. SNNAP: Approximate computing on programmable SoCs via neural acceleration . In Proceedings of the 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915) . IEEE, 603--614. T. Moreau, M. Wyse, J. Nelson, A. Sampson, H. Esmaeilzadeh, L. Ceze, and M. Oskin. 2015. SNNAP: Approximate computing on programmable SoCs via neural acceleration. In Proceedings of the 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915). IEEE, 603--614."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925952"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2003.1220583"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/1365490.1365500"},{"key":"e_1_2_1_39_1","unstructured":"ON Semiconductor. 2015. 1\/4-Inch 5 Mp System-on-a-Chip (SOC) CMOS Digital Image Sensor. Retrieved from http:\/\/www.onsemi.cn\/PowerSolutions\/document\/MT9P111-D.PDF.  ON Semiconductor. 2015. 1\/4-Inch 5 Mp System-on-a-Chip (SOC) CMOS Digital Image Sensor. Retrieved from http:\/\/www.onsemi.cn\/PowerSolutions\/document\/MT9P111-D.PDF."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2011.19"},{"key":"e_1_2_1_41_1","volume-title":"Proceedings of the 2016 26th International Conference on Field Programmable Logic and Applications (FPL\u201916)","author":"\u00d6zkan M. A.","unstructured":"M. A. \u00d6zkan , O. Reiche , F. Hannig , and J. Teich . 2016. FPGA-based accelerator design from a domain-specific language . In Proceedings of the 2016 26th International Conference on Field Programmable Logic and Applications (FPL\u201916) . IEEE, 1--9. M. A. \u00d6zkan, O. Reiche, F. Hannig, and J. Teich. 2016. FPGA-based accelerator design from a domain-specific language. In Proceedings of the 2016 26th International Conference on Field Programmable Logic and Applications (FPL\u201916). IEEE, 1--9."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/SASP.2009.5226333"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-007-0110-8"},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the 2011 IEEE 17th International Symposium on High Performance Computer Architecture. 406--417","author":"Pellauer M.","unstructured":"M. Pellauer , M. Adler , M. Kinsy , A. Parashar , and J. Emer . 2011. HAsim: FPGA-based high-detail multicore simulation using time-division multiplexing . In Proceedings of the 2011 IEEE 17th International Symposium on High Performance Computer Architecture. 406--417 . M. Pellauer, M. Adler, M. Kinsy, A. Parashar, and J. Emer. 2011. HAsim: FPGA-based high-detail multicore simulation using time-division multiplexing. In Proceedings of the 2011 IEEE 17th International Symposium on High Performance Computer Architecture. 406--417."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2872362.2872415"},{"key":"e_1_2_1_46_1","unstructured":"Qualcomm Inc. 2016. Snapdragon 800 Series Mobile Processors. Retrieved from https:\/\/www.qualcomm.com\/products\/snapdragon\/processors\/800-series.  Qualcomm Inc. 2016. Snapdragon 800 Series Mobile Processors. Retrieved from https:\/\/www.qualcomm.com\/products\/snapdragon\/processors\/800-series."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185528"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462176"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2005.1407713"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2656075.2656081"},{"key":"e_1_2_1_51_1","first-page":"148","article-title":"A realtime image processing chip set. In 1986 IEEE International Solid-State Circuits Conference","author":"Ruetz P.","year":"1986","unstructured":"P. Ruetz and R. Brodersen . 1986 . A realtime image processing chip set. In 1986 IEEE International Solid-State Circuits Conference . Digest of Technical Papers , Vol. XXIX. 148 -- 149 . P. Ruetz and R. Brodersen. 1986. A realtime image processing chip set. In 1986 IEEE International Solid-State Circuits Conference. Digest of Technical Papers, Vol. XXIX. 148--149.","journal-title":"Digest of Technical Papers"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2010.69"},{"key":"e_1_2_1_53_1","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML\u201911)","author":"Sujeeth A.","unstructured":"A. Sujeeth , H.-J. Lee , K. Brown , T. Rompf , H. Chafi , M. Wu , A. Atreya , M. Odersky , and K. Olukotun . 2011. OptiML: An implicitly parallel domain-specific language for machine learning . In Proceedings of the 28th International Conference on Machine Learning (ICML\u201911) . JMLR, Bellevue, WA, 609--616. A. Sujeeth, H.-J. Lee, K. Brown, T. Rompf, H. Chafi, M. Wu, A. Atreya, M. Odersky, and K. Olukotun. 2011. OptiML: An implicitly parallel domain-specific language for machine learning. In Proceedings of the 28th International Conference on Machine Learning (ICML\u201911). JMLR, Bellevue, WA, 609--616."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2008.240"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/143095.143131"},{"key":"e_1_2_1_56_1","unstructured":"Xilinx. 2016. AXI DMA v7.1 LogiCORE IPProduct Guide. Retrieved from http:\/\/www.xilinx.com\/support\/documentation\/ip_documentation\/axi_dma\/v7_1\/pg021_axi_dma.pdf.  Xilinx. 2016. AXI DMA v7.1 LogiCORE IPProduct Guide. Retrieved from http:\/\/www.xilinx.com\/support\/documentation\/ip_documentation\/axi_dma\/v7_1\/pg021_axi_dma.pdf."},{"key":"e_1_2_1_57_1","unstructured":"Xilinx. 2016. Vivado High-Level Synthesis. Retrieved from http:\/\/www.xilinx.com\/products\/design-tools\/vivado\/integration\/esl-design.html.  Xilinx. 2016. Vivado High-Level Synthesis. Retrieved from http:\/\/www.xilinx.com\/products\/design-tools\/vivado\/integration\/esl-design.html."},{"key":"e_1_2_1_58_1","unstructured":"Xilinx. 2016. Xilinx Wiki - Open Source Linux. Retrieved from http:\/\/www.wiki.xilinx.com\/Open+Source+Linux.  Xilinx. 2016. Xilinx Wiki - Open Source Linux. Retrieved from http:\/\/www.wiki.xilinx.com\/Open+Source+Linux."},{"key":"e_1_2_1_59_1","unstructured":"Xilinx. 2016. Zynq-7000 All Programmable SoCOverview. Retrieved from http:\/\/www.xilinx.com\/support\/documentation\/data_sheets\/ds190-Zynq-7000-Overview.pdf.  Xilinx. 2016. Zynq-7000 All Programmable SoCOverview. Retrieved from http:\/\/www.xilinx.com\/support\/documentation\/data_sheets\/ds190-Zynq-7000-Overview.pdf."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"e_1_2_1_61_1","doi-asserted-by":"crossref","unstructured":"Z. Zhang Y. Fan W. Jiang G. Han C. Yang and J. Cong. 2008. AutoPilot: A Platform-Based ESL Synthesis System. Springer Netherlands Dordrecht 99--112.  Z. Zhang Y. Fan W. Jiang G. Han C. Yang and J. Cong. 2008. AutoPilot: A Platform-Based ESL Synthesis System. Springer Netherlands Dordrecht 99--112.","DOI":"10.1007\/978-1-4020-8588-8_6"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3107953","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3107953","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:30:25Z","timestamp":1750217425000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3107953"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,8,16]]},"references-count":59,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2017,9,30]]}},"alternative-id":["10.1145\/3107953"],"URL":"https:\/\/doi.org\/10.1145\/3107953","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,8,16]]},"assertion":[{"value":"2016-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-08-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}