{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T18:35:50Z","timestamp":1780338950853,"version":"3.54.1"},"reference-count":32,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2017,8,30]],"date-time":"2017-08-30T00:00:00Z","timestamp":1504051200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2017,9,30]]},"abstract":"<jats:p>Specialized Digital Signal Processors (DSPs), which can be found in a wide range of modern devices, play an important role in power-efficient, high-performance image processing. Applications including camera sensor post-processing and computer vision benefit from being (partially) mapped onto such DSPs. However, due to their specialized instruction sets and dependence on low-level code optimization, developing applications for DSPs is more time-consuming and error-prone than for general-purpose processors. Halide is a domain-specific language (DSL) that enables low-effort development of portable, high-performance imaging pipelines\u2014a combination of qualities that is hard, if not impossible, to find among DSP programming models. We propose a set of extensions and modifications to Halide to generate code for DSP C compilers, focusing specifically on diverse SIMD target instruction sets and heterogeneous scratchpad memory hierarchies. We implement said techniques for a commercial DSP found in an Intel Image Processing Unit (IPU), demonstrating that this solution can be used to achieve performance within 20% of highly tuned, manually written C code, while leading to a reduction in code complexity. By comparing performance of Halide algorithms using our solution to results on CPU and GPU targets, we confirm the value of using DSP targets with Halide.<\/jats:p>","DOI":"10.1145\/3106343","type":"journal-article","created":{"date-parts":[[2017,8,30]],"date-time":"2017-08-30T12:52:18Z","timestamp":1504097538000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Extending Halide to Improve Software Development for Imaging DSPs"],"prefix":"10.1145","volume":"14","author":[{"given":"Sander","family":"Vocke","sequence":"first","affiliation":[{"name":"Eindhoven University of Technology, De Rondom, Eindhoven AP, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Henk","family":"Corporaal","sequence":"additional","affiliation":[{"name":"Eindhoven University of Technology, De Rondom, Eindhoven AP, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Roel","family":"Jordans","sequence":"additional","affiliation":[{"name":"Eindhoven University of Technology, De Rondom, Eindhoven AP, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rosilde","family":"Corvino","sequence":"additional","affiliation":[{"name":"Intel, Eindhoven AG, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rick","family":"Nas","sequence":"additional","affiliation":[{"name":"Intel, Eindhoven AG, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2017,8,30]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778766"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628092"},{"key":"e_1_2_1_3_1","unstructured":"Gary Bradski and Adrian Kaehler. 2008. Learning OpenCV: Computer vision with the OpenCV library. O\u2019Reilly Media Inc. Gary Bradski and Adrian Kaehler. 2008. Learning OpenCV: Computer vision with the OpenCV library. O\u2019Reilly Media Inc."},{"key":"e_1_2_1_4_1","volume-title":"Retrieved","year":"2015"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.1986.4767851"},{"key":"e_1_2_1_6_1","unstructured":"CEVA. 2015. CEVA-XM4 Intelligent Vision Processor White Paper. (Feb. 2015). CEVA. 2015. CEVA-XM4 Intelligent Vision Processor White Paper. (Feb. 2015)."},{"key":"e_1_2_1_7_1","first-page":"34","article-title":"Hexagon DSP: An architecture optimized for mobile multimedia and communications. IEEE Micro","volume":"34","author":"Codrescu Lucian","year":"2014","journal-title":"IEEE"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2851141.2851157"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/1949767.1949786"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1142\/S0129626412500107"},{"key":"e_1_2_1_11_1","unstructured":"Halide Project. 2016. Halide GitHub Repository. Retrieved July 2 2016 from http:\/\/github.com\/halide\/Halide. Halide Project. 2016. Halide GitHub Repository. Retrieved July 2 2016 from http:\/\/github.com\/halide\/Halide."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.2.23"},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Lee Howes. 2015. The OpenCL Specification v2.1 Rev 23. Specification. Khronos OpenCL Working Group. Lee Howes. 2015. The OpenCL Specification v2.1 Rev 23. Specification. Khronos OpenCL Working Group.","DOI":"10.1145\/2791321.2791337"},{"key":"e_1_2_1_14_1","unstructured":"Intel. 2016a. 6th Generation Intel Processor Datasheet for UY-Platforms. (March 2016). Intel. 2016a. 6th Generation Intel Processor Datasheet for UY-Platforms. (March 2016)."},{"key":"e_1_2_1_15_1","volume-title":"Retrieved","year":"2016"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2015.4"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 10th Conference on Pattern Languages of Programs (PLOP\u201903)","author":"Jones Joel","year":"2003"},{"key":"e_1_2_1_18_1","unstructured":"Roel Jordans. 2015. Instruction-set Architecture Synthesis for VLIW Processors. Ph.D. Dissertation. Eindhoven University of Technology. Roel Jordans. 2015. Instruction-set Architecture Synthesis for VLIW Processors. Ph.D. Dissertation. Eindhoven University of Technology."},{"key":"e_1_2_1_19_1","unstructured":"Tushar Kumar. 2016. Heterogeneous Computing Made Simpler with Symphony SDK. Retrieved April 2 2017 from https:\/\/developer.qualcomm.com\/blog\/heterogeneous-computing-made-simpler-symphony-sdk. Tushar Kumar. 2016. Heterogeneous Computing Made Simpler with Symphony SDK. Retrieved April 2 2017 from https:\/\/developer.qualcomm.com\/blog\/heterogeneous-computing-made-simpler-symphony-sdk."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/960116.54022"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/977395.977673"},{"key":"e_1_2_1_22_1","unstructured":"Zhihong Lin Jagadeesh Sankaran and Tom Flanagan. 2013. Empowering automotive vision with TIs vision accelerationpac. White Paper SPRY251. Texas Insruments. Zhihong Lin Jagadeesh Sankaran and Tom Flanagan. 2013. Empowering automotive vision with TIs vision accelerationpac. White Paper SPRY251. Texas Insruments."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2737924.2737974"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925952"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2775054.2694364"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2159430.2159431"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342004041291"},{"key":"e_1_2_1_28_1","unstructured":"Qualcomm. 2017. Symphony System Manager SDK. Retrieved April 2 2017 from http:\/\/developer.qualcomm.com\/ software\/symphony-system-manager-sdk. Qualcomm. 2017. Symphony System Manager SDK. Retrieved April 2 2017 from http:\/\/developer.qualcomm.com\/ software\/symphony-system-manager-sdk."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2499370.2462176"},{"key":"e_1_2_1_30_1","unstructured":"Jonathan Millard Ragan-Kelley. 2014. Decoupling Algorithms from the Organization of Computation for High Performance Image Processing. Ph.D. Dissertation. Massachusetts Institute of Technology. Cambridge MA. Jonathan Millard Ragan-Kelley. 2014. Decoupling Algorithms from the Organization of Computation for High Performance Image Processing. Ph.D. Dissertation. Massachusetts Institute of Technology. Cambridge MA."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2014.100"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0146-664X(78)80020-3"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3106343","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3106343","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,25]],"date-time":"2025-06-25T08:33:35Z","timestamp":1750840415000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3106343"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,8,30]]},"references-count":32,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2017,9,30]]}},"alternative-id":["10.1145\/3106343"],"URL":"https:\/\/doi.org\/10.1145\/3106343","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,8,30]]},"assertion":[{"value":"2016-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-08-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}