{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,2]],"date-time":"2025-12-02T15:07:36Z","timestamp":1764688056710,"version":"3.37.3"},"reference-count":25,"publisher":"Oxford University Press (OUP)","issue":"5","license":[{"start":{"date-parts":[[2023,8,11]],"date-time":"2023-08-11T00:00:00Z","timestamp":1691712000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2021ZD0110202"],"award-info":[{"award-number":["2021ZD0110202"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,6,22]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>While tensor accelerated compilers have proven effective in deploying deep neural networks (DNN) on general-purpose hardware, optimizing for FPGA remains challenging due to the complex DNN architectures and the heterogeneous, semi-open compute units. This paper introduces the Automatic Kernel Generation for DNN on CPU-FPGA (AKGF) framework for efficient deployment of DNN on heterogeneous CPU-FPGA platforms. AKGF generates an intermediate representation (IR) of the DNN using TVM\u2019s Halide IR, annotates the operators of model layers in the IR to compute them on the corresponding hardware cores, and further optimizes the operator code for CPU and FPGA using ARM\u2019s function library and the polyhedral model to enhance model inference speed and power consumption. The experimental tests conducted on a CPU-FPGA board validate the effectiveness of AKGF, demonstrating significant acceleration ratios (up to 6.7x) compared to state-of-the-art accelerators while achieving a 2x power optimization. AKGF effectively leverages the computational capabilities of both CPU and FPGA for high-performance deployment of DNN on CPU-FPGA platforms.<\/jats:p>","DOI":"10.1093\/comjnl\/bxad086","type":"journal-article","created":{"date-parts":[[2023,8,12]],"date-time":"2023-08-12T08:13:53Z","timestamp":1691828033000},"page":"1619-1627","source":"Crossref","is-referenced-by-count":4,"title":["AKGF: Automatic Kernel Generation for DNN on CPU-FPGA"],"prefix":"10.1093","volume":"67","author":[{"given":"Dong","family":"Dong","sequence":"first","affiliation":[{"name":"Beijing Key Laboratory of Digital Media, State Key Lab Virtual Real Technology and Systems, Beihang University , Beijing, 100191, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongxu","family":"Jiang","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Digital Media, State Key Lab Virtual Real Technology and Systems, Beihang University , Beijing, 100191, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Boyu","family":"Diao","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, Chinese Academy of Sciences , Beijing, 100086, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2023,8,11]]},"reference":[{"key":"2024062312365602700_ref1","doi-asserted-by":"crossref","first-page":"519","DOI":"10.1145\/2499370.2462176","article-title":"Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines","volume":"48","author":"Ragan-Kelley","year":"2013","journal-title":"ACM Sigplan Notices"},{"key":"2024062312365602700_ref2","doi-asserted-by":"crossref","first-page":"1400","DOI":"10.1093\/comjnl\/bxac017","article-title":"Simplified high level parallelism expression on heterogeneous systems through data partition pattern description","volume":"66","author":"Wu","year":"2023","journal-title":"Comput. J."},{"key":"2024062312365602700_ref3","doi-asserted-by":"crossref","first-page":"1236","DOI":"10.1109\/TCAD.2021.3089667","article-title":"A fast precision tuning solution for always-on DNN accelerators","volume":"41","author":"Wang","year":"2022","journal-title":"IEEE Trans. Comput. Aided Des. Integr. Circuits Syst"},{"key":"2024062312365602700_ref4","first-page":"948","article-title":"Towards Intelligent Compiler Optimization","volume-title":"2022 45th Jubilee Int. Convention on Information, Communication and Electronic Technology (MIPRO)","author":"Kovac","year":"2022"},{"key":"2024062312365602700_ref5","first-page":"316","article-title":"Warping cache simulation of polyhedral programs","volume-title":"Proc. 43rd ACM SIGPLAN Int. Conf. on Programming Language Design and Implementation","author":"Morelli"},{"key":"2024062312365602700_ref6","first-page":"299","article-title":"isl: an Integer Set Library for the Polyhedral Model","volume-title":"Mathematical Software\u2013ICMS 2010: Third Int. Congress on Mathematical Software","author":"Verdoolaege"},{"key":"2024062312365602700_ref7","doi-asserted-by":"crossref","first-page":"228","DOI":"10.1086\/513316","article-title":"Pluto: a numerical code for computational astrophysics","volume":"170","author":"Mignone","journal-title":"Astrophys. J. Suppl. Ser."},{"key":"2024062312365602700_ref8","first-page":"138","article-title":"Pencil: A Platform-neutral Compute Intermediate Language for Accelerator Programming","volume-title":"2015 Int. Conf. on Parallel Architecture and Compilation (PACT)","author":"Baghdadi","year":"2007"},{"key":"2024062312365602700_ref9","doi-asserted-by":"crossref","first-page":"1250010","DOI":"10.1142\/S0129626412500107","article-title":"Polly\u2014performing polyhedral optimizations on a low-level intermediate representation","volume":"22","author":"Grosser","year":"2012","journal-title":"Parallel Process. Lett."},{"key":"2024062312365602700_ref10"},{"key":"2024062312365602700_ref11","doi-asserted-by":"crossref","first-page":"429","DOI":"10.1145\/2786763.2694364","article-title":"Polymage: automatic optimization for image processing pipelines","volume":"43","author":"Mullapudi","journal-title":"ACM SIGARCH Comput. Architect. News"},{"key":"2024062312365602700_ref12","first-page":"193","article-title":"Tiramisu: A Polyhedral Compiler for Expressing Fast and Portable Code","volume-title":"2019 IEEE\/ACM Int. Symposium on Code Generation and Optimization (CGO)","author":"Baghdadi"},{"key":"2024062312365602700_ref13","first-page":"17","article-title":"Alphaz: A System for Design Space Exploration in the Polyhedral Model","volume-title":"Int. Workshop on Languages and Compilers for Parallel Computing","author":"Yuki"},{"journal-title":"Technical report. Citeseer.","article-title":"Chill: a framework for composing high-level loop transformations","author":"Chen","key":"2024062312365602700_ref14"},{"key":"2024062312365602700_ref15","first-page":"1","article-title":"Polysa: Polyhedral-based Systolic Array Auto-compilation","volume-title":"2018 IEEE\/ACM Int. Conf. Computer-Aided Design (ICCAD)","author":"Cong","year":"2013"},{"key":"2024062312365602700_ref16","first-page":"63","article-title":"Domain Specific Description in Halide for Randomized Image Convolution","volume-title":"2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conf. (APSIPA ASC)","author":"Takagi"},{"volume":"11","key":"2024062312365602700_ref17"},{"key":"2024062312365602700_ref18","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1109\/MM.2019.2928962","article-title":"A hardware\u2013software blueprint for flexible deep learning specialization","volume":"39","author":"Moreau","journal-title":"IEEE Micro"},{"key":"2024062312365602700_ref19","first-page":"51","article-title":"Heterohalide: From Image Processing DSL to Efficient FPGA Acceleration","volume-title":"Proc. 2020 ACM\/SIGDA Int. Symposium on Field-Programmable Gate Arrays","author":"Li"},{"key":"2024062312365602700_ref20","first-page":"1","article-title":"Design of a high-performance gemm-like tensor\u2013tensor multiplication","volume":"44","author":"Springer","journal-title":"ACM Trans. Math. Software (TOMS)"},{"key":"2024062312365602700_ref21","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"Imagenet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","journal-title":"Commun. ACM"},{"key":"2024062312365602700_ref22"},{"key":"2024062312365602700_ref23","first-page":"779","article-title":"You Only Look Once: Unified, Real-time Object Detection","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Redmon"},{"key":"2024062312365602700_ref24","first-page":"3","article-title":"An opencl-based parallel acceleration of a sobel edge detection algorithm using intel fpga technology","volume":"32","author":"Almomany","journal-title":"South African Comput. J."},{"key":"2024062312365602700_ref25","first-page":"79","article-title":"Research and Implementation of an Embedded Image Classification Method Based on ZYNQ","volume-title":"2022 11th Int. Conf. Information and Communication Technology (ICTech)","author":"Wang"}],"container-title":["The Computer Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/67\/5\/1619\/58307899\/bxad086.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/67\/5\/1619\/58307899\/bxad086.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,23]],"date-time":"2024-06-23T12:37:39Z","timestamp":1719146259000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/comjnl\/article\/67\/5\/1619\/7241314"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,11]]},"references-count":25,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2023,8,11]]},"published-print":{"date-parts":[[2024,6,22]]}},"URL":"https:\/\/doi.org\/10.1093\/comjnl\/bxad086","relation":{},"ISSN":["0010-4620","1460-2067"],"issn-type":[{"type":"print","value":"0010-4620"},{"type":"electronic","value":"1460-2067"}],"subject":[],"published-other":{"date-parts":[[2024,5]]},"published":{"date-parts":[[2023,8,11]]}}}