{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T15:40:49Z","timestamp":1759333249363,"version":"3.41.0"},"reference-count":26,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2019,3,28]],"date-time":"2019-03-28T00:00:00Z","timestamp":1553731200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004663","name":"Ministry of Science and Technology, Taiwan","doi-asserted-by":"crossref","award":["MOST-105-2218-E-006 -027-MY2"],"award-info":[{"award-number":["MOST-105-2218-E-006 -027-MY2"]}],"id":[{"id":"10.13039\/501100004663","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2019,5,31]]},"abstract":"<jats:p>Heterogeneous computing leverages more than one kind of processors to boost the performance of user-space applications with the heterogeneous programming languages, e.g., OpenCL. While some works have been done to accelerate the computations required by Linux kernel software, they are either application-specific solutions or tightly coupled with the certain computing platforms and are not able to support the general-purpose in-kernel accelerations using different types of processors. In this article, the general-purpose software framework called Kernel acceleration with OpenCL (KOCL), is proposed to tackle the problem. KOCL exposes a set of the high-level programming interfaces for the Linux kernel module developers to offload compute-intensive tasks on different hardware accelerators without managing and coordinating the platform-specific computing and memory resources. The simplified programming efforts are achieved by the developed platform management and memory models, which provide a systematic means of managing the heterogeneous hardware resources. In addition, the one- and zero-copy data-buffering schemes are offered by KOCL, so that the offloaded tasks deliver high performance on the platforms with different memory architectures. We have developed the prototype system to accelerate the Network-Attached Storage server applications. Significant performance improvements are achieved with the three different types of accelerators, i.e., the multicore processor, the integrated GPU, and the discrete GPU, respectively. We believe that KOCL is useful for the design of embedded appliances to evaluate the performance of design alternatives.<\/jats:p>","DOI":"10.1145\/3315569","type":"journal-article","created":{"date-parts":[[2019,3,29]],"date-time":"2019-03-29T12:45:01Z","timestamp":1553863501000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Augmenting Operating Systems with OpenCL Accelerators"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8967-1385","authenticated-orcid":false,"given":"Chia-Heng","family":"Tu","sequence":"first","affiliation":[{"name":"National Cheng Kung University, Taiwan (R.O.C.)"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Te-Sheng","family":"Lin","sequence":"additional","affiliation":[{"name":"National Cheng Kung University, Taiwan (R.O.C.)"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,3,28]]},"reference":[{"doi-asserted-by":"publisher","key":"e_1_2_1_1_1","DOI":"10.1016\/j.parco.2016.05.006"},{"doi-asserted-by":"publisher","key":"e_1_2_1_2_1","DOI":"10.1145\/3050748.3050760"},{"unstructured":"I. Eidus A. Arcangeli C. Wright and H. Dickins. 2018. Kernel samepage merging. Retrieved from https:\/\/www.linux-kvm.org\/page\/KSM.  I. Eidus A. Arcangeli C. Wright and H. Dickins. 2018. Kernel samepage merging. Retrieved from https:\/\/www.linux-kvm.org\/page\/KSM.","key":"e_1_2_1_3_1"},{"unstructured":"freedesktop. 2018. Beignet. Retrieved from https:\/\/www.freedesktop.org\/wiki\/Software\/Beignet\/.  freedesktop. 2018. Beignet. Retrieved from https:\/\/www.freedesktop.org\/wiki\/Software\/Beignet\/.","key":"e_1_2_1_4_1"},{"unstructured":"T. Hicks and D. Kirkland. 2018. Retrieved from eCryptfs. http:\/\/ecryptfs.org\/.  T. Hicks and D. Kirkland. 2018. Retrieved from eCryptfs. http:\/\/ecryptfs.org\/.","key":"e_1_2_1_5_1"},{"unstructured":"Intel. 2018. INTEL AES-NI. Retrieved from https:\/\/software.intel.com\/sites\/default\/files\/m\/d\/4\/1\/d\/8\/Introduction_to_Intel_Secure_Key_Instructions.pdf.  Intel. 2018. INTEL AES-NI. Retrieved from https:\/\/software.intel.com\/sites\/default\/files\/m\/d\/4\/1\/d\/8\/Introduction_to_Intel_Secure_Key_Instructions.pdf.","key":"e_1_2_1_6_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_7_1","DOI":"10.1007\/s10766-014-0320-y"},{"doi-asserted-by":"crossref","unstructured":"D. R. Kaeli P. Mistry D. Schaa and D. P. Zhang. 2015. Heterogeneous Computing with OpenCL 2.0 (1st ed.). Morgan Kaufmann Publishers Inc.   D. R. Kaeli P. Mistry D. Schaa and D. P. Zhang. 2015. Heterogeneous Computing with OpenCL 2.0 (1st ed.). Morgan Kaufmann Publishers Inc.","key":"e_1_2_1_8_1","DOI":"10.1016\/B978-0-12-801414-1.00001-6"},{"doi-asserted-by":"publisher","key":"e_1_2_1_9_1","DOI":"10.1145\/2304576.2304623"},{"doi-asserted-by":"crossref","unstructured":"W. C. Lin C. H. Tu C. W. Yeh and S. H. Hung. 2017. GPU acceleration for kernel samepage merging. In RTCSA. 1--6.  W. C. Lin C. H. Tu C. W. Yeh and S. H. Hung. 2017. GPU acceleration for kernel samepage merging. In RTCSA. 1--6.","key":"e_1_2_1_10_1","DOI":"10.1109\/RTCSA.2017.8046334"},{"doi-asserted-by":"crossref","unstructured":"Y. Luo S. Li K. Sun R. Renteria and K. Choi. 2017. Implementation of deep learning neural network for real-time object recognition in OpenCL framework. In ISOCC. 298--299.  Y. Luo S. Li K. Sun R. Renteria and K. Choi. 2017. Implementation of deep learning neural network for real-time object recognition in OpenCL framework. In ISOCC. 298--299.","key":"e_1_2_1_11_1","DOI":"10.1109\/ISOCC.2017.8368905"},{"key":"e_1_2_1_12_1","first-page":"2","article-title":"Multi-GPU implementation of machine learning algorithm using CUDA and OpenCL","volume":"5","author":"Masek J.","year":"2016","journal-title":"Int. J. Adv. Telecommun. Electrotech. Sign. Syst."},{"unstructured":"C. Nugteren. 2018. CLTune: An automatic OpenCL and CUDA kernel tuner. Retrieved from https:\/\/github.com\/CNugteren\/CLTune.  C. Nugteren. 2018. CLTune: An automatic OpenCL and CUDA kernel tuner. Retrieved from https:\/\/github.com\/CNugteren\/CLTune.","key":"e_1_2_1_13_1"},{"unstructured":"NVIDIA Corporation. 2018. CUDA:Compute Unified Device Architecture. Retrieved from https:\/\/developer.nvidia.com\/about-cuda.  NVIDIA Corporation. 2018. CUDA:Compute Unified Device Architecture. Retrieved from https:\/\/developer.nvidia.com\/about-cuda.","key":"e_1_2_1_14_1"},{"unstructured":"NXP Semiconductors. 2018. RD-IMX6Q-SABRE: SABRE Board for Smart Devices Based on the i.MX 6Quad Applications Processors. Retrieved from https:\/\/www.nxp.com\/support\/developer-resources\/evaluation-and-development-boards\/sabre-development-system\/sabre-board-for-smart-devices-based-on-the-i.mx-6quad-applications-processors:RD-IMX6Q-SABRE.  NXP Semiconductors. 2018. RD-IMX6Q-SABRE: SABRE Board for Smart Devices Based on the i.MX 6Quad Applications Processors. Retrieved from https:\/\/www.nxp.com\/support\/developer-resources\/evaluation-and-development-boards\/sabre-development-system\/sabre-board-for-smart-devices-based-on-the-i.mx-6quad-applications-processors:RD-IMX6Q-SABRE.","key":"e_1_2_1_15_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_16_1","DOI":"10.1145\/2581122.2544163"},{"volume-title":"SPIN: Seamless operating system integration of peer-to-peer DMA between SSDs and GPUs. In USENIX ATC. 167--179.","year":"2017","author":"Shai B.","key":"e_1_2_1_17_1"},{"volume-title":"Gdev: First-class GPU resource management in the operating system. In USENIX ATC. 37--37.","year":"2012","author":"Shinpei K.","key":"e_1_2_1_18_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_19_1","DOI":"10.1145\/2499368.2451169"},{"doi-asserted-by":"publisher","key":"e_1_2_1_20_1","DOI":"10.1145\/2963098"},{"unstructured":"W. Sun and R. Ricci. 2013. Augmenting operating systems with the GPU. http:\/\/arxiv.org\/abs\/1305.3345. CoRR (May 2013). arxiv:1305.3345  W. Sun and R. Ricci. 2013. Augmenting operating systems with the GPU. http:\/\/arxiv.org\/abs\/1305.3345. CoRR (May 2013). arxiv:1305.3345","key":"e_1_2_1_21_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_22_1","DOI":"10.1145\/2367589.2367595"},{"key":"e_1_2_1_23_1","first-page":"9","article-title":"The evolution of bitcoin hardware","volume":"50","author":"Taylor M. B.","year":"2017","journal-title":"Computer"},{"unstructured":"The Khronos Group Inc. 2018. OpenCL: The open standard for parallel programming of heterogeneous systems. Retrieved from http:\/\/www.khronos.org\/opencl\/.  The Khronos Group Inc. 2018. OpenCL: The open standard for parallel programming of heterogeneous systems. Retrieved from http:\/\/www.khronos.org\/opencl\/.","key":"e_1_2_1_24_1"},{"volume-title":"VOCL: An optimized environment for transparent virtualization of graphics processing units. In InPar. 1--12.","year":"2012","author":"Xiao S.","key":"e_1_2_1_25_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_26_1","DOI":"10.1145\/2688500.2688505"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3315569","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3315569","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:53:34Z","timestamp":1750204414000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3315569"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,3,28]]},"references-count":26,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2019,5,31]]}},"alternative-id":["10.1145\/3315569"],"URL":"https:\/\/doi.org\/10.1145\/3315569","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"type":"print","value":"1084-4309"},{"type":"electronic","value":"1557-7309"}],"subject":[],"published":{"date-parts":[[2019,3,28]]},"assertion":[{"value":"2018-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-03-28","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}