{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T11:10:38Z","timestamp":1784200238540,"version":"3.55.0"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"OOPSLA2","funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2023YFB4503204"],"award-info":[{"award-number":["2023YFB4503204"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Program. Lang."],"published-print":{"date-parts":[[2025,10,9]]},"abstract":"<jats:p>Practical encrypted neural network inference under the CKKS fully homomorphic encryption (FHE) scheme relies heavily on accelerating two key kernel operations: Matrix-Vector Multiplication (MVM) and Convolution (Conv). However, existing solutions\u2014such as expert-tuned libraries and domain-specific languages\u2014are designed in an ad hoc manner, leading to significant inefficiencies caused by excessive rotations.<\/jats:p>\n                  <jats:p>\n                    We introduce MKR, a novel composition-based compiler approach that optimizes MVM and Conv kernel operations for DNN models under CKKS within a unified framework. MKR decomposes each kernel into composable units, called\n                    <jats:italic toggle=\"yes\">MetaKernels<\/jats:italic>\n                    , to enhance SIMD parallelism within ciphertexts (via horizontal batching) and computational parallelism across them (via vertical batching). Our approach tackles previously unaddressed challenges, including reducing rotation overhead through a rotation-aware cost model for data packing, while also ensuring high slot utilization, uniform handling of inputs with arbitrary sizes, and compatibility with the output tensor layout. Implemented in a production-quality FHE compiler, MKR achieves inference time speedups of 10.08\u00d7\u2212185.60\u00d7 for individual MVM and Conv kernels and 1.75\u00d7\u221211.84\u00d7 for end-to-end inference compared to a state-of-the-art FHE compiler. Moreover, MKR enables homomorphic execution of large DNN models, where prior methods fail, significantly advancing the practicality of FHE compilers.\n                  <\/jats:p>","DOI":"10.1145\/3763095","type":"journal-article","created":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T08:51:31Z","timestamp":1759999891000},"page":"1261-1288","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["MetaKernel: Enabling Efficient Encrypted Neural Network Inference through Unified MVM and Convolution"],"prefix":"10.1145","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-7070-6828","authenticated-orcid":false,"given":"Peng","family":"Yuan","sequence":"first","affiliation":[{"name":"Ant Group, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-6845-7864","authenticated-orcid":false,"given":"Yan","family":"Liu","sequence":"additional","affiliation":[{"name":"Ant Group, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-5782-0454","authenticated-orcid":false,"given":"JianXin","family":"Lai","sequence":"additional","affiliation":[{"name":"Ant Group, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-2098-2193","authenticated-orcid":false,"given":"Long","family":"Li","sequence":"additional","affiliation":[{"name":"Ant Group, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-9692-5578","authenticated-orcid":false,"given":"Tianxiang","family":"Sui","sequence":"additional","affiliation":[{"name":"Ant Group, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-2778-7100","authenticated-orcid":false,"given":"Linjie","family":"Xiao","sequence":"additional","affiliation":[{"name":"Ant Group, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-8443-3803","authenticated-orcid":false,"given":"Xiaojing","family":"Zhang","sequence":"additional","affiliation":[{"name":"Ant Group, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-5144-5442","authenticated-orcid":false,"given":"Qing","family":"Zhu","sequence":"additional","affiliation":[{"name":"Ant Group, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0380-3506","authenticated-orcid":false,"given":"Jingling","family":"Xue","sequence":"additional","affiliation":[{"name":"UNSW, Sydney, Australia"},{"name":"Ant Group, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,10,9]]},"reference":[{"key":"e_1_3_2_2_1","doi-asserted-by":"publisher","unstructured":"Al Badawi Ahmad Bates Jack Bergamaschi Flavio Cousins David Bruce Erabelli Saroja Genise Nicholas Halevi Shai Hunt Hamish Kim Andrey Lee Yongwoo Liu Zeyu Micciancio Daniele Quah Ian Polyakov Yuriy R.V. Saraswathy Rohloff Kurt Saylor Jonathan Suponitsky Dmitriy Triplett Matthew Vaikuntanathan Vinod and Zucca Vincent. 2022. OpenFHE: Open-Source Fully Homomorphic Encryption Library. In Proceedings of the 10th Workshop on Encrypted Computing & Applied Homomorphic Cryptography (Los Angeles CA USA) (WAHC\u201922). Association for Computing Machinery Bates Jack NY USA 53\u201363. doi:10.1145\/3560827.3563379","DOI":"10.1145\/3560827.3563379"},{"key":"e_1_3_2_3_1","unstructured":"Martin R. Albrecht Melissa Chase Hao Chen Jintai Ding Shafi Goldwasser Sergey Gorbunov Shai Halevi Jeffrey Hoffstein Kim Laine Kristin E. Lauter Satya Lokam Daniele Micciancio Dustin Moody Travis Morrison Amit Sahai and Vinod Vaikuntanathan. 2019. Homomorphic Encryption Standard. IACR Cryptol. ePrint Arch. (2019) 939. https:\/\/eprint.iacr.org\/2019\/939"},{"key":"e_1_3_2_4_1","doi-asserted-by":"publisher","unstructured":"Fabian Boemer Anamaria Costache Rosario Cammarota and Casimir Wierzynski. 2019. nGraph-HE2: A High-Throughput Framework for Neural Network Inference on Encrypted Data. In Proceedings of the 7th ACM Workshop on Encrypted Computing & Applied Homomorphic Cryptography (London United Kingdom) (WAHC\u201919). Association for Computing Machinery Bates Jack NY USA 45\u201356. doi:10.1145\/3338469.3358944","DOI":"10.1145\/3338469.3358944"},{"key":"e_1_3_2_5_1","doi-asserted-by":"crossref","unstructured":"Jean-Philippe Bossuat Rosario Cammarota Jung Hee Cheon Ilaria Chillotti Benjamin R. Curtis Wei Dai Huijing Gong Erin Hales Duhyeong Kim Bryan Kumara Changmin Lee Xianhui Lu Carsten Maple Alberto Pedrouzo-Ulloa Rachel Player Luis Antonio Ruiz Lopez Yongsoo Song Donggeon Yhee and Bahattin Yildiz. 2011. Security Guidelines for Implementing Homomorphic Encryption. Cryptology ePrint Archive Paper 2024\/463. https:\/\/eprint.iacr.org\/2024\/463","DOI":"10.62056\/anxra69p1"},{"key":"e_1_3_2_6_1","unstructured":"Zvika Brakerski Craig Gentry and Vinod Vaikuntanathan. 2011. Fully Homomorphic Encryption without Bootstrapping. Cryptology ePrint Archive Paper 2011\/277. https:\/\/eprint.iacr.org\/2011\/277"},{"key":"e_1_3_2_7_1","doi-asserted-by":"crossref","unstructured":"Jung Hee Cheon Kyoohyung Han Andrey Kim Miran Kim and Yongsoo Song. 2019. A Full RNS Variant of Approximate Homomorphic Encryption. In Selected Areas in Cryptography \u2013 SAC 2018 Carlos Cid and Michael J. Jacobson Jr. (Eds.). Springer International Publishing Cham 347\u2013368.","DOI":"10.1007\/978-3-030-10970-7_16"},{"key":"e_1_3_2_8_1","doi-asserted-by":"publisher","unstructured":"Jung Hee Cheon Andrey Kim Miran Kim and Yongsoo Song. 2017. Homomorphic Encryption for Arithmetic of Approximate Numbers. In Advances in Cryptology - ASIACRYPT 2017. Springer International Publishing Cham 409\u2013437. doi:10.1007\/978-3-319-70694-8_15","DOI":"10.1007\/978-3-319-70694-8_15"},{"key":"e_1_3_2_9_1","unstructured":"Seonyoung Cheon Yongwoo Lee Dongkwan Kim Ju Min Lee Sunchul Jung Taekyung Kim Dongyoon Lee and Hanjun Kim. 2024. DaCapo: Automatic Bootstrapping Management for Efficient Fully Homomorphic Encryption. In 33rd USENIX Security Symposium (USENIX Security 24). USENIX Association Philadelphia PA 6993\u20137010. https:\/\/www.usenix.org\/conference\/usenixsecurity24\/presentation\/cheon"},{"key":"e_1_3_2_10_1","unstructured":"Ilaria Chillotti Nicolas Gama Mariya Georgieva and Malika Izabach\u00c3\u00a8ne. 2018. TFHE: Fast Fully Homomorphic Encryption over the Torus. Cryptology ePrint Archive Paper 2018\/421. https:\/\/eprint.iacr.org\/2018\/421"},{"key":"e_1_3_2_11_1","doi-asserted-by":"publisher","unstructured":"Roshan Dathathri Blagovesta Kostova Olli Saarikivi Wei Dai Kim Laine and Madanlal Musuvathi. 2020. EVA: An Encrypted Vector Arithmetic Language and Compiler for Efficient Homomorphic Computation. In Proceedings of the 41st ACM SIGPLAN Conference on Programming Language Design and Implementation. 546\u2013561. doi:10.1145\/3385412.3386023","DOI":"10.1145\/3385412.3386023"},{"key":"e_1_3_2_12_1","doi-asserted-by":"crossref","unstructured":"Roshan Dathathri Olli Saarikivi Hao Chen Kim Laine Kristin Lauter Saeed Maleki Madan Musuvathi and Todd Mytkowicz. 2019. CHET: An Optimizing Compiler for Fully-Homomorphic Neural-Network Inferencing. In PLDI 2019. ACM 142\u2013156. https:\/\/www.microsoft.com\/en-us\/research\/publication\/chet-an-optimizing-compiler-for-fully-homomorphic-neural-network-inferencing\/","DOI":"10.1145\/3314221.3314628"},{"key":"e_1_3_2_13_1","doi-asserted-by":"crossref","unstructured":"L\u00e9o Ducas and Daniele Micciancio. 2015. FHEW: Bootstrapping Homomorphic Encryption in Less Than a Second. In Advances in Cryptology \u2013 EUROCRYPT 2015 Elisabeth Oswald and Marc Fischlin (Eds.). Springer Berlin Heidelberg Berlin Heidelberg 617\u2013640.","DOI":"10.1007\/978-3-662-46800-5_24"},{"key":"e_1_3_2_14_1","doi-asserted-by":"publisher","unstructured":"Austin Ebel Karthik Garimella and Brandon Reagen. 2025. Orion: A Fully Homomorphic Encryption Framework for Deep Learning. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems Volume 2 (Rotterdam Netherlands) (ASPLOS \u201925). Association for Computing Machinery Bates Jack NY USA 734\u2013749. doi:10.1145\/3676641.3716008","DOI":"10.1145\/3676641.3716008"},{"key":"e_1_3_2_15_1","unstructured":"Junfeng Fan and Frederik Vercauteren. 2012. Somewhat Practical Fully Homomorphic Encryption. Cryptology ePrint Archive Paper 2012\/144. https:\/\/eprint.iacr.org\/2012\/144"},{"key":"e_1_3_2_16_1","doi-asserted-by":"publisher","unstructured":"S. Fan Z. Wang W. Xu R. Hou D. Meng and M. Zhang. 2023. TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPU. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE Computer Society Bates Jack CA USA 922\u2013934. doi:10.1109\/HPCA56546.2023.10071017","DOI":"10.1109\/HPCA56546.2023.10071017"},{"key":"e_1_3_2_17_1","doi-asserted-by":"crossref","unstructured":"Craig Gentry. 2009. A fully homomorphic encryption scheme. Ph. D. Dissertation. Stanford University.","DOI":"10.1145\/1536414.1536440"},{"key":"e_1_3_2_18_1","doi-asserted-by":"crossref","unstructured":"Craig Gentry Amit Sahai and Brent Waters. 2013. Homomorphic Encryption from Learning with Errors: ConceptuallySimpler Asymptotically-Faster Attribute-Based. In Advances in Cryptology \u2013 CRYPTO 2013 Ran Canetti and Juan A. Garay (Eds.). Springer Berlin Heidelberg Berlin Heidelberg 75\u201392.","DOI":"10.1007\/978-3-642-40041-4_5"},{"key":"e_1_3_2_19_1","doi-asserted-by":"publisher","unstructured":"Armin Gerami Monte Hoover Pranav S. Dulepet and Ramani Duraiswami. 2024. FAST: Factorizable Attention for Speeding up Transformers. CoRR abs\/2402.07901 (2024). arXiv:2402.07901 doi:10.48550\/ARXIV.2402.07901","DOI":"10.48550\/ARXIV.2402.07901"},{"key":"e_1_3_2_20_1","unstructured":"Ran Gilad-Bachrach Nathan Dowlin Kim Laine Kristin E. Lauter Michael Naehrig and John Wernsing. 2016. CryptoNets: Applying Neural Networks to Encrypted Data with High Throughput and Accuracy. In Proceedings of the 33nd International Conference on Machine Learning ICML 2016 Cousins David Bruce NY USA June 19\u201324 2016 (FMLR Workshop and Conference Proceedings Vol. 48) Maria-Florina Balcan and Kilian Q. Weinberger (Eds.). JMLR.org 201\u2013210. http:\/\/proceedings.mlr. press\/v48\/gilad-bachrach16.html"},{"key":"e_1_3_2_21_1","doi-asserted-by":"publisher","unstructured":"Gamze G\u00fcrsoy Eduardo Chielle Charlotte M. Brannon Michail Maniatakos and Mark Gerstein. 2020. Privacy-preserving genotype imputation with fully homomorphic encryption. bioRxiv (2020). doi:10.1101\/2020.05.29.124412","DOI":"10.1101\/2020.05.29.124412"},{"key":"e_1_3_2_22_1","doi-asserted-by":"publisher","unstructured":"Shai Halevi and Victor Shoup. 2018. Faster Homomorphic Linear Transformations in HElib. In Advances in Cryptology CRYPTO 2018: 38th Annual International Cryptology Conference Bates Jack CA USA August 19\u201323 2018 Proceedings Part I (Santa Barbara CA USA). Springer-Verlag Berlin Heidelberg 93\u2013120. doi:10.1007\/978-3-319-96884-1_4","DOI":"10.1007\/978-3-319-96884-1_4"},{"key":"e_1_3_2_23_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2015. Deep Residual Learning for Image Recognition. CoRR abs\/1512.03385 (2015). arXiv:1512.03385http:\/\/arxiv.org\/abs\/1512.03385"},{"key":"e_1_3_2_24_1","unstructured":"Andrew G. Howard Menglong Zhu Bo Chen Dmitry Kalenichenko Weijun Wang Tobias Weyand Marco Andreetto and Hartwig Adam. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv:1704.04861 [cs.CV] https:\/\/arxiv.org\/abs\/1704.04861"},{"key":"e_1_3_2_25_1","doi-asserted-by":"publisher","unstructured":"Siddharth Jayashankar Edward Chen Tom Tang Wenting Zheng and Dimitrios Skarlatos. 2025. Cinnamon: A Framework for Scale-Out Encrypted AI. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems Volume 1 (Rotterdam Netherlands) (ASPLOS \u201925). Association for Computing Machinery Bates Jack NY USA 133\u2013150. doi:10.1145\/3669940.3707260","DOI":"10.1145\/3669940.3707260"},{"key":"e_1_3_2_26_1","doi-asserted-by":"publisher","unstructured":"Yangqing Jia Evan Shelhamer Jeff Donahue Sergey Karayev Jonathan Long Ross Girshick Sergio Guadarrama and Trevor Darrell. 2014. Caffe: Convolutional Architecture for Fast Feature Embedding. In Proceedings of the 22nd ACM International Conference on Multimedia (Orlando Florida USA) (MM \u201914). Association for Computing Machinery Bates Jack NY USA 675\u2013678. doi:10.1145\/2647868.2654889","DOI":"10.1145\/2647868.2654889"},{"key":"e_1_3_2_27_1","unstructured":"Zhennan Qin Jianhui Li and Dan Lavery. 2024. oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation. In 2024 IEEE\/ACM International Symposium on Code Generation and Optimization (CGO). IEEE 193\u2013204."},{"key":"e_1_3_2_28_1","unstructured":"Chiraag Juvekar Vinod Vaikuntanathan and Anantha Chandrakasan. 2018. GAZELLE: a low latency framework for secure neural network inference. In Proceedings of the 27th USENIX Conference on Security Symposium (Baltimore MD USA) (SEC\u201918). USENIX Association USA 1651\u20131668."},{"key":"e_1_3_2_29_1","doi-asserted-by":"publisher","unstructured":"Aleksandar Krastev Nikola Samardzic Simon Langowski Srinivas Devadas and Daniel Sanchez. 2024. A Tensor Compiler with Automatic Data Packing for Simple and Efficient Fully Homomorphic Encryption. In Proceedings of the 45st ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI \u201924). ACM. doi:10.1145\/3656382","DOI":"10.1145\/3656382"},{"key":"e_1_3_2_30_1","unstructured":"Eunsang Lee Joon-Woo Lee Junghyun Lee Young-Sik Kim Yongjune Kim Jong-Seon No and Woosuk Choi. 2022b. Low-Complexity Deep Convolutional Neural Networks on Fully Homomorphic Encryption Using Multiplexed Parallel Convolutions. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research Vol. 162) Bates Jack Bergamaschi Flavio Bates Jack Bergamaschi Flavio Bates Jack and Sivan Sabato (Eds.). PMLR 12403\u201312422. https:\/\/proceedings.mlr.press\/v162\/lee22e.html"},{"key":"e_1_3_2_31_1","doi-asserted-by":"crossref","unstructured":"Eunsang Lee Joon-Woo Lee Jong-Seon No and Young-Sik Kim. 2021. Minimax Approximation of Sign Function by Composite Polynomial for Homomorphic Comparison. IEEE Transactions on Dependable and Secure Computing 19(6) 3711\u20133727.","DOI":"10.1109\/TDSC.2021.3105111"},{"key":"e_1_3_2_32_1","doi-asserted-by":"publisher","unstructured":"Joon-Woo Lee Hyungchul Kang Yongwoo Lee Woosuk Choi Jieun Eom Maxim Deryabin Eunsang Lee Junghyun Lee Donghoon Yoo Young-Sik Kim and Jong-Seon No. 2022a. Privacy-Preserving Machine Learning With Fully Homomorphic Encryption for Deep Neural Network. IEEE Access 10 (2022) 30039\u201330054. doi:10.1109\/ACCESS.2022.3159694","DOI":"10.1109\/ACCESS.2022.3159694"},{"key":"e_1_3_2_33_1","unstructured":"Long Li Jianxin Lai Peng Yuan Tianxiang Sui Yan Liu Qing Zhu Xiaojing Zhang Linjie Xiao Wenguang Chen and Jingling Xue. 2025. ANT-ACE: An FHE Compiler Framework for Automating Neural Network Inference. In 2025 IEEE\/ACM International Symposium on Code Generation and Optimization (CGO)."},{"key":"e_1_3_2_34_1","doi-asserted-by":"publisher","unstructured":"YanLiu Jianxin Lai Long Li Tianxiang Sui Linjie Xiao Peng Yuan Xiaojing Zhang Qing Zhu Wenguang Chen and Jingling Xue. 2025. ReSBM: Region-based Scale and Minimal-Level Bootstrapping Management for FHE via Min-Cut. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems Volume 1 (Rotterdam Netherlands) (ASPLOS \u201925). Association for Computing Machinery Bates Jack NY USA 924\u2013939. doi:10.1145\/3669940.3707276","DOI":"10.1145\/3669940.3707276"},{"key":"e_1_3_2_35_1","doi-asserted-by":"publisher","unstructured":"Guiwen Luo Shihui Fu and Guang Gong. 2023. Speeding Up Multi-Scalar Multiplication over Fixed Points Towards Efficient zkSNARKs. IACR Trans. Cryptogr. Hardw. Embed. Syst. 2023 2 (2023) 358\u2013380. doi:10.46586\/TCHES.V2023.I2.358-380","DOI":"10.46586\/TCHES.V2023.I2.358-380"},{"key":"e_1_3_2_36_1","unstructured":"Meta. 2024. Build the future of AI with Meta Llama 3. https:\/\/llama.meta.com\/llama3\/"},{"key":"e_1_3_2_37_1","doi-asserted-by":"publisher","unstructured":"Nikola Samardzic Axel Feldmann Aleksandar Krastev Srinivas Devadas Ronald Dreslinski Christopher Peikert and Daniel Sanchez. 2021. F1: A Fast and Programmable Accelerator for Fully Homomorphic Encryption. In MICRO-54: 54th Annual IEEE\/ACM International Symposium on Microarchitecture (Virtual Event Greece) (MICRO \u201921). Association for Computing Machinery Bates Jack NY USA 238\u2013252. doi:10.1145\/3466752.3480070","DOI":"10.1145\/3466752.3480070"},{"key":"e_1_3_2_38_1","doi-asserted-by":"publisher","unstructured":"Nikola Samardzic Axel Feldmann Aleksandar Krastev Nathan Manohar Nicholas Genise Srinivas Devadas Karim Eldefrawy Chris Peikert and Daniel Sanchez. 2022. CraterLake: a hardware accelerator for efficient unbounded computation on encrypted data. In Proceedings of the 49th Annual International Symposium on Computer Architecture (New York New York) (ISCA \u201922). Association for Computing Machinery Bates Jack NY USA 173\u2013187. doi:10.1145\/3470496.3527393","DOI":"10.1145\/3470496.3527393"},{"key":"e_1_3_2_39_1","unstructured":"SEAL2020. Microsoft SEAL (release 3.6). https:\/\/github.com\/Microsoft\/SEAL. Microsoft Research Redmond WA.."},{"key":"e_1_3_2_40_1","doi-asserted-by":"crossref","unstructured":"Halevi Shai and Shoup Victor. 2014. Algorithms in HElib. In Advances in Cryptology - CRYPTO 2014 Juan A. Garay and Rosario Gennaro (Eds.). Springer Berlin Heidelberg Berlin Heidelberg 554\u2013571.","DOI":"10.1007\/978-3-662-44371-2_31"},{"key":"e_1_3_2_41_1","unstructured":"Xing Su Xiangke Liao and Jingling Xue. 2017. Automatic Generation of Fast BLAS3-GEMM: A Portable Compiler Approach. In 2017 IEEE\/ACM International Symposium on Code Generation and Optimization (CGO). IEEE 193\u2013204."},{"key":"e_1_3_2_42_1","unstructured":"Alexander Viand Patrick Jattke Miro Haller and Anwar Hithnawi. 2022. HECO: Automatic Code Optimizations for Efficient Fully Homomorphic Encryption. (2022). https:\/\/arxiv.org\/abs\/2202.01649"},{"key":"e_1_3_2_43_1","doi-asserted-by":"publisher","unstructured":"Lee Yongwoo Heo Seonyeong Cheon Seonyoung Jeong Shinnung Kim Changsu Kim Eunkyung Lee Dongyoon and Kim Hanjun. 2022. HECATE: Performance-Aware Scale Optimization for Homomorphic Encryption Compiler. In 2022 IEEE\/ACM International Symposium on Code Generation and Optimization (CGO). 193\u2013204. doi:10.1109\/CGO53902.2022.9741265","DOI":"10.1109\/CGO53902.2022.9741265"},{"key":"e_1_3_2_44_1","doi-asserted-by":"publisher","unstructured":"Feng Yu Guangli Li Jiacheng Zhao Huimin Cui Xiaobing Feng and Jingling Xue. 2024. Optimizing Dynamic-Shape Neural Networks on Accelerators via On-the-Fly Micro-Kernel Polymerization. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems Volume 2 (La Jolla CA USA) (ASPLOS \u201924). Association for Computing Machinery Bates Jack NY USA 797\u2013812. doi:10.1145\/3620665.3640390","DOI":"10.1145\/3620665.3640390"},{"key":"e_1_3_2_45_1","doi-asserted-by":"publisher","unstructured":"Peng Yuan Yan Liu JianXin Lai Long Li Tianxiang Sui Linjie Xiao Xiaojing Zhang Qing Zhu and Jingling Xue. 2025. MetaKernel Artifact. doi:10.5281\/zenodo.16911192","DOI":"10.5281\/zenodo.16911192"},{"key":"e_1_3_2_46_1","doi-asserted-by":"publisher","unstructured":"Y. Zhai M. Ibrahim Y. Qiu F. Boemer Z. Chen A. Titov and A. Lyashevsky. 2022. Accelerating Encrypted Computing on Intel GPUs. In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE Computer Society Bates Jack CA USA 705\u2013716. doi:10.1109\/IPDPS53621.2022.00074","DOI":"10.1109\/IPDPS53621.2022.00074"},{"key":"e_1_3_2_47_1","doi-asserted-by":"publisher","unstructured":"ZhongchengZhang Ying Liu Yuyang Zhang Zhenchuan Chen Jiacheng Zhao Xiaobing Feng Huimin Cui and Jingling Xue. 2025. Qiwu: Exploiting Ciphertext-Level SIMD Parallelism in Homomorphic Encryption Programs. In Proceedings of the 23rd ACM\/IEEE International Symposium on Code Generation and Optimization (Las Vegas NV USA) (CGO \u201925). Association for Computing Machinery Bates Jack NY USA 523\u2013537. doi:10.1145\/3696443.3708917","DOI":"10.1145\/3696443.3708917"}],"container-title":["Proceedings of the ACM on Programming Languages"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3763095","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T10:12:50Z","timestamp":1784196770000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3763095"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,9]]},"references-count":46,"journal-issue":{"issue":"OOPSLA2","published-print":{"date-parts":[[2025,10,9]]}},"alternative-id":["10.1145\/3763095"],"URL":"https:\/\/doi.org\/10.1145\/3763095","relation":{},"ISSN":["2475-1421"],"issn-type":[{"value":"2475-1421","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,9]]},"assertion":[{"value":"2025-03-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-12","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}