{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T04:56:37Z","timestamp":1783572997976,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":52,"publisher":"ACM","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,2,22]]},"DOI":"10.1145\/3748173.3779188","type":"proceedings-article","created":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T21:17:35Z","timestamp":1770326255000},"page":"56-66","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-6815-8297","authenticated-orcid":false,"given":"Dong","family":"Liu","sequence":"first","affiliation":[{"name":"Yale University, New Haven, Connecticut, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7096-2438","authenticated-orcid":false,"given":"Yanxuan","family":"Yu","sequence":"additional","affiliation":[{"name":"Columbia University, New York, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,2,21]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Marcos K Aguilera Nadav Amit William J Bolosky Atul Chaugule Camille Coppens J\u00e9r\u00e9mie Duchesne Chaoran Guo et al. 2019. Remote regions: a simple abstraction for remote memory. arXiv preprint arXiv:1908.06189 (2019)."},{"key":"e_1_3_2_1_2_1","volume-title":"The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783","author":"Meta AI.","year":"2024","unstructured":"Meta AI. 2024. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783 (2024)."},{"key":"e_1_3_2_1_3_1","volume-title":"GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.","author":"Ainslie Joshua","year":"2023","unstructured":"Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebr\u00f3n, and Sumit Sanghai. 2023. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. (2023). arXiv:2305.13245 [cs.CL] https:\/\/arxiv.org\/abs\/2305.13245"},{"key":"e_1_3_2_1_4_1","volume-title":"Cheng Li, Du Li, Elton Zheng, Jeff Rasley, Shaden Smith, Olatunji Ruwase, et al.","author":"Aminabadi Reza Yazdani","year":"2022","unstructured":"Reza Yazdani Aminabadi, Samyam Rajbhandari, Minjia Zhang, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Jeff Rasley, Shaden Smith, Olatunji Ruwase, et al., 2022. Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale. arXiv preprint arXiv:2207.00032 (2022)."},{"key":"e_1_3_2_1_5_1","unstructured":"Jinze Bai Shuai Bai Yunfei Chu Zeyu Cui Kai Dang Xiaodong Deng Yang Fan Wenbin Ge Yu Han Fei Huang et al. 2023. Qwen Technical Report. arXiv preprint arXiv:2309.16609 (2023)."},{"key":"e_1_3_2_1_6_1","volume-title":"Proceedings of the 2021 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays. 102-113","author":"Boyd-Wickizer Silas","year":"2021","unstructured":"Silas Boyd-Wickizer and Jintao Zhai. 2021. Memory controller design for FPGA-based systems. In Proceedings of the 2021 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays. 102-113."},{"key":"e_1_3_2_1_7_1","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. Advances in neural information processing systems Vol. 33 (2020) 1877-1901."},{"key":"e_1_3_2_1_8_1","volume-title":"Medusa: Simple llm inference acceleration framework with multiple decoding heads. arXiv preprint arXiv:2401.10774","author":"Cai Tianle","year":"2024","unstructured":"Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao. 2024. Medusa: Simple llm inference acceleration framework with multiple decoding heads. arXiv preprint arXiv:2401.10774 (2024)."},{"key":"e_1_3_2_1_9_1","volume-title":"Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al.","author":"Chen Mark","year":"2021","unstructured":"Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al., 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)."},{"key":"e_1_3_2_1_10_1","unstructured":"ShareGPT Community. 2023. ShareGPT: Share your wildest ChatGPT conversations. https:\/\/sharegpt.com."},{"key":"e_1_3_2_1_11_1","unstructured":"CXL Consortium. 2022. Compute Express Link Specification Revision 2.0."},{"key":"e_1_3_2_1_12_1","unstructured":"Intel Corporation. 2024. Intel Agilex 7 FPGA Architecture. (2024)."},{"key":"e_1_3_2_1_13_1","first-page":"287","volume-title":"2023 USENIX Annual Technical Conference (USENIX ATC 23)","author":"Gouk Donghyun","year":"2023","unstructured":"Donghyun Gouk, Sangwon Lee, Miryeong Kwon, and Myoungsoo Shin. 2023. Direct access, high-performance memory disaggregation with directcxl. In 2023 USENIX Annual Technical Conference (USENIX ATC 23). 287-294."},{"key":"e_1_3_2_1_14_1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1-14","author":"Ham Tae Jun","year":"2020","unstructured":"Tae Jun Ham, Yejin Lee, et al., 2020. FPGA-based hardware acceleration of transformer attention modules. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1-14."},{"key":"e_1_3_2_1_15_1","volume-title":"International Conference on Machine Learning. 1919-1928","author":"Hashemi Milad","year":"2018","unstructured":"Milad Hashemi, Kevin Swersky, Jamie A Smith, Grant Ayers, Heiner Litz, Jichuan Chang, Christos Kozyrakis, and Parthasarathy Ranganathan. 2018. Learning memory access patterns. In International Conference on Machine Learning. 1919-1928."},{"key":"e_1_3_2_1_16_1","volume-title":"Kurt Keutzer, and Amir Gholami.","author":"Hooper Coleman","year":"2024","unstructured":"Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. 2024. Kvquant: Towards 10 million context length llm inference with kv cache quantization. arXiv preprint arXiv:2401.18079 (2024)."},{"key":"e_1_3_2_1_17_1","unstructured":"Junhyeok Jang Hanjin Park Jinsoo Choi Jaehoon Lee Suhwan Lee et al. 2023. CXL-ANNS: Software-Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search. arXiv preprint arXiv:2305.15838 (2023)."},{"key":"e_1_3_2_1_18_1","volume-title":"Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, et al.","author":"Jiang Albert Q","year":"2023","unstructured":"Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, et al., 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_3_2_1_20_1","volume-title":"Fast inference from transformers via speculative decoding. arXiv preprint arXiv:2211.17192","author":"Leviathan Yaniv","year":"2023","unstructured":"Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. arXiv preprint arXiv:2211.17192 (2023)."},{"key":"e_1_3_2_1_21_1","unstructured":"Zhe Li Yuxuan Yang Yiqi Xing Zhipeng Zhang Qingyang Wang et al. 2024. CXL-MEM: Enabling Cost-Effective and Performant Memory Expansion for Deep Learning. arXiv preprint arXiv:2401.13841 (2024)."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555815.1555789"},{"key":"e_1_3_2_1_23_1","unstructured":"Aixin Liu Bei Feng Bin Wang Bingxuan Wang Bo Liu Chenggang Zhao Chengqi Dengr Chong Ruan Damai Dai Daya Guo et al. 2024a. Deepseek-v2: A strong economical and efficient mixture-of-experts language model. arXiv preprint arXiv:2405.04434 (2024)."},{"key":"e_1_3_2_1_24_1","unstructured":"Dong Liu and Yanxuan Yu. 2025a. LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference. (2025). arXiv:2406.19657 [cs.LG] https:\/\/arxiv.org\/abs\/2406.19657"},{"key":"e_1_3_2_1_25_1","unstructured":"Dong Liu and Yanxuan Yu. 2025b. \u03c0-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling. (2025). arXiv:2511.10696 [cs.CL] https:\/\/arxiv.org\/abs\/2511.10696"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3746027.3758181"},{"key":"e_1_3_2_1_27_1","volume-title":"CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference. arXiv preprint arXiv:2511.21702","author":"Liu Dong","year":"2025","unstructured":"Dong Liu, Yanxuan Yu, and Ben Lengerich. 2025. CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference. arXiv preprint arXiv:2511.21702 (2025)."},{"key":"e_1_3_2_1_28_1","volume-title":"MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning. In ICML 2025 Workshop on Long-Context Foundation Models.","author":"Liu Dong","unstructured":"Dong Liu, Yanxuan Yu, Xuhong Wang, Ben Lengerich, and Ying Nian Wu. [n.d.]. MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning. In ICML 2025 Workshop on Long-Context Foundation Models."},{"key":"e_1_3_2_1_29_1","volume-title":"Kivi: A tuning-free asymmetric 2bit quantization for kv cache. arXiv preprint arXiv:2402.02750","author":"Liu Zirui","year":"2024","unstructured":"Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu. 2024b. Kivi: A tuning-free asymmetric 2bit quantization for kv cache. arXiv preprint arXiv:2402.02750 (2024)."},{"key":"e_1_3_2_1_30_1","volume-title":"Specinfer: Accelerating generative llm serving with speculative inference and token tree verification. arXiv preprint arXiv:2305.09781","author":"Miao Xupeng","year":"2024","unstructured":"Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Wong, Yaohui Zhu, and Zhihao Jia. 2024. Specinfer: Accelerating generative llm serving with speculative inference and token tree verification. arXiv preprint arXiv:2305.09781 (2024)."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/K16-1028"},{"key":"e_1_3_2_1_32_1","unstructured":"NVIDIA. 2023. FasterTransformer: Transformer Inference Library. (2023)."},{"key":"e_1_3_2_1_33_1","volume-title":"GPT-4 technical report. arXiv preprint arXiv:2303.08774","author":"AI.","year":"2023","unstructured":"OpenAI. 2023. GPT-4 technical report. arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_3_2_1_34_1","volume-title":"2020 IEEE International Conference on Big Data (Big Data). 1533-1542","author":"Patel Tirthak","year":"2020","unstructured":"Tirthak Patel, Sunita Potluri, and Devesh Tiwari. 2020. Prefetching for CNNs: Exploiting spatial and temporal locality in modern DNN inference. In 2020 IEEE International Conference on Big Data (Big Data). 1533-1542."},{"key":"e_1_3_2_1_35_1","volume-title":"FPGA Acceleration of Transformer-based Large Language Models. In 2023 IEEE 31st Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM). 84-94","author":"Peng Hongyi","year":"2023","unstructured":"Hongyi Peng, Sitao Huang, Tong Zhou, et al., 2023. FPGA Acceleration of Transformer-based Large Language Models. In 2023 IEEE 31st Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM). 84-94."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2847263.2847265"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_3_2_1_38_1","volume-title":"Yossi Adi, Jingyu Liu, Tal Remez, J\u00e9r\u00e9my Rapin, et al.","author":"Rozi\u00e8re Baptiste","year":"2023","unstructured":"Baptiste Rozi\u00e8re, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, J\u00e9r\u00e9my Rapin, et al., 2023. Code Llama: Open Foundation Models for Code. arXiv preprint arXiv:2308.12950 (2023)."},{"key":"e_1_3_2_1_39_1","volume-title":"Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. 574-587","author":"Ruan Zheng","year":"2023","unstructured":"Zheng Ruan, Malte Schwarzkopf, Marcos K Aguilera, and Adam Belay. 2023. POND: CXL-Based Memory Pooling Systems for Cloud Platforms. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. 574-587."},{"key":"e_1_3_2_1_40_1","first-page":"69","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Shan Yizhou","year":"2018","unstructured":"Yizhou Shan, Yutong Huang, Yilun Chen, and Yiying Zhang. 2018. LegoOS: A disseminated, distributed OS for hardware resource disaggregation. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 69-87."},{"key":"e_1_3_2_1_41_1","volume-title":"Flexgen: High-throughput generative inference of large language models with a single gpu. arXiv preprint arXiv:2303.06865","author":"Sheng Ying","year":"2023","unstructured":"Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher R\u00e9, Ion Stoica, and Ce Zhang. 2023. Flexgen: High-throughput generative inference of large language models with a single gpu. arXiv preprint arXiv:2303.06865 (2023)."},{"key":"e_1_3_2_1_42_1","first-page":"1145","article-title":"FPGA-based memory controllers for non-volatile memory systems","volume":"69","author":"Swamy Shriram","year":"2020","unstructured":"Shriram Swamy, Jian Liu, et al., 2020. FPGA-based memory controllers for non-volatile memory systems. IEEE Trans. Comput., Vol. 69, 8 (2020), 1145-1158.","journal-title":"IEEE Trans. Comput."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3706628.3708867"},{"key":"e_1_3_2_1_44_1","volume-title":"Gemma: Open Models Based on Gemini Research and Technology. arXiv preprint arXiv:2403.08295","author":"Team Gemma","year":"2024","unstructured":"Gemma Team. 2024a. Gemma: Open Models Based on Gemini Research and Technology. arXiv preprint arXiv:2403.08295 (2024)."},{"key":"e_1_3_2_1_45_1","volume-title":"Powering Code Intelligence at Scale. arXiv preprint","author":"Team Qwen","year":"2024","unstructured":"Qwen Team. 2024b. CodeQwen1.5: Powering Code Intelligence at Scale. arXiv preprint (2024)."},{"key":"e_1_3_2_1_46_1","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)."},{"key":"e_1_3_2_1_47_1","volume-title":"Smoothquant: Accurate and efficient post-training quantization for large language models. arXiv preprint arXiv:2211.10438","author":"Xiao Guangxuan","year":"2023","unstructured":"Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023a. Smoothquant: Accurate and efficient post-training quantization for large language models. arXiv preprint arXiv:2211.10438 (2023)."},{"key":"e_1_3_2_1_48_1","volume-title":"Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453","author":"Xiao Guangxuan","year":"2023","unstructured":"Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2023b. Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453 (2023)."},{"key":"e_1_3_2_1_49_1","unstructured":"AMD Xilinx. 2023. Versal AI Core Series Architecture Manual. (2023)."},{"key":"e_1_3_2_1_50_1","first-page":"521","volume-title":"16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Yu Gyeong-In","year":"2022","unstructured":"Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022. Orca: A distributed serving system for Transformer-Based generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). 521-538."},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"e_1_3_2_1_52_1","unstructured":"Zhenyu Zhang Ying Sheng Tianyi Zhou Tianlong Chen Lianmin Zheng Ruisi Cai Zhao Song Yuandong Tian Christopher R\u00e9 Clark Barrett et al. 2023. H2O: Heavy-hitter oracle for efficient generative inference of large language models. arXiv preprint arXiv:2306.14048 (2023)."}],"event":{"name":"FPGA '26:The 2026 ACM\/SIGDA International Symposium on Field Programmable Gate Arrays","location":"Seaside CA USA","sponsor":["SIGDA ACM Special Interest Group on Design Automation"]},"container-title":["Proceedings of the 2026 ACM\/SIGDA International Symposium on Field Programmable Gate Arrays"],"original-title":[],"deposited":{"date-parts":[[2026,2,9]],"date-time":"2026-02-09T16:17:39Z","timestamp":1770653859000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3748173.3779188"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,21]]},"references-count":52,"alternative-id":["10.1145\/3748173.3779188","10.1145\/3748173"],"URL":"https:\/\/doi.org\/10.1145\/3748173.3779188","relation":{},"subject":[],"published":{"date-parts":[[2026,2,21]]},"assertion":[{"value":"2026-02-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}