{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T14:30:52Z","timestamp":1787495452971,"version":"build-2736575974"},"publisher-location":"Singapore","reference-count":28,"publisher":"Springer Nature Singapore","isbn-type":[{"value":"9789819248049","type":"print"},{"value":"9789819248056","type":"electronic"}],"license":[{"start":{"date-parts":[[2026,8,24]],"date-time":"2026-08-24T00:00:00Z","timestamp":1787529600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2026,8,24]],"date-time":"2026-08-24T00:00:00Z","timestamp":1787529600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2027]]},"DOI":"10.1007\/978-981-92-4805-6_5","type":"book-chapter","created":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T13:47:12Z","timestamp":1787492832000},"page":"67-81","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["JSQKV: Joint Sparsification and\u00a0Quantization for\u00a0KV-Cache Compression and\u00a0Decode Acceleration"],"prefix":"10.1007","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-1044-0312","authenticated-orcid":false,"given":"Hao","family":"Zhang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9836-558X","authenticated-orcid":false,"given":"Xiaoli","family":"Gong","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haoran","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huayou","family":"Su","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qingxia","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9086-1178","authenticated-orcid":false,"given":"Jin","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,8,24]]},"reference":[{"key":"5_CR1","unstructured":"Agrawal, A., Kedia, N., Panwar, A., et al.: Taming throughput-latency tradeoff in LLM inference with sarathi-serve. arXiv preprint arXiv:2403.02310 (2024). https:\/\/arxiv.org\/abs\/2403.02310"},{"key":"5_CR2","doi-asserted-by":"crossref","unstructured":"Ashkboos, S., et al.: Quarot: outlier-free 4-bit inference in rotated LLMs. arXiv preprint arXiv:2404.00456 (2024)","DOI":"10.52202\/079017-3180"},{"key":"5_CR3","unstructured":"Bai, Y., et al.: Longbench: a bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508 (2023). https:\/\/arxiv.org\/abs\/2308.14508"},{"key":"5_CR4","unstructured":"Cai, Z., et al.: Pyramidkv: dynamic KV cache compression based on pyramidal information funneling. arXiv preprint arXiv:2406.02069 (2024). https:\/\/arxiv.org\/abs\/2406.02069"},{"key":"5_CR5","unstructured":"Chen, Z., et al.: Rotatekv: accurate and robust 2-bit kv cache quantization for LLMs via outlier-aware adaptive rotations (2025). https:\/\/arxiv.org\/abs\/2501.16383"},{"key":"5_CR6","unstructured":"Cobbe, K., et al.: Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168 (2021). https:\/\/arxiv.org\/abs\/2110.14168"},{"key":"5_CR7","unstructured":"Ding, Y., et al.: Longrope: extending LLM context window beyond 2 million tokens. arXiv preprint arXiv:2402.13753 (2024). https:\/\/arxiv.org\/abs\/2402.13753"},{"key":"5_CR8","unstructured":"Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A.: The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)"},{"key":"5_CR9","unstructured":"Hong, K., et al.: FlashDecoding++: faster large language model inference with asynchronization, flat GEMM optimization, and heuristics. In: Proceedings of Machine Learning and Systems (2024)"},{"key":"5_CR10","doi-asserted-by":"crossref","unstructured":"Hooper, C., et al.: Kvquant: towards 10 million context length LLM inference with kv cache quantization. arXiv preprint arXiv:2401.18079 (2024)","DOI":"10.52202\/079017-0040"},{"key":"5_CR11","unstructured":"Jiang, A.Q., et al.: Mistral 7b. arXiv preprint arXiv:2310.06825 (2023). https:\/\/arxiv.org\/abs\/2310.06825"},{"key":"5_CR12","doi-asserted-by":"publisher","unstructured":"Joo, D., Hosseini, H., Hadidi, R., Asgari, B.: Coruscant: co-designing GPU kernel and sparse tensor core to advocate unstructured sparsity in efficient LLM inference. In: Proceedings of the 58th IEEE\/ACM International Symposium on Microarchitecture, pp. 232\u2013245 (2025). https:\/\/doi.org\/10.1145\/3725843.3756065","DOI":"10.1145\/3725843.3756065"},{"key":"5_CR13","doi-asserted-by":"crossref","unstructured":"Joo, D., Hosseini, H., Hadidi, R., Asgari, B.: Mustafar: promoting unstructured sparsity for kv cache pruning in LLM inference. arXiv preprint arXiv:2505.22913 (2025). https:\/\/arxiv.org\/abs\/2505.22913","DOI":"10.52202\/085713-2564"},{"key":"5_CR14","doi-asserted-by":"crossref","unstructured":"Kwon, W., et al.: Efficient memory management for large language model serving with pagedattention. In: Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles (2023). https:\/\/arxiv.org\/abs\/2309.06180","DOI":"10.1145\/3600006.3613165"},{"key":"5_CR15","unstructured":"Li, Y., et al.: Snapkv: LLM knows what you are looking for before generation. arXiv preprint arXiv:2404.14469 (2024). https:\/\/arxiv.org\/abs\/2404.14469, version 2, June 17, 2024"},{"key":"5_CR16","unstructured":"Liu, Z., et al.: Kivi: a tuning-free asymmetric 2bit quantization for kv cache. arXiv preprint arXiv:2402.02750 (2024)"},{"key":"5_CR17","unstructured":"Lv, B., Zhou, Q., Ding, X., Wang, Y., Ma, Z.: Kvpruner: structural pruning for faster and memory-efficient large language models. arXiv preprint arXiv:2409.11057 (2024). https:\/\/arxiv.org\/abs\/2409.11057"},{"key":"5_CR18","unstructured":"Merity, S., Xiong, C., Bradbury, J., Socher, R.: Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843 (2016). https:\/\/arxiv.org\/abs\/1609.07843"},{"key":"5_CR19","doi-asserted-by":"crossref","unstructured":"Singhania, P., Singh, S., He, S., et al.: Loki: low-rank keys for efficient sparse attention. arXiv preprint arXiv:2406.02542 (2024). https:\/\/arxiv.org\/abs\/2406.02542","DOI":"10.52202\/079017-0532"},{"key":"5_CR20","unstructured":"Touvron, H., et al.: Llama 2: open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023). https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"5_CR21","doi-asserted-by":"crossref","unstructured":"Trivedi, H., Balasubramanian, N., Khot, T., Sabharwal, A.: Musique: multihop questions via single-hop question composition. Trans. Assoc. Comput. Linguist. 10, 539\u2013554 (2022). https:\/\/aclanthology.org\/2022.tacl-1.31\/","DOI":"10.1162\/tacl_a_00475"},{"key":"5_CR22","unstructured":"Xia, H., et al.: Flash-LLM: enabling cost-effective and highly-efficient large generative model inference with unstructured sparsity. arXiv preprint arXiv:2309.10285 (2023). https:\/\/arxiv.org\/abs\/2309.10285"},{"key":"5_CR23","unstructured":"Xiao, G., Tian, Y., Chen, B., Han, S., Lewis, M.: Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453 (2024). https:\/\/arxiv.org\/abs\/2309.17453"},{"key":"5_CR24","unstructured":"Xu, Y., et al.: Think: thinner key cache by query-driven pruning. In: International Conference on Learning Representations (2025). https:\/\/arxiv.org\/abs\/2407.21018, spotlight"},{"key":"5_CR25","unstructured":"Yang, A., et al.: Qwen2.5 technical report. arXiv preprint arXiv:2412.15115 (2024). https:\/\/arxiv.org\/abs\/2412.15115"},{"key":"5_CR26","unstructured":"Ye, Z., et al.: Flashinfer: efficient and customizable attention engine for LLM inference serving. In: Proceedings of Machine Learning and Systems (2025)"},{"key":"5_CR27","unstructured":"Zhang, Z., et al.: H$$_2$$o: heavy-hitter oracle for efficient generative inference of large language models. arXiv preprint arXiv:2306.14048 (2023). https:\/\/arxiv.org\/abs\/2306.14048"},{"key":"5_CR28","unstructured":"Zhong, Y., Liu, S., Chen, J., et al.: Distserve: disaggregating prefill and decoding for goodput-optimized large language model serving. In: 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2024) (2024). https:\/\/www.usenix.org\/conference\/osdi24\/presentation\/zhong-yinmin"}],"container-title":["Lecture Notes in Computer Science","Advanced Parallel Processing Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-981-92-4805-6_5","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T13:47:16Z","timestamp":1787492836000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-981-92-4805-6_5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8,24]]},"ISBN":["9789819248049","9789819248056"],"references-count":28,"URL":"https:\/\/doi.org\/10.1007\/978-981-92-4805-6_5","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"value":"0302-9743","type":"print"},{"value":"1611-3349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,8,24]]},"assertion":[{"value":"24 August 2026","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"The authors have no competing interests to declare that are relevant to the content of this article.","order":1,"name":"Ethics","label":"Disclosure of Interests","group":{"name":"EthicsHeading","label":"Ethics"}},{"value":"The technical solution and experimental design presented in this paper were independently completed by the authors. AI tools were used solely for language polishing\u00a0and formatting optimization, and did not participate in the development of the research ideas or core content.","order":2,"name":"Ethics","label":"Use of Artificial Intelligence Tools","group":{"name":"EthicsHeading","label":"Ethics"}},{"value":"APPT","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"International Symposium on Advanced Parallel Processing Technologies","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Brussels","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Belgium","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2026","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"27 July 2026","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"27 July 2026","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"17","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"appt2026","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/www.appt-conference.com\/2026","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}