{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T16:51:58Z","timestamp":1784652718769,"version":"3.55.0"},"reference-count":47,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T00:00:00Z","timestamp":1732147200000},"content-version":"vor","delay-in-days":325,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,11,18]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Recent large language model applications, such as Retrieval-Augmented Generation and chatbots, have led to an increased need to process longer input contexts. However, this requirement is hampered by inherent limitations. Architecturally, models are constrained by a context window defined during training. Additionally, processing extensive texts requires substantial GPU memory. We propose a novel approach, Finch, to compress the input context by leveraging the pre-trained model weights of the self-attention. Given a prompt and a long text, Finch iteratively identifies the most relevant Key (K) and Value (V) pairs over chunks of the text conditioned on the prompt. Only such pairs are stored in the KV cache, which, within the space constrained by the context window, ultimately contains a compressed version of the long text. Our proposal enables models to consume large inputs even with high compression (up to 93x) while preserving semantic integrity without the need for fine-tuning.<\/jats:p>","DOI":"10.1162\/tacl_a_00716","type":"journal-article","created":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T19:15:57Z","timestamp":1732216557000},"page":"1517-1532","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":4,"title":["FINCH: Prompt-guided Key-Value Cache Compression for Large Language Models"],"prefix":"10.1162","volume":"12","author":[{"given":"Giulio","family":"Corallo","sequence":"first","affiliation":[{"name":"SAP Labs, France, EURECOM, France. giulio.corallo@sap.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Paolo","family":"Papotti","sequence":"additional","affiliation":[{"name":"EURECOM, France. papotti@eurecom.fr"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,11,18]]},"reference":[{"key":"2024112119155177300_bib1","doi-asserted-by":"publisher","first-page":"227","DOI":"10.1162\/tacl_a_00544","article-title":"Transformers for tabular data representation: A survey of models and applications","volume":"11","author":"Badaro","year":"2023","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024112119155177300_bib2","doi-asserted-by":"publisher","first-page":"3119","DOI":"10.18653\/v1\/2024.acl-long.172","article-title":"LongBench: A bilingual, multitask benchmark for long context understanding","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Bai","year":"2024"},{"key":"2024112119155177300_bib3","article-title":"Memory GAPS: Would LLM pass the Tulving Test?","author":"Chauvet","year":"2024","journal-title":"arXiv:2402.16505"},{"key":"2024112119155177300_bib4","doi-asserted-by":"publisher","first-page":"276","DOI":"10.18653\/v1\/W19-4828","article-title":"What does BERT look at? An analysis of BERT\u2019s attention","volume-title":"Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Clark","year":"2019"},{"key":"2024112119155177300_bib5","article-title":"Finch: Prompt-guided key-value cache compression","author":"Corallo","year":"2024","journal-title":"arXiv:2408.00167"},{"key":"2024112119155177300_bib6","first-page":"4271","article-title":"Funnel-Transformer: Filtering out sequential redundancy for efficient language processing","volume-title":"Advances in Neural Information Processing Systems","author":"Dai","year":"2020"},{"key":"2024112119155177300_bib7","article-title":"Language modeling is compression","volume-title":"The Twelfth International Conference on Learning Representations","author":"Deletang","year":"2024"},{"key":"2024112119155177300_bib8","article-title":"GPT3.int8(): 8-bit matrix multiplication for transformers at scale","volume-title":"Advances in Neural Information Processing Systems","author":"Dettmers","year":"2022"},{"key":"2024112119155177300_bib9","article-title":"QLoRA: Efficient finetuning of quantized LLMs","volume-title":"Thirty-seventh Conference on Neural Information Processing Systems","author":"Dettmers","year":"2023"},{"key":"2024112119155177300_bib10","article-title":"A survey on in-context learning","author":"Dong","year":"2022","journal-title":"arXiv:2301.00234"},{"key":"2024112119155177300_bib11","first-page":"10323","article-title":"SparseGPT: Massive language models can be accurately pruned in one-shot","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Frantar","year":"2023"},{"key":"2024112119155177300_bib12","article-title":"OPTQ: Accurate quantization for generative pre-trained transformers","volume-title":"The Eleventh International Conference on Learning Representations","author":"Frantar","year":"2023"},{"key":"2024112119155177300_bib13","article-title":"Model tells you what to discard: Adaptive KV cache compression for llMs","volume-title":"The Twelfth International Conference on Learning Representations","author":"Ge","year":"2024"},{"key":"2024112119155177300_bib14","article-title":"In-context autoencoder for context compression in a large language model","volume-title":"The Twelfth International Conference on Learning Representations","author":"Ge","year":"2024"},{"key":"2024112119155177300_bib15","first-page":"3690","article-title":"PoWER-BERT: Accelerating BERT inference via progressive word-vector elimination","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Goyal","year":"2020"},{"key":"2024112119155177300_bib16","doi-asserted-by":"publisher","first-page":"7275","DOI":"10.18653\/v1\/2022.acl-long.502","article-title":"Transkimmer: Transformer learns to layer-wise skim","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Guan","year":"2022"},{"key":"2024112119155177300_bib17","doi-asserted-by":"publisher","first-page":"3991","DOI":"10.18653\/v1\/2024.naacl-long.222","article-title":"LM-Infinite: Zero-shot extreme length generalization for large language models","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Han","year":"2024"},{"key":"2024112119155177300_bib18","doi-asserted-by":"publisher","first-page":"8798","DOI":"10.18653\/v1\/2022.acl-long.602","article-title":"Pyramid-BERT: Reducing complexity via successive core-set based token selection","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Huang","year":"2022"},{"key":"2024112119155177300_bib19","article-title":"Mistral 7b","author":"Jiang","year":"2023","journal-title":"arXiv:2310.06825"},{"key":"2024112119155177300_bib20","doi-asserted-by":"publisher","first-page":"13358","DOI":"10.18653\/v1\/2023.emnlp-main.825","article-title":"LLMLingua: Compressing prompts for accelerated inference of large language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Jiang","year":"2023"},{"key":"2024112119155177300_bib21","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.91","article-title":"LongLLMLingua: Accelerating and enhancing LLMs in long context scenarios via prompt compression","volume-title":"ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models","author":"Jiang","year":"2024"},{"key":"2024112119155177300_bib22","article-title":"Learning to remember rare events","volume-title":"International Conference on Learning Representations","author":"Kaiser","year":"2017"},{"key":"2024112119155177300_bib23","doi-asserted-by":"publisher","first-page":"784","DOI":"10.1145\/3534678.3539260","article-title":"Learned token pruning for transformers","volume-title":"Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Kim","year":"2022"},{"key":"2024112119155177300_bib24","article-title":"Retrieval-augmented generation for knowledge-intensive NLP tasks","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Lewis","year":"2020"},{"key":"2024112119155177300_bib25","article-title":"Textbooks are all you need II: phi-1.5 technical report","author":"Li","year":"2023","journal-title":"arXiv:2309.05463"},{"key":"2024112119155177300_bib26","article-title":"Unlocking context constraints of LLMs: Enhancing context efficiency of LLMs with self-information-based content filtering","author":"Li","year":"2023","journal-title":"arXiv:2304.12102"},{"key":"2024112119155177300_bib27","article-title":"Ring attention with blockwise transformers for near-infinite context","volume-title":"NeurIPS 2023 Foundation Models for Decision Making Workshop","author":"Liu","year":"2023"},{"key":"2024112119155177300_bib28","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1162\/tacl_a_00638","article-title":"Lost in the middle: How language models use long contexts","volume":"12","author":"Liu","year":"2024","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024112119155177300_bib29","article-title":"Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time","volume-title":"Thirty-seventh Conference on Neural Information Processing Systems","author":"Liu","year":"2023"},{"key":"2024112119155177300_bib30","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/2022.acl-long.1","article-title":"AdapLeR: Speeding up inference by adaptive length reduction","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Modarressi","year":"2022"},{"key":"2024112119155177300_bib31","article-title":"Learning to compress prompts with gist tokens","volume-title":"Thirty-seventh Conference on Neural Information Processing Systems","author":"Jesse","year":"2023"},{"key":"2024112119155177300_bib32","article-title":"Transformers are multi-state RNNs","author":"Oren","year":"2024","journal-title":"arXiv:2401.06104"},{"key":"2024112119155177300_bib33","article-title":"Training language models to follow instructions with human feedback","volume-title":"Advances in Neural Information Processing Systems","author":"Ouyang","year":"2022"},{"key":"2024112119155177300_bib34","doi-asserted-by":"publisher","first-page":"784","DOI":"10.18653\/v1\/P18-2124","article-title":"Know what you don\u2019t know: Unanswerable questions for SQuAD","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Rajpurkar","year":"2018"},{"key":"2024112119155177300_bib35","doi-asserted-by":"publisher","first-page":"3982","DOI":"10.18653\/v1\/D19-1410","article-title":"Sentence-BERT: Sentence embeddings using Siamese BERT-networks","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Reimers","year":"2019"},{"key":"2024112119155177300_bib36","article-title":"On the efficacy of eviction policy for key-value constrained generative language model inference","author":"Ren","year":"2024","journal-title":"arXiv:2402.06262"},{"key":"2024112119155177300_bib37","doi-asserted-by":"publisher","first-page":"2210","DOI":"10.18653\/v1\/D17-1235","article-title":"Generating high-quality and informative conversation responses with sequence-to- sequence models","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing","author":"Shao","year":"2017"},{"key":"2024112119155177300_bib38","doi-asserted-by":"publisher","first-page":"127063","DOI":"10.1016\/j.neucom.2023.127063","article-title":"RoFormer: Enhanced transformer with rotary position embedding","volume":"568","author":"Jianlin","year":"2024","journal-title":"Neurocomputing"},{"key":"2024112119155177300_bib39","article-title":"Llama: Open and efficient foundation language models","author":"Touvron","year":"2023","journal-title":"arXiv:2302.13971"},{"key":"2024112119155177300_bib40","article-title":"Llama 2: Open foundation and fine-tuned chat models","author":"Touvron","year":"2023","journal-title":"arXiv:2307.09288"},{"key":"2024112119155177300_bib41","article-title":"Attention is all you need","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani","year":"2017"},{"key":"2024112119155177300_bib42","article-title":"Diverse beam search: Decoding diverse solutions from neural sequence models","author":"Vijayakumar","year":"2016","journal-title":"arXiv:1610.02424"},{"key":"2024112119155177300_bib43","article-title":"Chain of thought prompting elicits reasoning in large language models","volume-title":"Advances in Neural Information Processing Systems","author":"Wei","year":"2022"},{"key":"2024112119155177300_bib44","doi-asserted-by":"publisher","first-page":"5621","DOI":"10.18653\/v1\/2022.findings-emnlp.412","article-title":"Prompt compression and contrastive conditioning for controllability and toxicity reduction in language models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Wingate","year":"2022"},{"key":"2024112119155177300_bib45","article-title":"Efficient streaming language models with attention sinks","volume-title":"The Twelfth International Conference on Learning Representations","author":"Xiao","year":"2024"},{"key":"2024112119155177300_bib46","article-title":"H2O: Heavy-hitter oracle for efficient generative inference of large language models","volume-title":"Thirty-seventh Conference on Neural Information Processing Systems","author":"Zhang","year":"2023"},{"key":"2024112119155177300_bib47","first-page":"18330","article-title":"BERT loses patience: Fast and robust inference with early exit","volume-title":"Advances in Neural Information Processing Systems","author":"Zhou","year":"2020"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00716\/2480391\/tacl_a_00716.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00716\/2480391\/tacl_a_00716.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T19:16:06Z","timestamp":1732216566000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00716\/125280\/FINCH-Prompt-guided-Key-Value-Cache-Compression"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":47,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00716","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}