AnchorKV: Anchor-Residual KV Cache Compression
9/10AnchorKV presents an innovative key-value cache compression method for large language model inference with long contexts, significantly reducing memory requirements and mitigating token eviction losses. The approach improved inference speed and memory efficiency across different large models, providing a practical technique for deploying LLMs with extended context windows in production.
