iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD
arXiv:2610.02815v1 Announce Type: new Abstract: Long chain-of-thought reasoning substantially increases KV-cache memory during autoregressive decoding, as every generated token introduces new key and value states and causes the cache to grow linearly with decoding length. Existing KV-cache compression methods typically control this growth through token eviction, but irreversible deletion can remove historical states that later reasoning may need to revisit. SVD-based low-rank compression provid…