Pyramid Cache: Layer-Adaptive KV Cache Compression with Signature-Based Cold Storage
Fuente:
Zenodo
Enregistré dans:
| Auteur principal: | |
|---|---|
| Format: | Recurso digital |
| Publié: |
Zenodo
2026
|
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866901841238818816 |
|---|---|
| author | Sergio dj |
| author_facet | Sergio dj |
| contents | <p class="p1">We present Pyramid Cache, a layer-adaptive compression architecture for transformer KV</p> <p class="p1">caches combined with a signature-based cold storage mechanism. Our approach exploits</p> <p class="p1">three properties: (1) token-identity structure varies systematically across layers, enabling</p> <p class="p1">layer-adaptive compression; (2) direction-only signatures predict attention relevance with</p> <p class="p1">0.926 Spearman correlation; and (3) only 0.4% of tokens receive meaningful attention at</p> <p class="p1">any step. A proof-of-concept implementation demonstrates that the system correctly</p> <p class="p1">answers both factual retrieval and complex reasoning questions — including multi-hop</p> <p class="p1">reasoning, causal chains, and contradiction detection — while skipping 99% of tokens</p> <p class="p1">(attending to only 5-8 out of 500+). Complex reasoning tasks show a similarity gap of only</p> <p class="p1">0.0215 compared to simple retrieval, with an EXCELLENT verdict at 1% reconstruction</p> <p class="p1">budget. Our approach is complementary to existing quantization methods like TurboQuant</p> <p class="p1">and represents a fundamentally different strategy: rather than compressing all vectors</p> <p class="p1">equally, we identify and skip the irrelevant ones entirely.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19536239 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Pyramid Cache: Layer-Adaptive KV Cache Compression with Signature-Based Cold Storage Sergio dj <p class="p1">We present Pyramid Cache, a layer-adaptive compression architecture for transformer KV</p> <p class="p1">caches combined with a signature-based cold storage mechanism. Our approach exploits</p> <p class="p1">three properties: (1) token-identity structure varies systematically across layers, enabling</p> <p class="p1">layer-adaptive compression; (2) direction-only signatures predict attention relevance with</p> <p class="p1">0.926 Spearman correlation; and (3) only 0.4% of tokens receive meaningful attention at</p> <p class="p1">any step. A proof-of-concept implementation demonstrates that the system correctly</p> <p class="p1">answers both factual retrieval and complex reasoning questions — including multi-hop</p> <p class="p1">reasoning, causal chains, and contradiction detection — while skipping 99% of tokens</p> <p class="p1">(attending to only 5-8 out of 500+). Complex reasoning tasks show a similarity gap of only</p> <p class="p1">0.0215 compared to simple retrieval, with an EXCELLENT verdict at 1% reconstruction</p> <p class="p1">budget. Our approach is complementary to existing quantization methods like TurboQuant</p> <p class="p1">and represents a fundamentally different strategy: rather than compressing all vectors</p> <p class="p1">equally, we identify and skip the irrelevant ones entirely.</p> |
| title | Pyramid Cache: Layer-Adaptive KV Cache Compression with Signature-Based Cold Storage |
| url | https://doi.org/10.5281/zenodo.19536239 |