X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Yixiao, Zheng, Jianlei, Zheng, Chaoda, Chen, Shijia, Liu, Mingdian, Liu, Tongping, Luo, Tengwei, Zhang, Yu, Wang, Boyang, Xu, Linkun, Lu, Siyuan, Tian, Bo, Liu, Xianming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
by: Wu, Shunlong, et al.
Published: (2026)
by: Wu, Shunlong, et al.
Published: (2026)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
Motion-Aware Caching for Efficient Autoregressive Video Generation
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
Beyond Few-Step Inference: Accelerating Video Diffusion Transformer Model Serving with Inter-Request Caching Reuse
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
by: Zheng, Chaoda, et al.
Published: (2026)
by: Zheng, Chaoda, et al.
Published: (2026)
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
by: Agarwal, Shubham, et al.
Published: (2025)
by: Agarwal, Shubham, et al.
Published: (2025)
FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
by: Lin, Hongbin, et al.
Published: (2025)
by: Lin, Hongbin, et al.
Published: (2025)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
FlexCache: Flexible Approximate Cache System for Video Diffusion
by: Sun, Desen, et al.
Published: (2024)
by: Sun, Desen, et al.
Published: (2024)
TrimCaching: Parameter-sharing AI Model Caching in Wireless Edge Networks
by: Qu, Guanqiao, et al.
Published: (2024)
by: Qu, Guanqiao, et al.
Published: (2024)
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
by: Qin, Ziran, et al.
Published: (2025)
by: Qin, Ziran, et al.
Published: (2025)
Graph-Guided Adaptive Channel Elimination for KV Cache Compression
by: Tong, Enwei, et al.
Published: (2026)
by: Tong, Enwei, et al.
Published: (2026)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
TrimCaching: Parameter-sharing Edge Caching for AI Model Downloading
by: Qu, Guanqiao, et al.
Published: (2024)
by: Qu, Guanqiao, et al.
Published: (2024)
Enhancing Cross‐Domain Book Classification Through Caching‐Enabled Networks and Transformer Technology
by: Qiang Li, et al.
Published: (2024)
by: Qiang Li, et al.
Published: (2024)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
by: Liu, Kaiwei, et al.
Published: (2025)
by: Liu, Kaiwei, et al.
Published: (2025)
Aligning with Human Values to Enhance Interaction: An eHMI-Mediated Lane-Changing Negotiation Strategy Using Bayesian Inference
by: Peng, Boyao, et al.
Published: (2025)
by: Peng, Boyao, et al.
Published: (2025)
CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
by: Wu, Guanlong, et al.
Published: (2026)
by: Wu, Guanlong, et al.
Published: (2026)
InstCache: A Predictive Cache for LLM Serving
by: Zou, Longwei, et al.
Published: (2024)
by: Zou, Longwei, et al.
Published: (2024)
Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
by: Lin, Hongzhan, et al.
Published: (2025)
by: Lin, Hongzhan, et al.
Published: (2025)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
by: Tian, Yuyang, et al.
Published: (2025)
by: Tian, Yuyang, et al.
Published: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
by: Sun, Hanshi, et al.
Published: (2024)
by: Sun, Hanshi, et al.
Published: (2024)
The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference
by: Chodavarapu, Ranjith, et al.
Published: (2026)
by: Chodavarapu, Ranjith, et al.
Published: (2026)
WorldCache: Content-Aware Caching for Accelerated Video World Models
by: Nawaz, Umair, et al.
Published: (2026)
by: Nawaz, Umair, et al.
Published: (2026)
Cache Me if You Can: Accelerating Diffusion Models through Block Caching
by: Wimbauer, Felix, et al.
Published: (2023)
by: Wimbauer, Felix, et al.
Published: (2023)
Demand Private Coded Caching: Small Cache Size
by: Lu, Qinyi, et al.
Published: (2025)
by: Lu, Qinyi, et al.
Published: (2025)
SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
KV Cache Compression for Inference Efficiency in LLMs: A Review
by: Liu, Yanyu, et al.
Published: (2025)
by: Liu, Yanyu, et al.
Published: (2025)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
by: Geng, Yingsheng, et al.
Published: (2026)
by: Geng, Yingsheng, et al.
Published: (2026)
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
by: Liu, Zuyan, et al.
Published: (2024)
by: Liu, Zuyan, et al.
Published: (2024)
Compiler-First State Space Duality and Portable $O(1)$ Autoregressive Caching for Inference
by: Santoni, Cosmo
Published: (2026)
by: Santoni, Cosmo
Published: (2026)
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
by: Feng, Weilun, et al.
Published: (2026)
by: Feng, Weilun, et al.
Published: (2026)
FedCache 2.0: Federated Edge Learning with Knowledge Caching and Dataset Distillation
by: Pan, Quyang, et al.
Published: (2024)
by: Pan, Quyang, et al.
Published: (2024)
ReFrame: Layer Caching for Accelerated Inference in Real-Time Rendering
by: Liu, Lufei, et al.
Published: (2025)
by: Liu, Lufei, et al.
Published: (2025)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
by: Liu, Akide, et al.
Published: (2024)
by: Liu, Akide, et al.
Published: (2024)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
by: Zhang, Huawei, et al.
Published: (2025)
by: Zhang, Huawei, et al.
Published: (2025)
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
by: Samuel, Dvir, et al.
Published: (2026)
by: Samuel, Dvir, et al.
Published: (2026)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
by: Nadali, Alireza, et al.
Published: (2026)
by: Nadali, Alireza, et al.
Published: (2026)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
by: Chen, Jinhan, et al.
Published: (2025)
by: Chen, Jinhan, et al.
Published: (2025)
vCache: Verified Semantic Prompt Caching
by: Schroeder, Luis Gaspar, et al.
Published: (2025)
by: Schroeder, Luis Gaspar, et al.
Published: (2025)
Similar Items
-
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
by: Wu, Shunlong, et al.
Published: (2026) -
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025) -
Motion-Aware Caching for Efficient Autoregressive Video Generation
by: Xu, Jing, et al.
Published: (2026) -
Beyond Few-Step Inference: Accelerating Video Diffusion Transformer Model Serving with Inter-Request Caching Reuse
by: Liu, Hao, et al.
Published: (2026) -
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
by: Zheng, Chaoda, et al.
Published: (2026)