XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Su, Zunhai, Ye, Weihao, Feng, Hansen, Fan, Keyu, Zhang, Jing, Yu, Dahai, Liu, Zhengwu, Wong, Ngai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
von: Liu, Xuanyi, et al.
Veröffentlicht: (2026)
von: Liu, Xuanyi, et al.
Veröffentlicht: (2026)
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
CodeComp: Structural KV Cache Compression for Agentic Coding
von: Chen, Qiujiang, et al.
Veröffentlicht: (2026)
von: Chen, Qiujiang, et al.
Veröffentlicht: (2026)
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
von: Xu, Zhisong, et al.
Veröffentlicht: (2026)
von: Xu, Zhisong, et al.
Veröffentlicht: (2026)
VGGT: Visual Geometry Grounded Transformer
von: Wang, Jianyuan, et al.
Veröffentlicht: (2025)
von: Wang, Jianyuan, et al.
Veröffentlicht: (2025)
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
von: Kim, Minsoo, et al.
Veröffentlicht: (2025)
von: Kim, Minsoo, et al.
Veröffentlicht: (2025)
Training Transformers for KV Cache Compressibility
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
von: Shu, Zhijian, et al.
Veröffentlicht: (2025)
von: Shu, Zhijian, et al.
Veröffentlicht: (2025)
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
Decomposing Densification in Gaussian Splatting for Faster 3D Scene Reconstruction
von: Huang, Binxiao, et al.
Veröffentlicht: (2025)
von: Huang, Binxiao, et al.
Veröffentlicht: (2025)
DoPE: Denoising Rotary Position Embedding
von: Xiong, Jing, et al.
Veröffentlicht: (2025)
von: Xiong, Jing, et al.
Veröffentlicht: (2025)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
von: Sun, Xiangyu, et al.
Veröffentlicht: (2026)
von: Sun, Xiangyu, et al.
Veröffentlicht: (2026)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
von: Liu, Yuhan, et al.
Veröffentlicht: (2023)
von: Liu, Yuhan, et al.
Veröffentlicht: (2023)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
von: Wang, Zipeng, et al.
Veröffentlicht: (2025)
von: Wang, Zipeng, et al.
Veröffentlicht: (2025)
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
von: Wu, Xianfeng, et al.
Veröffentlicht: (2025)
von: Wu, Xianfeng, et al.
Veröffentlicht: (2025)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
von: Zhang, Haowei, et al.
Veröffentlicht: (2026)
von: Zhang, Haowei, et al.
Veröffentlicht: (2026)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
Unveiling Super Experts in Mixture-of-Experts Large Language Models
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
von: Xiao, He, et al.
Veröffentlicht: (2025)
von: Xiao, He, et al.
Veröffentlicht: (2025)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory
von: Deng, Tianchen, et al.
Veröffentlicht: (2026)
von: Deng, Tianchen, et al.
Veröffentlicht: (2026)
CaliDrop: KV Cache Compression with Calibration
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
CktFormalizer: Autoformalization of Natural Language into Circuit Representations
von: Xiong, Jing, et al.
Veröffentlicht: (2026)
von: Xiong, Jing, et al.
Veröffentlicht: (2026)
GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression
von: Goldstein, Daniel, et al.
Veröffentlicht: (2024)
von: Goldstein, Daniel, et al.
Veröffentlicht: (2024)
SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes
von: Lee, Jungho, et al.
Veröffentlicht: (2025)
von: Lee, Jungho, et al.
Veröffentlicht: (2025)
StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference
von: Huang, Kai, et al.
Veröffentlicht: (2025)
von: Huang, Kai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026) -
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
von: Liu, Xuanyi, et al.
Veröffentlicht: (2026) -
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
von: Su, Zunhai, et al.
Veröffentlicht: (2025) -
CodeComp: Structural KV Cache Compression for Agentic Coding
von: Chen, Qiujiang, et al.
Veröffentlicht: (2026) -
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)