StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xuanyi, Yu, Chunan, Ji, Deyi, Zhu, Qi, Sun, Lingyun, Li, Xuanfu, Ma, Jin, Chen, Tianrun, Zhu, Lanyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HD-VGGT: High-Resolution Visual Geometry Transformer
von: Chen, Tianrun, et al.
Veröffentlicht: (2026)
von: Chen, Tianrun, et al.
Veröffentlicht: (2026)
Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors
von: Zang, Ying, et al.
Veröffentlicht: (2026)
von: Zang, Ying, et al.
Veröffentlicht: (2026)
4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
von: Zang, Ying, et al.
Veröffentlicht: (2026)
von: Zang, Ying, et al.
Veröffentlicht: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
von: Su, Zunhai, et al.
Veröffentlicht: (2026)
SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning
von: Ji, Deyi, et al.
Veröffentlicht: (2025)
von: Ji, Deyi, et al.
Veröffentlicht: (2025)
Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
LLaFS: When Large Language Models Meet Few-Shot Segmentation
von: Zhu, Lanyun, et al.
Veröffentlicht: (2023)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2023)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
von: Shu, Zhijian, et al.
Veröffentlicht: (2025)
von: Shu, Zhijian, et al.
Veröffentlicht: (2025)
POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
von: Liu, Yuhan, et al.
Veröffentlicht: (2023)
von: Liu, Yuhan, et al.
Veröffentlicht: (2023)
From Air to Wear: Personalized 3D Digital Fashion with AR/VR Immersive 3D Sketching
von: Zang, Ying, et al.
Veröffentlicht: (2025)
von: Zang, Ying, et al.
Veröffentlicht: (2025)
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
von: Xu, Zhisong, et al.
Veröffentlicht: (2026)
von: Xu, Zhisong, et al.
Veröffentlicht: (2026)
IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding
von: Zhu, Lanyun, et al.
Veröffentlicht: (2024)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2024)
Img2CAD: Conditioned 3D CAD Model Generation from Single Image with Structured Visual Geometry
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)
VGGT: Visual Geometry Grounded Transformer
von: Wang, Jianyuan, et al.
Veröffentlicht: (2025)
von: Wang, Jianyuan, et al.
Veröffentlicht: (2025)
Video-Zero: Self-Evolution Video Understanding
von: Zhang, Ruixu, et al.
Veröffentlicht: (2026)
von: Zhang, Ruixu, et al.
Veröffentlicht: (2026)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
CamGeo: Sparse Camera-Conditioned Image-to-Video Generation with 3D Geometry Priors
von: Liu, Xuanyi, et al.
Veröffentlicht: (2026)
von: Liu, Xuanyi, et al.
Veröffentlicht: (2026)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
von: Wang, Zipeng, et al.
Veröffentlicht: (2025)
von: Wang, Zipeng, et al.
Veröffentlicht: (2025)
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
von: Shen, You, et al.
Veröffentlicht: (2025)
von: Shen, You, et al.
Veröffentlicht: (2025)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
von: Kim, Minsoo, et al.
Veröffentlicht: (2025)
von: Kim, Minsoo, et al.
Veröffentlicht: (2025)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
von: Sun, Xiangyu, et al.
Veröffentlicht: (2026)
von: Sun, Xiangyu, et al.
Veröffentlicht: (2026)
The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026)
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026)
xLSTM-UNet can be an Effective 2D & 3D Medical Image Segmentation Backbone with Vision-LSTM (ViL) better than its Mamba Counterpart
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
STAC: Plug-and-Play Spatio-Temporal Aware Cache Compression for Streaming 3D Reconstruction
von: Wang, Runze, et al.
Veröffentlicht: (2026)
von: Wang, Runze, et al.
Veröffentlicht: (2026)
Breaking the Box: Enhancing Remote Sensing Image Segmentation with Freehand Sketches
von: Zang, Ying, et al.
Veröffentlicht: (2025)
von: Zang, Ying, et al.
Veröffentlicht: (2025)
Syllables to Scenes: Literary-Guided Free-Viewpoint 3D Scene Synthesis from Japanese Haiku
von: Yu, Chunan, et al.
Veröffentlicht: (2025)
von: Yu, Chunan, et al.
Veröffentlicht: (2025)
Streaming 4D Visual Geometry Transformer
von: Zhuo, Dong, et al.
Veröffentlicht: (2025)
von: Zhuo, Dong, et al.
Veröffentlicht: (2025)
Training Transformers for KV Cache Compressibility
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HD-VGGT: High-Resolution Visual Geometry Transformer
von: Chen, Tianrun, et al.
Veröffentlicht: (2026) -
Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors
von: Zang, Ying, et al.
Veröffentlicht: (2026) -
4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
von: Zang, Ying, et al.
Veröffentlicht: (2026) -
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026) -
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
von: Su, Zunhai, et al.
Veröffentlicht: (2026)