Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
Fuente:
arXiv
Guardado en:
| Autores principales: | Xi, Haocheng, Yang, Shuo, Zhao, Yilong, Li, Muyang, Cai, Han, Li, Xingyang, Lin, Yujun, Zhang, Zhuoyang, Zhang, Jintao, Li, Xiuyu, Xu, Zhiying, Wu, Jun, Xu, Chenfeng, Stoica, Ion, Han, Song, Keutzer, Kurt |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
por: Xi, Haocheng, et al.
Publicado: (2025)
por: Xi, Haocheng, et al.
Publicado: (2025)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
por: Yang, Shuo, et al.
Publicado: (2025)
por: Yang, Shuo, et al.
Publicado: (2025)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
por: Yang, Shuo, et al.
Publicado: (2026)
por: Yang, Shuo, et al.
Publicado: (2026)
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
por: Chen, Junyu, et al.
Publicado: (2025)
por: Chen, Junyu, et al.
Publicado: (2025)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
por: Li, Xingyang, et al.
Publicado: (2025)
por: Li, Xingyang, et al.
Publicado: (2025)
StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
por: Feng, Tianrui, et al.
Publicado: (2025)
por: Feng, Tianrui, et al.
Publicado: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
por: Tiwari, Rishabh, et al.
Publicado: (2025)
por: Tiwari, Rishabh, et al.
Publicado: (2025)
VideoGen-Eval: Agent-based System for Video Generation Evaluation
por: Yang, Yuhang, et al.
Publicado: (2025)
por: Yang, Yuhang, et al.
Publicado: (2025)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
por: Han, Insu, et al.
Publicado: (2025)
por: Han, Insu, et al.
Publicado: (2025)
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
por: Liang, Feng, et al.
Publicado: (2024)
por: Liang, Feng, et al.
Publicado: (2024)
SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing
por: Zhou, Xuanyi, et al.
Publicado: (2026)
por: Zhou, Xuanyi, et al.
Publicado: (2026)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
por: Chen, Han, et al.
Publicado: (2025)
por: Chen, Han, et al.
Publicado: (2025)
PolarQuant: Quantizing KV Caches with Polar Transformation
por: Han, Insu, et al.
Publicado: (2025)
por: Han, Insu, et al.
Publicado: (2025)
Segment Any Motion in Videos
por: Huang, Nan, et al.
Publicado: (2025)
por: Huang, Nan, et al.
Publicado: (2025)
Sparse Refinement for Efficient High-Resolution Semantic Segmentation
por: Liu, Zhijian, et al.
Publicado: (2024)
por: Liu, Zhijian, et al.
Publicado: (2024)
S*: Test Time Scaling for Code Generation
por: Li, Dacheng, et al.
Publicado: (2025)
por: Li, Dacheng, et al.
Publicado: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
por: Zheng, Mingzhe, et al.
Publicado: (2025)
por: Zheng, Mingzhe, et al.
Publicado: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
por: Zheng, Mingzhe, et al.
Publicado: (2024)
por: Zheng, Mingzhe, et al.
Publicado: (2024)
SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
por: Li, Muyang, et al.
Publicado: (2024)
por: Li, Muyang, et al.
Publicado: (2024)
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
por: Tomar, Aditya, et al.
Publicado: (2025)
por: Tomar, Aditya, et al.
Publicado: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
por: Jiang, Jindong, et al.
Publicado: (2025)
por: Jiang, Jindong, et al.
Publicado: (2025)
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
por: He, Xingyang, et al.
Publicado: (2025)
por: He, Xingyang, et al.
Publicado: (2025)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
por: Chen, Yilong, et al.
Publicado: (2025)
por: Chen, Yilong, et al.
Publicado: (2025)
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
por: Khaki, Samir, et al.
Publicado: (2025)
por: Khaki, Samir, et al.
Publicado: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
por: Di, Shangzhe, et al.
Publicado: (2025)
por: Di, Shangzhe, et al.
Publicado: (2025)
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
por: Xi, Haocheng, et al.
Publicado: (2026)
por: Xi, Haocheng, et al.
Publicado: (2026)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
por: Jia, Jinda, et al.
Publicado: (2026)
por: Jia, Jinda, et al.
Publicado: (2026)
QuantCache: Adaptive Importance-Guided Quantization with Hierarchical Latent and Layer Caching for Video Generation
por: Wu, Junyi, et al.
Publicado: (2025)
por: Wu, Junyi, et al.
Publicado: (2025)
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
por: Li, Weiqing, et al.
Publicado: (2025)
por: Li, Weiqing, et al.
Publicado: (2025)
Immiscible Diffusion: Accelerating Diffusion Training with Noise Assignment
por: Li, Yiheng, et al.
Publicado: (2024)
por: Li, Yiheng, et al.
Publicado: (2024)
Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility
por: Li, Yiheng, et al.
Publicado: (2025)
por: Li, Yiheng, et al.
Publicado: (2025)
Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
por: Tao, Keda, et al.
Publicado: (2025)
por: Tao, Keda, et al.
Publicado: (2025)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
por: Zandieh, Amir, et al.
Publicado: (2024)
por: Zandieh, Amir, et al.
Publicado: (2024)
Magic-Me: Identity-Specific Video Customized Diffusion
por: Ma, Ze, et al.
Publicado: (2024)
por: Ma, Ze, et al.
Publicado: (2024)
HallE-Control: Controlling Object Hallucination in Large Multimodal Models
por: Zhai, Bohan, et al.
Publicado: (2023)
por: Zhai, Bohan, et al.
Publicado: (2023)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
por: Yesiltepe, Hidir, et al.
Publicado: (2026)
por: Yesiltepe, Hidir, et al.
Publicado: (2026)
QuantVSR: Low-Bit Post-Training Quantization for Real-World Video Super-Resolution
por: Chai, Bowen, et al.
Publicado: (2025)
por: Chai, Bowen, et al.
Publicado: (2025)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
por: Pang, Zhanzhong, et al.
Publicado: (2026)
por: Pang, Zhanzhong, et al.
Publicado: (2026)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
por: Sharma, Akshat, et al.
Publicado: (2024)
por: Sharma, Akshat, et al.
Publicado: (2024)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
por: Ji, Yicheng, et al.
Publicado: (2026)
por: Ji, Yicheng, et al.
Publicado: (2026)
Ejemplares similares
-
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
por: Xi, Haocheng, et al.
Publicado: (2025) -
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
por: Yang, Shuo, et al.
Publicado: (2025) -
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
por: Yang, Shuo, et al.
Publicado: (2026) -
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
por: Chen, Junyu, et al.
Publicado: (2025) -
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
por: Li, Xingyang, et al.
Publicado: (2025)