SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Xuanyi, Mang, Qiuyang, Yang, Shuo, Xi, Haocheng, Zhang, Jintao, Mao, Huanzhi, Gonzalez, Joseph E., Keutzer, Kurt, Stoica, Ion, Cheung, Alvin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
von: Li, Hanchen, et al.
Veröffentlicht: (2025)
von: Li, Hanchen, et al.
Veröffentlicht: (2025)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Qrita: High-performance Top-k and Top-p using Pivot-based Truncation and Selection
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation
von: Zhou, Shang, et al.
Veröffentlicht: (2026)
von: Zhou, Shang, et al.
Veröffentlicht: (2026)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Automated Discovery of Test Oracles for Database Management Systems Using LLMs
von: Mang, Qiuyang, et al.
Veröffentlicht: (2025)
von: Mang, Qiuyang, et al.
Veröffentlicht: (2025)
EvoX: Meta-Evolution for Automated Discovery
von: Liu, Shu, et al.
Veröffentlicht: (2026)
von: Liu, Shu, et al.
Veröffentlicht: (2026)
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
Some Present-Day Problems of Romanian Library Science
von: Stoica, Ion
Veröffentlicht: (1973)
von: Stoica, Ion
Veröffentlicht: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
von: Stoica, Ion
Veröffentlicht: (1972)
von: Stoica, Ion
Veröffentlicht: (1972)
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
PLOP: Cost-Based Placement of Semantic Operators in Hybrid Query Plans
von: Mang, Qiuyang, et al.
Veröffentlicht: (2026)
von: Mang, Qiuyang, et al.
Veröffentlicht: (2026)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Combee: Scaling Prompt Learning for Self-Improving Language Model Agents
von: Li, Hanchen, et al.
Veröffentlicht: (2026)
von: Li, Hanchen, et al.
Veröffentlicht: (2026)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
EAR
Veröffentlicht: (2025)
Veröffentlicht: (2025)
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan
von: Nehrdich, Sebastian, et al.
Veröffentlicht: (2026)
von: Nehrdich, Sebastian, et al.
Veröffentlicht: (2026)
S*: Test Time Scaling for Code Generation
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale
von: He, Runyuan, et al.
Veröffentlicht: (2026)
von: He, Runyuan, et al.
Veröffentlicht: (2026)
Revisiting Cache Freshness for Emerging Real-Time Applications
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching
von: Zhao, Yilong, et al.
Veröffentlicht: (2024)
von: Zhao, Yilong, et al.
Veröffentlicht: (2024)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
Shangri-La, paraíso terrenal / Liu Huanzhi y Zhang Tao
von: Liu Huanzhi
Veröffentlicht: (2008)
von: Liu Huanzhi
Veröffentlicht: (2008)
Campesinos buscan una vida modestamente acomodada / Liu Huanzhi
von: Liu Huanzhi
von: Liu Huanzhi
HTM-EAR: Importance-Preserving Tiered Memory with Hybrid Routing under Saturation
von: Singh, Shubham Kumar
Veröffentlicht: (2026)
von: Singh, Shubham Kumar
Veröffentlicht: (2026)
vAttention: Verified Sparse Attention
von: Desai, Aditya, et al.
Veröffentlicht: (2025)
von: Desai, Aditya, et al.
Veröffentlicht: (2025)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
von: Li, Hanchen, et al.
Veröffentlicht: (2025) -
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
von: Yang, Shuo, et al.
Veröffentlicht: (2025) -
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
von: Xi, Haocheng, et al.
Veröffentlicht: (2025) -
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
von: Zhang, Jintao, et al.
Veröffentlicht: (2026) -
Qrita: High-performance Top-k and Top-p using Pivot-based Truncation and Selection
von: Park, Jongseok, et al.
Veröffentlicht: (2026)