FastCar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shen, Xuan, Ma, Weize, Zhou, Yufa, Tang, Enhao, Xie, Yanyue, Li, Zhengang, Gong, Yifan, Wang, Quanyi, Ding, Henghui, Wang, Yiwei, Wang, Yanzhi, Zhao, Pu, Lin, Jun, Gu, Jiuxiang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance
par: Shen, Xuan, et autres
Publié: (2025)
par: Shen, Xuan, et autres
Publié: (2025)
QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge
par: Shen, Xuan, et autres
Publié: (2025)
par: Shen, Xuan, et autres
Publié: (2025)
Efficient Reasoning with Hidden Thinking
par: Shen, Xuan, et autres
Publié: (2025)
par: Shen, Xuan, et autres
Publié: (2025)
LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
par: Shen, Xuan, et autres
Publié: (2024)
par: Shen, Xuan, et autres
Publié: (2024)
Squat: Quant Small Language Models on the Edge
par: Shen, Xuan, et autres
Publié: (2024)
par: Shen, Xuan, et autres
Publié: (2024)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
par: Li, Bohan, et autres
Publié: (2024)
par: Li, Bohan, et autres
Publié: (2024)
Collaborative Compression for Large-Scale MoE Deployment on Edge
par: Chen, Yixiao, et autres
Publié: (2025)
par: Chen, Yixiao, et autres
Publié: (2025)
Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
par: Gu, Yuxuan, et autres
Publié: (2025)
par: Gu, Yuxuan, et autres
Publié: (2025)
Fast and Memory-Efficient Video Diffusion Using Streamlined Inference
par: Zhan, Zheng, et autres
Publié: (2024)
par: Zhan, Zheng, et autres
Publié: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
par: Shen, Xuan, et autres
Publié: (2023)
par: Shen, Xuan, et autres
Publié: (2023)
FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation
par: Liu, Dong, et autres
Publié: (2025)
par: Liu, Dong, et autres
Publié: (2025)
Fast Solve of Broadband Electromagnetic Scattering Problems Based on Krylov Subspace Basis Functions Combining With Compressive Sensing
par: Zhonggen Wang, et autres
Publié: (2025)
par: Zhonggen Wang, et autres
Publié: (2025)
OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation
par: Zhao, Lin, et autres
Publié: (2026)
par: Zhao, Lin, et autres
Publié: (2026)
fastkqr: A Fast Algorithm for Kernel Quantile Regression
par: Tang, Qian, et autres
Publié: (2024)
par: Tang, Qian, et autres
Publié: (2024)
HybridFlow: Infusing Continuity into Masked Codebook for Extreme Low-Bitrate Image Compression
par: Lu, Lei, et autres
Publié: (2024)
par: Lu, Lei, et autres
Publié: (2024)
OBIFormer: A Fast Attentive Denoising Framework for Oracle Bone Inscriptions
par: Li, Jinhao, et autres
Publié: (2025)
par: Li, Jinhao, et autres
Publié: (2025)
MagCache: Fast Video Generation with Magnitude-Aware Cache
par: Ma, Zehong, et autres
Publié: (2025)
par: Ma, Zehong, et autres
Publié: (2025)
Semi-supervised Method for Risk Prediction with Doubly Censored EHR Data
par: Zhou, Jie, et autres
Publié: (2026)
par: Zhou, Jie, et autres
Publié: (2026)
Numerical Pruning for Efficient Autoregressive Models
par: Shen, Xuan, et autres
Publié: (2024)
par: Shen, Xuan, et autres
Publié: (2024)
AyE-Edge: Automated Deployment Space Search Empowering Accuracy yet Efficient Real-Time Object Detection on the Edge
par: Wu, Chao, et autres
Publié: (2024)
par: Wu, Chao, et autres
Publié: (2024)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
par: Li, Bingxuan, et autres
Publié: (2025)
par: Li, Bingxuan, et autres
Publié: (2025)
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
par: Xie, Yanyue, et autres
Publié: (2024)
par: Xie, Yanyue, et autres
Publié: (2024)
Fast Multichannel Topology Discovery in Cognitive Radio Networks
par: Wang, Yung-Li, et autres
Publié: (2025)
par: Wang, Yung-Li, et autres
Publié: (2025)
Rethinking Token Reduction for State Space Models
par: Zhan, Zheng, et autres
Publié: (2024)
par: Zhan, Zheng, et autres
Publié: (2024)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
par: Li, Zhengang, et autres
Publié: (2024)
par: Li, Zhengang, et autres
Publié: (2024)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
par: Zhu, Jianian, et autres
Publié: (2025)
par: Zhu, Jianian, et autres
Publié: (2025)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
par: Wang, Xuan, et autres
Publié: (2024)
par: Wang, Xuan, et autres
Publié: (2024)
EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
par: Wang, Chen, et autres
Publié: (2025)
par: Wang, Chen, et autres
Publié: (2025)
Calibrating Car-Following Models via Bayesian Dynamic Regression
par: Zhang, Chengyuan, et autres
Publié: (2023)
par: Zhang, Chengyuan, et autres
Publié: (2023)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
par: li, Fei, et autres
Publié: (2026)
par: li, Fei, et autres
Publié: (2026)
Differential Privacy Mechanisms in Neural Tangent Kernel Regression
par: Gu, Jiuxiang, et autres
Publié: (2024)
par: Gu, Jiuxiang, et autres
Publié: (2024)
Fast Authenticated and Interoperable Multimedia Healthcare Data over Hybrid-Storage Blockchains
par: Yang, Jucai, et autres
Publié: (2025)
par: Yang, Jucai, et autres
Publié: (2025)
CacheFlow: Fast Human Motion Prediction by Cached Normalizing Flow
par: Maeda, Takahiro, et autres
Publié: (2025)
par: Maeda, Takahiro, et autres
Publié: (2025)
Updatable Balanced Index for Fast On-device Search with Auto-selection Model
par: Ji, Yushuai, et autres
Publié: (2025)
par: Ji, Yushuai, et autres
Publié: (2025)
In-situ Autoguidance: Eliciting Self-Correction in Diffusion Models
par: Gu, Enhao, et autres
Publié: (2025)
par: Gu, Enhao, et autres
Publié: (2025)
SuperFlow: A Fully-Customized RTL-to-GDS Design Automation Flow for Adiabatic Quantum-Flux-Parametron Superconducting Circuits
par: Xie, Yanyue, et autres
Publié: (2024)
par: Xie, Yanyue, et autres
Publié: (2024)
Search for Efficient Large Language Models
par: Shen, Xuan, et autres
Publié: (2024)
par: Shen, Xuan, et autres
Publié: (2024)
Lotus: learning-based online thermal and latency variation management for two-stage detectors on edge devices
par: Gong, Yifan, et autres
Publié: (2024)
par: Gong, Yifan, et autres
Publié: (2024)
SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model
par: Li, Zhengang, et autres
Publié: (2024)
par: Li, Zhengang, et autres
Publié: (2024)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
par: Zhang, Dingyan, et autres
Publié: (2024)
par: Zhang, Dingyan, et autres
Publié: (2024)
Documents similaires
-
DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance
par: Shen, Xuan, et autres
Publié: (2025) -
QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge
par: Shen, Xuan, et autres
Publié: (2025) -
Efficient Reasoning with Hidden Thinking
par: Shen, Xuan, et autres
Publié: (2025) -
LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
par: Shen, Xuan, et autres
Publié: (2024) -
Squat: Quant Small Language Models on the Edge
par: Shen, Xuan, et autres
Publié: (2024)