FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Akide, Zhang, Zeyu, Li, Zhexin, Bai, Xuehai, Han, Yizeng, Tang, Jiasheng, Xing, Yuanjie, Wu, Jichao, Yang, Mingyang, Chen, Weihua, He, Jiahao, He, Yuanyu, Wang, Fan, Haffari, Gholamreza, Zhuang, Bohan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
CoV: Chain-of-View Prompting for Spatial Reasoning
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
von: Inferix Team, et al.
Veröffentlicht: (2025)
von: Inferix Team, et al.
Veröffentlicht: (2025)
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
von: Liu, Akide, et al.
Veröffentlicht: (2026)
von: Liu, Akide, et al.
Veröffentlicht: (2026)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
von: Chen, Feng, et al.
Veröffentlicht: (2025)
von: Chen, Feng, et al.
Veröffentlicht: (2025)
Motion Mamba: Efficient and Long Sequence Motion Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
Few-Step Distillation for Text-to-Image Generation: A Practical Guide
von: Pu, Yifan, et al.
Veröffentlicht: (2025)
von: Pu, Yifan, et al.
Veröffentlicht: (2025)
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
von: Wang, Weijie, et al.
Veröffentlicht: (2025)
von: Wang, Weijie, et al.
Veröffentlicht: (2025)
Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
von: Chen, Feng, et al.
Veröffentlicht: (2025)
von: Chen, Feng, et al.
Veröffentlicht: (2025)
(Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
Neighboring Autoregressive Modeling for Efficient Visual Generation
von: He, Yefei, et al.
Veröffentlicht: (2025)
von: He, Yefei, et al.
Veröffentlicht: (2025)
ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality
von: He, Yefei, et al.
Veröffentlicht: (2024)
von: He, Yefei, et al.
Veröffentlicht: (2024)
Evidence-based Distributional Alignment for Large Language Models
von: Pham, Viet-Thanh, et al.
Veröffentlicht: (2026)
von: Pham, Viet-Thanh, et al.
Veröffentlicht: (2026)
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues
von: Hua, Yuncheng, et al.
Veröffentlicht: (2024)
von: Hua, Yuncheng, et al.
Veröffentlicht: (2024)
Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach
von: Li, Zhuang, et al.
Veröffentlicht: (2024)
von: Li, Zhuang, et al.
Veröffentlicht: (2024)
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
von: Bai, Xuehai, et al.
Veröffentlicht: (2026)
von: Bai, Xuehai, et al.
Veröffentlicht: (2026)
Bridging the Gap Between Multimodal Foundation Models and World Models
von: He, Xuehai
Veröffentlicht: (2025)
von: He, Xuehai
Veröffentlicht: (2025)
Discrete Minds in a Continuous World: Do Language Models Know Time Passes?
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler
von: Wang, Minghan, et al.
Veröffentlicht: (2026)
von: Wang, Minghan, et al.
Veröffentlicht: (2026)
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
On the Reliability of Large Language Models for Causal Discovery
von: Feng, Tao, et al.
Veröffentlicht: (2024)
von: Feng, Tao, et al.
Veröffentlicht: (2024)
IMO: Greedy Layer-Wise Sparse Representation Learning for Out-of-Distribution Text Classification with Pre-trained Models
von: Feng, Tao, et al.
Veröffentlicht: (2024)
von: Feng, Tao, et al.
Veröffentlicht: (2024)
An Empirical Study on How Video-LLMs Answer Video Questions
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
Less Detail, Better Answers: Degradation-Driven Prompting for VQA
von: Han, Haoxuan, et al.
Veröffentlicht: (2026)
von: Han, Haoxuan, et al.
Veröffentlicht: (2026)
Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Zero-Shot Privacy-Aware Text Rewriting via Iterative Tree Search
von: Huang, Shuo, et al.
Veröffentlicht: (2025)
von: Huang, Shuo, et al.
Veröffentlicht: (2025)
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
von: Yang, Hao, et al.
Veröffentlicht: (2025)
von: Yang, Hao, et al.
Veröffentlicht: (2025)
IRIS: An Iterative and Integrated Framework for Verifiable Causal Discovery in the Absence of Tabular Data
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue Systems
von: Feng, Tao, et al.
Veröffentlicht: (2024)
von: Feng, Tao, et al.
Veröffentlicht: (2024)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models
von: Li, Zhuang, et al.
Veröffentlicht: (2024)
von: Li, Zhuang, et al.
Veröffentlicht: (2024)
RIDE: Enhancing Large Language Model Alignment through Restyled In-Context Learning Demonstration Exemplars
von: Hua, Yuncheng, et al.
Veröffentlicht: (2025)
von: Hua, Yuncheng, et al.
Veröffentlicht: (2025)
ME-Switch: A Memory-Efficient Expert Switching Framework for Large Language Models
von: Liu, Jing, et al.
Veröffentlicht: (2024)
von: Liu, Jing, et al.
Veröffentlicht: (2024)
FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation
von: Zhou, Junkang, et al.
Veröffentlicht: (2026)
von: Zhou, Junkang, et al.
Veröffentlicht: (2026)
AIPO: Learning to Reason from Active Interaction
von: Liu, Junnan, et al.
Veröffentlicht: (2026)
von: Liu, Junnan, et al.
Veröffentlicht: (2026)
The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite Graph
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024) -
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025) -
CoV: Chain-of-View Prompting for Spatial Reasoning
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026) -
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
von: Inferix Team, et al.
Veröffentlicht: (2025) -
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
von: Liu, Akide, et al.
Veröffentlicht: (2026)