FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Akide, Zhang, Zeyu, Li, Zhexin, Bai, Xuehai, Han, Yizeng, Tang, Jiasheng, Xing, Yuanjie, Wu, Jichao, Yang, Mingyang, Chen, Weihua, He, Jiahao, He, Yuanyu, Wang, Fan, Haffari, Gholamreza, Zhuang, Bohan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
por: Liu, Akide, et al.
Publicado: (2024)
por: Liu, Akide, et al.
Publicado: (2024)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
por: Zhang, Zeyu, et al.
Publicado: (2025)
por: Zhang, Zeyu, et al.
Publicado: (2025)
CoV: Chain-of-View Prompting for Spatial Reasoning
por: Zhao, Haoyu, et al.
Publicado: (2026)
por: Zhao, Haoyu, et al.
Publicado: (2026)
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
por: Inferix Team, et al.
Publicado: (2025)
por: Inferix Team, et al.
Publicado: (2025)
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
por: Liu, Akide, et al.
Publicado: (2026)
por: Liu, Akide, et al.
Publicado: (2026)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
por: Chen, Feng, et al.
Publicado: (2025)
por: Chen, Feng, et al.
Publicado: (2025)
Motion Mamba: Efficient and Long Sequence Motion Generation
por: Zhang, Zeyu, et al.
Publicado: (2024)
por: Zhang, Zeyu, et al.
Publicado: (2024)
Few-Step Distillation for Text-to-Image Generation: A Practical Guide
por: Pu, Yifan, et al.
Publicado: (2025)
por: Pu, Yifan, et al.
Publicado: (2025)
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
por: Wang, Weijie, et al.
Publicado: (2025)
por: Wang, Weijie, et al.
Publicado: (2025)
Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
por: Chen, Feng, et al.
Publicado: (2025)
por: Chen, Feng, et al.
Publicado: (2025)
(Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
por: Wu, Minghao, et al.
Publicado: (2024)
por: Wu, Minghao, et al.
Publicado: (2024)
Neighboring Autoregressive Modeling for Efficient Visual Generation
por: He, Yefei, et al.
Publicado: (2025)
por: He, Yefei, et al.
Publicado: (2025)
ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality
por: He, Yefei, et al.
Publicado: (2024)
por: He, Yefei, et al.
Publicado: (2024)
Evidence-based Distributional Alignment for Large Language Models
por: Pham, Viet-Thanh, et al.
Publicado: (2026)
por: Pham, Viet-Thanh, et al.
Publicado: (2026)
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues
por: Hua, Yuncheng, et al.
Publicado: (2024)
por: Hua, Yuncheng, et al.
Publicado: (2024)
Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach
por: Li, Zhuang, et al.
Publicado: (2024)
por: Li, Zhuang, et al.
Publicado: (2024)
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
por: Zhang, Zeyu, et al.
Publicado: (2024)
por: Zhang, Zeyu, et al.
Publicado: (2024)
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
por: Bai, Xuehai, et al.
Publicado: (2026)
por: Bai, Xuehai, et al.
Publicado: (2026)
Bridging the Gap Between Multimodal Foundation Models and World Models
por: He, Xuehai
Publicado: (2025)
por: He, Xuehai
Publicado: (2025)
Discrete Minds in a Continuous World: Do Language Models Know Time Passes?
por: Wang, Minghan, et al.
Publicado: (2025)
por: Wang, Minghan, et al.
Publicado: (2025)
GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler
por: Wang, Minghan, et al.
Publicado: (2026)
por: Wang, Minghan, et al.
Publicado: (2026)
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
por: Zhao, Wangbo, et al.
Publicado: (2025)
por: Zhao, Wangbo, et al.
Publicado: (2025)
On the Reliability of Large Language Models for Causal Discovery
por: Feng, Tao, et al.
Publicado: (2024)
por: Feng, Tao, et al.
Publicado: (2024)
IMO: Greedy Layer-Wise Sparse Representation Learning for Out-of-Distribution Text Classification with Pre-trained Models
por: Feng, Tao, et al.
Publicado: (2024)
por: Feng, Tao, et al.
Publicado: (2024)
An Empirical Study on How Video-LLMs Answer Video Questions
por: Gou, Chenhui, et al.
Publicado: (2025)
por: Gou, Chenhui, et al.
Publicado: (2025)
Less Detail, Better Answers: Degradation-Driven Prompting for VQA
por: Han, Haoxuan, et al.
Publicado: (2026)
por: Han, Haoxuan, et al.
Publicado: (2026)
Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
por: Yang, Hao, et al.
Publicado: (2024)
por: Yang, Hao, et al.
Publicado: (2024)
Zero-Shot Privacy-Aware Text Rewriting via Iterative Tree Search
por: Huang, Shuo, et al.
Publicado: (2025)
por: Huang, Shuo, et al.
Publicado: (2025)
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
por: Yang, Hao, et al.
Publicado: (2024)
por: Yang, Hao, et al.
Publicado: (2024)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
por: Yang, Hao, et al.
Publicado: (2024)
por: Yang, Hao, et al.
Publicado: (2024)
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
por: Yang, Hao, et al.
Publicado: (2025)
por: Yang, Hao, et al.
Publicado: (2025)
IRIS: An Iterative and Integrated Framework for Verifiable Causal Discovery in the Absence of Tabular Data
por: Feng, Tao, et al.
Publicado: (2025)
por: Feng, Tao, et al.
Publicado: (2025)
CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue Systems
por: Feng, Tao, et al.
Publicado: (2024)
por: Feng, Tao, et al.
Publicado: (2024)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
por: Xia, Haojun, et al.
Publicado: (2024)
por: Xia, Haojun, et al.
Publicado: (2024)
SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models
por: Li, Zhuang, et al.
Publicado: (2024)
por: Li, Zhuang, et al.
Publicado: (2024)
RIDE: Enhancing Large Language Model Alignment through Restyled In-Context Learning Demonstration Exemplars
por: Hua, Yuncheng, et al.
Publicado: (2025)
por: Hua, Yuncheng, et al.
Publicado: (2025)
ME-Switch: A Memory-Efficient Expert Switching Framework for Large Language Models
por: Liu, Jing, et al.
Publicado: (2024)
por: Liu, Jing, et al.
Publicado: (2024)
FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation
por: Zhou, Junkang, et al.
Publicado: (2026)
por: Zhou, Junkang, et al.
Publicado: (2026)
AIPO: Learning to Reason from Active Interaction
por: Liu, Junnan, et al.
Publicado: (2026)
por: Liu, Junnan, et al.
Publicado: (2026)
The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite Graph
por: Wu, Minghao, et al.
Publicado: (2024)
por: Wu, Minghao, et al.
Publicado: (2024)
Ejemplares similares
-
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
por: Liu, Akide, et al.
Publicado: (2024) -
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
por: Zhang, Zeyu, et al.
Publicado: (2025) -
CoV: Chain-of-View Prompting for Spatial Reasoning
por: Zhao, Haoyu, et al.
Publicado: (2026) -
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
por: Inferix Team, et al.
Publicado: (2025) -
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
por: Liu, Akide, et al.
Publicado: (2026)