APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Yuxiang, Li, Mingye, Han, Xu, Xiao, Chaojun, Zhao, Weilin, Sun, Ao, Yuan, Ziqi, Zhou, Hao, Meng, Fandong, Liu, Zhiyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
von: Sun, Ao, et al.
Veröffentlicht: (2024)
von: Sun, Ao, et al.
Veröffentlicht: (2024)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
NOSA: Native and Offloadable Sparse Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model Training
von: Sun, Ao, et al.
Veröffentlicht: (2024)
von: Sun, Ao, et al.
Veröffentlicht: (2024)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
von: Song, Chenyang, et al.
Veröffentlicht: (2025)
von: Song, Chenyang, et al.
Veröffentlicht: (2025)
BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
von: Sun, Ao, et al.
Veröffentlicht: (2025)
von: Sun, Ao, et al.
Veröffentlicht: (2025)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
von: Zhao, Weilin, et al.
Veröffentlicht: (2024)
von: Zhao, Weilin, et al.
Veröffentlicht: (2024)
MR. Video: "MapReduce" is the Principle for Long Video Understanding
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
Accelerating Inference in Large Language Models with a Unified Layer Skipping Strategy
von: Liu, Yijin, et al.
Veröffentlicht: (2024)
von: Liu, Yijin, et al.
Veröffentlicht: (2024)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
Linear Attention Sequence Parallelism
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
End-to-End Dense Video Grounding via Parallel Regression
von: Shi, Fengyuan, et al.
Veröffentlicht: (2021)
von: Shi, Fengyuan, et al.
Veröffentlicht: (2021)
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
TASP: Topology-aware Sequence Parallelism
von: Wang, Yida, et al.
Veröffentlicht: (2025)
von: Wang, Yida, et al.
Veröffentlicht: (2025)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
Understanding and Addressing the Under-Translation Problem from the Perspective of Decoding Objective
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
von: Mu, Yongyu, et al.
Veröffentlicht: (2026)
von: Mu, Yongyu, et al.
Veröffentlicht: (2026)
DRT: Deep Reasoning Translation via Long Chain-of-Thought
von: Wang, Jiaan, et al.
Veröffentlicht: (2024)
von: Wang, Jiaan, et al.
Veröffentlicht: (2024)
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
von: Zhang, Zheyu, et al.
Veröffentlicht: (2026)
von: Zhang, Zheyu, et al.
Veröffentlicht: (2026)
db-SP: Accelerating Sparse Attention for Visual Generative Models with Dual-Balanced Sequence Parallelism
von: Chen, Siqi, et al.
Veröffentlicht: (2025)
von: Chen, Siqi, et al.
Veröffentlicht: (2025)
Densing Law of LLMs
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
von: Chen, Yukang, et al.
Veröffentlicht: (2026)
von: Chen, Yukang, et al.
Veröffentlicht: (2026)
MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
Long-VMNet: Accelerating Long-Form Video Understanding via Fixed Memory
von: Gurukar, Saket, et al.
Veröffentlicht: (2025)
von: Gurukar, Saket, et al.
Veröffentlicht: (2025)
Variator: Accelerating Pre-trained Models with Plug-and-Play Compression Modules
von: Xiao, Chaojun, et al.
Veröffentlicht: (2023)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2023)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms
von: Yan, Shilin, et al.
Veröffentlicht: (2025)
von: Yan, Shilin, et al.
Veröffentlicht: (2025)
S$^3$Attention: Improving Long Sequence Attention with Smoothed Skeleton Sketching
von: Wang, Xue, et al.
Veröffentlicht: (2024)
von: Wang, Xue, et al.
Veröffentlicht: (2024)
LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information
von: Ping, Bowen, et al.
Veröffentlicht: (2025)
von: Ping, Bowen, et al.
Veröffentlicht: (2025)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
von: Zhong, Yangyang, et al.
Veröffentlicht: (2025)
von: Zhong, Yangyang, et al.
Veröffentlicht: (2025)
Accelerate Parallelizable Reasoning via Parallel Decoding within One Sequence
von: Yu, Yijiong
Veröffentlicht: (2025)
von: Yu, Yijiong
Veröffentlicht: (2025)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025) -
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
von: Sun, Ao, et al.
Veröffentlicht: (2024) -
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024) -
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
von: Zhao, Weilin, et al.
Veröffentlicht: (2025) -
NOSA: Native and Offloadable Sparse Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)