AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Shuheng, Liu, Yuqi, Zhou, Hongbo, Peng, Jun, Zhou, Yiyi, Sun, Xiaoshuai, Ji, Rongrong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
von: Chen, Tao, et al.
Veröffentlicht: (2026)
von: Chen, Tao, et al.
Veröffentlicht: (2026)
AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies
von: Hu, Xixi, et al.
Veröffentlicht: (2024)
von: Hu, Xixi, et al.
Veröffentlicht: (2024)
Image Captioning via Dynamic Path Customization
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
von: Zhang, Xian, et al.
Veröffentlicht: (2025)
von: Zhang, Xian, et al.
Veröffentlicht: (2025)
Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
von: Zhang, Jinlu, et al.
Veröffentlicht: (2024)
von: Zhang, Jinlu, et al.
Veröffentlicht: (2024)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
von: Zou, Siyu, et al.
Veröffentlicht: (2024)
von: Zou, Siyu, et al.
Veröffentlicht: (2024)
Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2024)
von: Luo, Gen, et al.
Veröffentlicht: (2024)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
von: Luo, Yaxin, et al.
Veröffentlicht: (2024)
von: Luo, Yaxin, et al.
Veröffentlicht: (2024)
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
von: Chen, Tao, et al.
Veröffentlicht: (2025)
von: Chen, Tao, et al.
Veröffentlicht: (2025)
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
von: Song, Baiyang, et al.
Veröffentlicht: (2026)
von: Song, Baiyang, et al.
Veröffentlicht: (2026)
Global2Local: A Joint-Hierarchical Attention for Video Captioning
von: Dai, Chengpeng, et al.
Veröffentlicht: (2022)
von: Dai, Chengpeng, et al.
Veröffentlicht: (2022)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
von: Li, Handong, et al.
Veröffentlicht: (2026)
von: Li, Handong, et al.
Veröffentlicht: (2026)
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
von: Tong, Bo, et al.
Veröffentlicht: (2024)
von: Tong, Bo, et al.
Veröffentlicht: (2024)
Deep Instruction Tuning for Segment Anything Model
von: Huang, Xiaorui, et al.
Veröffentlicht: (2024)
von: Huang, Xiaorui, et al.
Veröffentlicht: (2024)
Adaptive Keyframe Sampling for Long Video Understanding
von: Tang, Xi, et al.
Veröffentlicht: (2025)
von: Tang, Xi, et al.
Veröffentlicht: (2025)
Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing
von: Liu, Lin, et al.
Veröffentlicht: (2026)
von: Liu, Lin, et al.
Veröffentlicht: (2026)
Grounded Chain-of-Thought for Multimodal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
Omni-Referring Image Segmentation
von: Zheng, Qiancheng, et al.
Veröffentlicht: (2025)
von: Zheng, Qiancheng, et al.
Veröffentlicht: (2025)
Adaptive Keyframe Selection for Scalable 3D Scene Reconstruction in Dynamic Environments
von: Jha, Raman, et al.
Veröffentlicht: (2025)
von: Jha, Raman, et al.
Veröffentlicht: (2025)
AdaFlow: Opportunistic Inference on Asynchronous Mobile Data with Generalized Affinity Control
von: Wu, Fenmin, et al.
Veröffentlicht: (2024)
von: Wu, Fenmin, et al.
Veröffentlicht: (2024)
Any-to-3D Generation via Hybrid Diffusion Supervision
von: Fan, Yijun, et al.
Veröffentlicht: (2024)
von: Fan, Yijun, et al.
Veröffentlicht: (2024)
I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
AdaEdit: Adaptive Temporal and Channel Modulation for Flow-Based Image Editing
von: Li, Guandong, et al.
Veröffentlicht: (2026)
von: Li, Guandong, et al.
Veröffentlicht: (2026)
Routing Experts: Learning to Route Dynamic Experts in Multi-modal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
StealthDiffusion: Towards Evading Diffusion Forensic Detection through Diffusion Model
von: Zhou, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhou, Ziyin, et al.
Veröffentlicht: (2024)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
Mixed Degradation Image Restoration via Local Dynamic Optimization and Conditional Embedding
von: Gu, Yubin, et al.
Veröffentlicht: (2024)
von: Gu, Yubin, et al.
Veröffentlicht: (2024)
M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning
von: Ma, Yiwei, et al.
Veröffentlicht: (2025)
von: Ma, Yiwei, et al.
Veröffentlicht: (2025)
HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention
von: Zheng, Xuzhe, et al.
Veröffentlicht: (2026)
von: Zheng, Xuzhe, et al.
Veröffentlicht: (2026)
MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
von: Fang, Bo, et al.
Veröffentlicht: (2025)
von: Fang, Bo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
von: Chen, Tao, et al.
Veröffentlicht: (2026) -
AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies
von: Hu, Xixi, et al.
Veröffentlicht: (2024) -
Image Captioning via Dynamic Path Customization
von: Ma, Yiwei, et al.
Veröffentlicht: (2024) -
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
von: Zhang, Xian, et al.
Veröffentlicht: (2025) -
Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
von: Zhang, Jinlu, et al.
Veröffentlicht: (2024)