Improving LLM Video Understanding with 16 Frames Per Second
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yixuan, Tang, Changli, Zhuang, Jimin, Yang, Yudong, Sun, Guangzhi, Li, Wei, Ma, Zejun, Zhang, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Audio-centric Video Understanding Benchmark without Text Shortcut
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
by: Sun, Guangzhi, et al.
Published: (2025)
by: Sun, Guangzhi, et al.
Published: (2025)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
by: Tang, Changli, et al.
Published: (2025)
by: Tang, Changli, et al.
Published: (2025)
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
by: Tang, Changli, et al.
Published: (2024)
by: Tang, Changli, et al.
Published: (2024)
video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
by: Sun, Guangzhi, et al.
Published: (2025)
by: Sun, Guangzhi, et al.
Published: (2025)
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
by: Sun, Guangzhi, et al.
Published: (2024)
by: Sun, Guangzhi, et al.
Published: (2024)
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
by: Chen, Joya, et al.
Published: (2025)
by: Chen, Joya, et al.
Published: (2025)
JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues
by: Ji, Jiayi, et al.
Published: (2023)
by: Ji, Jiayi, et al.
Published: (2023)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding
by: Ma, Junpeng, et al.
Published: (2026)
by: Ma, Junpeng, et al.
Published: (2026)
M-LLM Based Video Frame Selection for Efficient Video Understanding
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
Event-Anchored Frame Selection for Effective Long-Video Understanding
by: Chen, Wang, et al.
Published: (2026)
by: Chen, Wang, et al.
Published: (2026)
FANeRV: Frequency Separation and Augmentation based Neural Representation for Video
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
by: Rahman, Aimon, et al.
Published: (2024)
by: Rahman, Aimon, et al.
Published: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
Velocity Disambiguation for Video Frame Interpolation
by: Zhong, Zhihang, et al.
Published: (2023)
by: Zhong, Zhihang, et al.
Published: (2023)
VideoScan: Enabling Efficient Streaming Video Understanding via Frame-level Semantic Carriers
by: Li, Ruanjun, et al.
Published: (2025)
by: Li, Ruanjun, et al.
Published: (2025)
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
by: Tang, Changli, et al.
Published: (2025)
by: Tang, Changli, et al.
Published: (2025)
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
by: Chen, Lin, et al.
Published: (2024)
by: Chen, Lin, et al.
Published: (2024)
Shot-Aware Frame Sampling for Video Understanding
by: Zhao, Mengyu, et al.
Published: (2026)
by: Zhao, Mengyu, et al.
Published: (2026)
Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
by: Cai, Peiliang, et al.
Published: (2026)
by: Cai, Peiliang, et al.
Published: (2026)
Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial Frame
by: Cao, Qinglong, et al.
Published: (2024)
by: Cao, Qinglong, et al.
Published: (2024)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
by: Sun, Shangkun, et al.
Published: (2024)
by: Sun, Shangkun, et al.
Published: (2024)
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy
by: Tang, Zuojin, et al.
Published: (2026)
by: Tang, Zuojin, et al.
Published: (2026)
Generalizable Implicit Motion Modeling for Video Frame Interpolation
by: Guo, Zujin, et al.
Published: (2024)
by: Guo, Zujin, et al.
Published: (2024)
WikiStyle+: A Multimodal Approach to Content-Style Representation Disentanglement for Artistic Image Stylization
by: Zhuoqi, Ma, et al.
Published: (2024)
by: Zhuoqi, Ma, et al.
Published: (2024)
OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data
by: Wei, Runpu, et al.
Published: (2025)
by: Wei, Runpu, et al.
Published: (2025)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
by: Ghazanfari, Sara, et al.
Published: (2025)
by: Ghazanfari, Sara, et al.
Published: (2025)
Video Frame Interpolation for Polarization via Swin-Transformer
by: Huang, Feng, et al.
Published: (2024)
by: Huang, Feng, et al.
Published: (2024)
Video Finetuning Improves Reasoning Between Frames
by: Yang, Ruiqi, et al.
Published: (2025)
by: Yang, Ruiqi, et al.
Published: (2025)
Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos
by: Luo, Rundong, et al.
Published: (2025)
by: Luo, Rundong, et al.
Published: (2025)
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
by: Liu, Ruyang, et al.
Published: (2025)
by: Liu, Ruyang, et al.
Published: (2025)
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
by: Zhao, Henghao, et al.
Published: (2025)
by: Zhao, Henghao, et al.
Published: (2025)
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
by: Tang, Changli, et al.
Published: (2026)
by: Tang, Changli, et al.
Published: (2026)
Video2LoRA: Unified Semantic-Controlled Video Generation via Per-Reference-Video LoRA
by: Wu, Zexi, et al.
Published: (2026)
by: Wu, Zexi, et al.
Published: (2026)
Video Prediction Transformers without Recurrence or Convolution
by: Tang, Yujin, et al.
Published: (2024)
by: Tang, Yujin, et al.
Published: (2024)
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
by: Bao, Xiaoyi, et al.
Published: (2025)
by: Bao, Xiaoyi, et al.
Published: (2025)
DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
by: Song, Zhende, et al.
Published: (2024)
by: Song, Zhende, et al.
Published: (2024)
GenCompositor: Generative Video Compositing with Diffusion Transformer
by: Yang, Shuzhou, et al.
Published: (2025)
by: Yang, Shuzhou, et al.
Published: (2025)
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
by: Tan, Wenhui, et al.
Published: (2026)
by: Tan, Wenhui, et al.
Published: (2026)
Similar Items
-
Audio-centric Video Understanding Benchmark without Text Shortcut
by: Yang, Yudong, et al.
Published: (2025) -
video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
by: Sun, Guangzhi, et al.
Published: (2025) -
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
by: Tang, Changli, et al.
Published: (2025) -
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
by: Tang, Changli, et al.
Published: (2024) -
video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
by: Sun, Guangzhi, et al.
Published: (2025)