Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Ghazanfari, Sara, Croce, Francesco, Flammarion, Nicolas, Krishnamurthy, Prashanth, Khorrami, Farshad, Garg, Siddharth |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
by: Ghazanfari, Sara, et al.
Published: (2024)
by: Ghazanfari, Sara, et al.
Published: (2024)
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
by: Ghazanfari, Sara, et al.
Published: (2026)
by: Ghazanfari, Sara, et al.
Published: (2026)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
by: Ghazanfari, Sara, et al.
Published: (2024)
by: Ghazanfari, Sara, et al.
Published: (2024)
LipSim: A Provably Robust Perceptual Similarity Metric
by: Ghazanfari, Sara, et al.
Published: (2023)
by: Ghazanfari, Sara, et al.
Published: (2023)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
by: Neuhaus, Yannic, et al.
Published: (2026)
by: Neuhaus, Yannic, et al.
Published: (2026)
RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
by: Patel, Naman, et al.
Published: (2025)
by: Patel, Naman, et al.
Published: (2025)
CLIPScope: Enhancing Zero-Shot OOD Detection with Bayesian Scoring
by: Fu, Hao, et al.
Published: (2024)
by: Fu, Hao, et al.
Published: (2024)
FlashMix: Fast Map-Free LiDAR Localization via Feature Mixing and Contrastive-Constrained Accelerated Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
by: Schlarmann, Christian, et al.
Published: (2025)
by: Schlarmann, Christian, et al.
Published: (2025)
Efficient and Distributed Large-Scale 3D Map Registration using Tomographic Features
by: Unlu, Halil Utku, et al.
Published: (2024)
by: Unlu, Halil Utku, et al.
Published: (2024)
SALSA: Swift Adaptive Lightweight Self-Attention for Enhanced LiDAR Place Recognition
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
Shot-Aware Frame Sampling for Video Understanding
by: Zhao, Mengyu, et al.
Published: (2026)
by: Zhao, Mengyu, et al.
Published: (2026)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
On the Adversarial Robustness of Discrete Image Tokenizers
by: Bhagwatkar, Rishika, et al.
Published: (2026)
by: Bhagwatkar, Rishika, et al.
Published: (2026)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
by: Rahman, Aimon, et al.
Published: (2024)
by: Rahman, Aimon, et al.
Published: (2024)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
by: Ge, Haonan, et al.
Published: (2025)
by: Ge, Haonan, et al.
Published: (2025)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
Motion-Aware Video Frame Interpolation
by: Han, Pengfei, et al.
Published: (2024)
by: Han, Pengfei, et al.
Published: (2024)
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
by: Zhang, Shaojie, et al.
Published: (2025)
by: Zhang, Shaojie, et al.
Published: (2025)
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
by: Lin, Shih-Yao, et al.
Published: (2025)
by: Lin, Shih-Yao, et al.
Published: (2025)
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
by: Wang, Yuan, et al.
Published: (2026)
by: Wang, Yuan, et al.
Published: (2026)
Object-Aware Video Matting with Cross-Frame Guidance
by: Zhang, Huayu, et al.
Published: (2025)
by: Zhang, Huayu, et al.
Published: (2025)
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
VideoScan: Enabling Efficient Streaming Video Understanding via Frame-level Semantic Carriers
by: Li, Ruanjun, et al.
Published: (2025)
by: Li, Ruanjun, et al.
Published: (2025)
HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
by: Yang, Yiqing, et al.
Published: (2025)
by: Yang, Yiqing, et al.
Published: (2025)
GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding
by: Ma, Junpeng, et al.
Published: (2026)
by: Ma, Junpeng, et al.
Published: (2026)
Improving LLM Video Understanding with 16 Frames Per Second
by: Li, Yixuan, et al.
Published: (2025)
by: Li, Yixuan, et al.
Published: (2025)
Event-Anchored Frame Selection for Effective Long-Video Understanding
by: Chen, Wang, et al.
Published: (2026)
by: Chen, Wang, et al.
Published: (2026)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
Video Finetuning Improves Reasoning Between Frames
by: Yang, Ruiqi, et al.
Published: (2025)
by: Yang, Ruiqi, et al.
Published: (2025)
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
Benchmarking Video Frame Interpolation
by: Kiefhaber, Simon, et al.
Published: (2024)
by: Kiefhaber, Simon, et al.
Published: (2024)
Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026)
AceVFI: A Comprehensive Survey of Advances in Video Frame Interpolation
by: Kye, Dahyeon, et al.
Published: (2025)
by: Kye, Dahyeon, et al.
Published: (2025)
Similar Items
-
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
by: Ghazanfari, Sara, et al.
Published: (2024) -
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
by: Ghazanfari, Sara, et al.
Published: (2026) -
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
by: Ghazanfari, Sara, et al.
Published: (2024) -
LipSim: A Provably Robust Perceptual Similarity Metric
by: Ghazanfari, Sara, et al.
Published: (2023) -
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
by: Neuhaus, Yannic, et al.
Published: (2026)