Video Finetuning Improves Reasoning Between Frames
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Ruiqi, Yun, Tian, Wang, Zihan, Pavlick, Ellie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Do Vision-Language Models Process Conflicting Information Across Modalities?
by: Hua, Tianze, et al.
Published: (2025)
by: Hua, Tianze, et al.
Published: (2025)
Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
by: Lepori, Michael A., et al.
Published: (2024)
by: Lepori, Michael A., et al.
Published: (2024)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
by: Ge, Haonan, et al.
Published: (2025)
by: Ge, Haonan, et al.
Published: (2025)
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
by: Tian, Shulin, et al.
Published: (2025)
by: Tian, Shulin, et al.
Published: (2025)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models
by: Jeon, Wooseok, et al.
Published: (2026)
by: Jeon, Wooseok, et al.
Published: (2026)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
by: Yang, Bowen, et al.
Published: (2025)
by: Yang, Bowen, et al.
Published: (2025)
Deep Neural Networks Can Learn Generalizable Same-Different Visual Relations
by: Tartaglini, Alexa R., et al.
Published: (2023)
by: Tartaglini, Alexa R., et al.
Published: (2023)
Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
by: Cai, Peiliang, et al.
Published: (2026)
by: Cai, Peiliang, et al.
Published: (2026)
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
by: Zhong, Tianxiong, et al.
Published: (2025)
by: Zhong, Tianxiong, et al.
Published: (2025)
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
by: Lewis, Martha, et al.
Published: (2022)
by: Lewis, Martha, et al.
Published: (2022)
AFFAKT: A Hierarchical Optimal Transport based Method for Affective Facial Knowledge Transfer in Video Deception Detection
by: Ji, Zihan, et al.
Published: (2024)
by: Ji, Zihan, et al.
Published: (2024)
FrameBridge: Improving Image-to-Video Generation with Bridge Models
by: Wang, Yuji, et al.
Published: (2024)
by: Wang, Yuji, et al.
Published: (2024)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
BehaviorVLM: Unified Finetuning-Free Behavioral Understanding with Vision-Language Reasoning
by: Ke, Jingyang, et al.
Published: (2026)
by: Ke, Jingyang, et al.
Published: (2026)
A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
by: Yi, Faliu, et al.
Published: (2025)
by: Yi, Faliu, et al.
Published: (2025)
M-LLM Based Video Frame Selection for Efficient Video Understanding
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
VFIMamba: Video Frame Interpolation with State Space Models
by: Zhang, Guozhen, et al.
Published: (2024)
by: Zhang, Guozhen, et al.
Published: (2024)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
by: Jang, Sangwon, et al.
Published: (2025)
by: Jang, Sangwon, et al.
Published: (2025)
Directing the Narrative: A Finetuning Method for Controlling Coherence and Style in Story Generation
by: Zhang, Jianzhang, et al.
Published: (2026)
by: Zhang, Jianzhang, et al.
Published: (2026)
Demystifying Video Reasoning
by: Wang, Ruisi, et al.
Published: (2026)
by: Wang, Ruisi, et al.
Published: (2026)
Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI
by: Moukheiber, Lama, et al.
Published: (2026)
by: Moukheiber, Lama, et al.
Published: (2026)
DLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation
by: Yuan, Zhihang, et al.
Published: (2025)
by: Yuan, Zhihang, et al.
Published: (2025)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
by: Yang, Xuyi, et al.
Published: (2025)
by: Yang, Xuyi, et al.
Published: (2025)
Detecting AI-Generated Video via Frame Consistency
by: Ma, Long, et al.
Published: (2024)
by: Ma, Long, et al.
Published: (2024)
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
by: Zeng, Qinglin, et al.
Published: (2025)
by: Zeng, Qinglin, et al.
Published: (2025)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
by: Ding, Yang, et al.
Published: (2025)
by: Ding, Yang, et al.
Published: (2025)
Boundary Matters: A Bi-Level Active Finetuning Framework
by: Lu, Han, et al.
Published: (2024)
by: Lu, Han, et al.
Published: (2024)
What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models
by: Yun, Tian, et al.
Published: (2025)
by: Yun, Tian, et al.
Published: (2025)
Process-of-Thought Reasoning for Videos
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
by: Han, Songhao, et al.
Published: (2024)
by: Han, Songhao, et al.
Published: (2024)
MISS: A Generative Pretraining and Finetuning Approach for Med-VQA
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Investigation of Frame Differences as Motion Cues for Video Object Segmentation
by: Kawamura, Sota, et al.
Published: (2025)
by: Kawamura, Sota, et al.
Published: (2025)
TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning
by: Liu, Junhua, et al.
Published: (2026)
by: Liu, Junhua, et al.
Published: (2026)
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
by: Liu, Chengwen, et al.
Published: (2026)
by: Liu, Chengwen, et al.
Published: (2026)
Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation
by: Wang, Hengyi, et al.
Published: (2026)
by: Wang, Hengyi, et al.
Published: (2026)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
by: Deng, Andong, et al.
Published: (2024)
by: Deng, Andong, et al.
Published: (2024)
VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
by: Cao, Xinye, et al.
Published: (2025)
by: Cao, Xinye, et al.
Published: (2025)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
by: Deng, Andong, et al.
Published: (2025)
by: Deng, Andong, et al.
Published: (2025)
Re-Identifying Kākā with AI-Automated Video Key Frame Extraction
by: Maddigan, Paula, et al.
Published: (2025)
by: Maddigan, Paula, et al.
Published: (2025)
Similar Items
-
How Do Vision-Language Models Process Conflicting Information Across Modalities?
by: Hua, Tianze, et al.
Published: (2025) -
Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
by: Lepori, Michael A., et al.
Published: (2024) -
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
by: Ge, Haonan, et al.
Published: (2025) -
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
by: Tian, Shulin, et al.
Published: (2025) -
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)