LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Shenghao, Yang, Qize, Li, Yuan-Ming, Wei, Xihan, Xie, Xiaohua, Zheng, Wei-Shi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
by: Fu, Shenghao, et al.
Published: (2024)
by: Fu, Shenghao, et al.
Published: (2024)
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
by: Peng, Yi-Xing, et al.
Published: (2025)
by: Peng, Yi-Xing, et al.
Published: (2025)
ViSpeak: Visual Instruction Feedback in Streaming Videos
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
by: Li, Yuan-Ming, et al.
Published: (2025)
by: Li, Yuan-Ming, et al.
Published: (2025)
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
by: Xie, Yuan, et al.
Published: (2025)
by: Xie, Yuan, et al.
Published: (2025)
DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation
by: Yan, Junkai, et al.
Published: (2024)
by: Yan, Junkai, et al.
Published: (2024)
ObjEmbed: Towards Universal Multimodal Object Embeddings
by: Fu, Shenghao, et al.
Published: (2026)
by: Fu, Shenghao, et al.
Published: (2026)
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding
by: Li, Chenglin, et al.
Published: (2025)
by: Li, Chenglin, et al.
Published: (2025)
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
by: Sun, Boyuan, et al.
Published: (2025)
by: Sun, Boyuan, et al.
Published: (2025)
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
by: Yang, Zhongyu, et al.
Published: (2026)
by: Yang, Zhongyu, et al.
Published: (2026)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
by: Li, Chenglin, et al.
Published: (2025)
by: Li, Chenglin, et al.
Published: (2025)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025)
by: Zhang, Xintong, et al.
Published: (2025)
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
by: Jiang, Zhiyuan, et al.
Published: (2025)
by: Jiang, Zhiyuan, et al.
Published: (2025)
TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
by: Pan, Junwen, et al.
Published: (2025)
by: Pan, Junwen, et al.
Published: (2025)
ZoomR: Memory Efficient Reasoning through Multi-Granularity Key Value Retrieval
by: Yang, David H., et al.
Published: (2026)
by: Yang, David H., et al.
Published: (2026)
Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
by: Yang, Jiashu, et al.
Published: (2025)
by: Yang, Jiashu, et al.
Published: (2025)
Adaptive Keyframe Sampling for Long Video Understanding
by: Tang, Xi, et al.
Published: (2025)
by: Tang, Xi, et al.
Published: (2025)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
by: Yu, Jiashuo, et al.
Published: (2025)
by: Yu, Jiashuo, et al.
Published: (2025)
Advancing Adaptive Multi-Stage Video Anomaly Reasoning: A Benchmark Dataset and Method
by: Huang, Chao, et al.
Published: (2026)
by: Huang, Chao, et al.
Published: (2026)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
by: Wu, Jiaxin, et al.
Published: (2025)
by: Wu, Jiaxin, et al.
Published: (2025)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
by: Qiu, Jihao, et al.
Published: (2026)
by: Qiu, Jihao, et al.
Published: (2026)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
Semantic Step Prediction: Multi-Step Latent Forecasting in LLM Reasoning Trajectories via Step Sampling
by: Yuan, Yidi
Published: (2026)
by: Yuan, Yidi
Published: (2026)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
by: Jin, Hongbo, et al.
Published: (2025)
by: Jin, Hongbo, et al.
Published: (2025)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
by: Sun, Boyuan, et al.
Published: (2026)
by: Sun, Boyuan, et al.
Published: (2026)
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation
by: Tang, Zichen, et al.
Published: (2025)
by: Tang, Zichen, et al.
Published: (2025)
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
by: Yu, Xuan, et al.
Published: (2025)
by: Yu, Xuan, et al.
Published: (2025)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
by: Ma, David, et al.
Published: (2025)
by: Ma, David, et al.
Published: (2025)
RenderFlow: Single-Step Neural Rendering via Flow Matching
by: Zhang, Shenghao, et al.
Published: (2026)
by: Zhang, Shenghao, et al.
Published: (2026)
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
by: Wang, Zheng, et al.
Published: (2026)
by: Wang, Zheng, et al.
Published: (2026)
Similar Items
-
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
by: Fu, Shenghao, et al.
Published: (2024) -
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
by: Fu, Shenghao, et al.
Published: (2025) -
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
by: Peng, Yi-Xing, et al.
Published: (2025) -
ViSpeak: Visual Instruction Feedback in Streaming Videos
by: Fu, Shenghao, et al.
Published: (2025) -
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
by: Fu, Shenghao, et al.
Published: (2025)