Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Zirun, Hong, Minjie, Jin, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
by: Hong, Minjie, et al.
Published: (2025)
by: Hong, Minjie, et al.
Published: (2025)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
by: Deng, Naihao, et al.
Published: (2024)
by: Deng, Naihao, et al.
Published: (2024)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
by: Xiao, Tong, et al.
Published: (2025)
by: Xiao, Tong, et al.
Published: (2025)
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
by: Lillemark, Hansen Jin, et al.
Published: (2026)
by: Lillemark, Hansen Jin, et al.
Published: (2026)
Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptation
by: He, Yongbo, et al.
Published: (2026)
by: He, Yongbo, et al.
Published: (2026)
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
by: Yang, Enneng, et al.
Published: (2024)
by: Yang, Enneng, et al.
Published: (2024)
Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
by: Xu, Chenhui, et al.
Published: (2025)
by: Xu, Chenhui, et al.
Published: (2025)
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
by: Duan, Chengqi, et al.
Published: (2025)
by: Duan, Chengqi, et al.
Published: (2025)
MLLMs-Augmented Visual-Language Representation Learning
by: Liu, Yanqing, et al.
Published: (2023)
by: Liu, Yanqing, et al.
Published: (2023)
Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
A More Word-like Image Tokenization for MLLMs
by: Lee, Hyun, et al.
Published: (2026)
by: Lee, Hyun, et al.
Published: (2026)
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
by: Liu, Ziyu, et al.
Published: (2024)
by: Liu, Ziyu, et al.
Published: (2024)
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
by: V Team, et al.
Published: (2025)
by: V Team, et al.
Published: (2025)
Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning
by: Kim, Kyungsoo, et al.
Published: (2025)
by: Kim, Kyungsoo, et al.
Published: (2025)
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner
by: Zhang, Wenchuan, et al.
Published: (2025)
by: Zhang, Wenchuan, et al.
Published: (2025)
HueManity: Probing Fine-Grained Visual Perception in MLLMs
by: Grover, Rynaa, et al.
Published: (2025)
by: Grover, Rynaa, et al.
Published: (2025)
MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models
by: Shi, Yang, et al.
Published: (2026)
by: Shi, Yang, et al.
Published: (2026)
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
by: Zhang, Jingyi, et al.
Published: (2025)
by: Zhang, Jingyi, et al.
Published: (2025)
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
by: Chen, Shuang, et al.
Published: (2025)
by: Chen, Shuang, et al.
Published: (2025)
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
by: Zhu, Yinglun, et al.
Published: (2025)
by: Zhu, Yinglun, et al.
Published: (2025)
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
by: Barrios, Wayner, et al.
Published: (2025)
by: Barrios, Wayner, et al.
Published: (2025)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
by: Zheng, Naishan, et al.
Published: (2025)
by: Zheng, Naishan, et al.
Published: (2025)
Multimodal Information Fusion for Chart Understanding: A Survey of MLLMs -- Evolution, Limitations, and Cognitive Enhancement
by: Yi, Zhihang, et al.
Published: (2026)
by: Yi, Zhihang, et al.
Published: (2026)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
by: Wu, Boyong, et al.
Published: (2026)
by: Wu, Boyong, et al.
Published: (2026)
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
by: Li, Shaoxuan, et al.
Published: (2026)
by: Li, Shaoxuan, et al.
Published: (2026)
Reinforcement Learning with Generalizable Gaussian Splatting
by: Wang, Jiaxu, et al.
Published: (2024)
by: Wang, Jiaxu, et al.
Published: (2024)
Tiny Machine Learning: Progress and Futures
by: Lin, Ji, et al.
Published: (2024)
by: Lin, Ji, et al.
Published: (2024)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs
by: Hu, Zixuan, et al.
Published: (2024)
by: Hu, Zixuan, et al.
Published: (2024)
VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
Similar Items
-
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
by: Hong, Minjie, et al.
Published: (2025) -
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
by: Deng, Naihao, et al.
Published: (2024) -
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
by: Xiao, Tong, et al.
Published: (2025) -
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
by: Guo, Zirun, et al.
Published: (2025) -
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
by: Guo, Zirun, et al.
Published: (2025)