Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Guo, Zirun, Hong, Minjie, Jin, Tao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
por: Hong, Minjie, et al.
Publicado: (2025)
por: Hong, Minjie, et al.
Publicado: (2025)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
por: Deng, Naihao, et al.
Publicado: (2024)
por: Deng, Naihao, et al.
Publicado: (2024)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
por: Xiao, Tong, et al.
Publicado: (2025)
por: Xiao, Tong, et al.
Publicado: (2025)
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
por: Guo, Zirun, et al.
Publicado: (2025)
por: Guo, Zirun, et al.
Publicado: (2025)
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
por: Guo, Zirun, et al.
Publicado: (2025)
por: Guo, Zirun, et al.
Publicado: (2025)
R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning
por: Yang, Qi, et al.
Publicado: (2025)
por: Yang, Qi, et al.
Publicado: (2025)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
por: Li, Chenghao, et al.
Publicado: (2026)
por: Li, Chenghao, et al.
Publicado: (2026)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
por: AI, Inclusion, et al.
Publicado: (2025)
por: AI, Inclusion, et al.
Publicado: (2025)
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
por: Lillemark, Hansen Jin, et al.
Publicado: (2026)
por: Lillemark, Hansen Jin, et al.
Publicado: (2026)
Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptation
por: He, Yongbo, et al.
Publicado: (2026)
por: He, Yongbo, et al.
Publicado: (2026)
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
por: Yang, Enneng, et al.
Publicado: (2024)
por: Yang, Enneng, et al.
Publicado: (2024)
Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
por: Xu, Chenhui, et al.
Publicado: (2025)
por: Xu, Chenhui, et al.
Publicado: (2025)
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
por: Guo, Zirun, et al.
Publicado: (2024)
por: Guo, Zirun, et al.
Publicado: (2024)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
por: Duan, Chengqi, et al.
Publicado: (2025)
por: Duan, Chengqi, et al.
Publicado: (2025)
MLLMs-Augmented Visual-Language Representation Learning
por: Liu, Yanqing, et al.
Publicado: (2023)
por: Liu, Yanqing, et al.
Publicado: (2023)
Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
por: Guo, Zirun, et al.
Publicado: (2024)
por: Guo, Zirun, et al.
Publicado: (2024)
A More Word-like Image Tokenization for MLLMs
por: Lee, Hyun, et al.
Publicado: (2026)
por: Lee, Hyun, et al.
Publicado: (2026)
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
por: Liu, Ziyu, et al.
Publicado: (2024)
por: Liu, Ziyu, et al.
Publicado: (2024)
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
por: V Team, et al.
Publicado: (2025)
por: V Team, et al.
Publicado: (2025)
Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning
por: Kim, Kyungsoo, et al.
Publicado: (2025)
por: Kim, Kyungsoo, et al.
Publicado: (2025)
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
por: Guo, Zirun, et al.
Publicado: (2025)
por: Guo, Zirun, et al.
Publicado: (2025)
Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner
por: Zhang, Wenchuan, et al.
Publicado: (2025)
por: Zhang, Wenchuan, et al.
Publicado: (2025)
HueManity: Probing Fine-Grained Visual Perception in MLLMs
por: Grover, Rynaa, et al.
Publicado: (2025)
por: Grover, Rynaa, et al.
Publicado: (2025)
MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models
por: Shi, Yang, et al.
Publicado: (2026)
por: Shi, Yang, et al.
Publicado: (2026)
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
por: Zhang, Jingyi, et al.
Publicado: (2025)
por: Zhang, Jingyi, et al.
Publicado: (2025)
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
por: Chen, Shuang, et al.
Publicado: (2025)
por: Chen, Shuang, et al.
Publicado: (2025)
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
por: Zhu, Yinglun, et al.
Publicado: (2025)
por: Zhu, Yinglun, et al.
Publicado: (2025)
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
por: Barrios, Wayner, et al.
Publicado: (2025)
por: Barrios, Wayner, et al.
Publicado: (2025)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
por: Zheng, Naishan, et al.
Publicado: (2025)
por: Zheng, Naishan, et al.
Publicado: (2025)
Multimodal Information Fusion for Chart Understanding: A Survey of MLLMs -- Evolution, Limitations, and Cognitive Enhancement
por: Yi, Zhihang, et al.
Publicado: (2026)
por: Yi, Zhihang, et al.
Publicado: (2026)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
por: Wu, Boyong, et al.
Publicado: (2026)
por: Wu, Boyong, et al.
Publicado: (2026)
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
por: Li, Shaoxuan, et al.
Publicado: (2026)
por: Li, Shaoxuan, et al.
Publicado: (2026)
Reinforcement Learning with Generalizable Gaussian Splatting
por: Wang, Jiaxu, et al.
Publicado: (2024)
por: Wang, Jiaxu, et al.
Publicado: (2024)
Tiny Machine Learning: Progress and Futures
por: Lin, Ji, et al.
Publicado: (2024)
por: Lin, Ji, et al.
Publicado: (2024)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
por: Qiu, Yansheng, et al.
Publicado: (2025)
por: Qiu, Yansheng, et al.
Publicado: (2025)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
por: Chen, Yi, et al.
Publicado: (2025)
por: Chen, Yi, et al.
Publicado: (2025)
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
por: Wei, Lai, et al.
Publicado: (2025)
por: Wei, Lai, et al.
Publicado: (2025)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
por: Chen, Yang, et al.
Publicado: (2024)
por: Chen, Yang, et al.
Publicado: (2024)
Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs
por: Hu, Zixuan, et al.
Publicado: (2024)
por: Hu, Zixuan, et al.
Publicado: (2024)
VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
por: Yu, Xinlei, et al.
Publicado: (2025)
por: Yu, Xinlei, et al.
Publicado: (2025)
Ejemplares similares
-
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
por: Hong, Minjie, et al.
Publicado: (2025) -
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
por: Deng, Naihao, et al.
Publicado: (2024) -
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
por: Xiao, Tong, et al.
Publicado: (2025) -
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
por: Guo, Zirun, et al.
Publicado: (2025) -
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
por: Guo, Zirun, et al.
Publicado: (2025)