Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chi, Qiu, Haibo, Zhang, Qiming, Xu, Yufei, Zeng, Zhixiong, Yang, Siqi, Shi, Peng, Ma, Lin, Zhang, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
by: Yang, Siqi, et al.
Published: (2025)
by: Yang, Siqi, et al.
Published: (2025)
Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
by: Chen, Kun, et al.
Published: (2025)
by: Chen, Kun, et al.
Published: (2025)
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2025)
by: Ma, Xiaoxiao, et al.
Published: (2025)
Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
by: Lan, Xiaohan, et al.
Published: (2025)
by: Lan, Xiaohan, et al.
Published: (2025)
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
by: Zhao, Xuanle, et al.
Published: (2025)
by: Zhao, Xuanle, et al.
Published: (2025)
Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
by: Chen, Lei, et al.
Published: (2025)
by: Chen, Lei, et al.
Published: (2025)
Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
by: Chen, Lei, et al.
Published: (2025)
by: Chen, Lei, et al.
Published: (2025)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
by: Qiu, Haibo, et al.
Published: (2025)
by: Qiu, Haibo, et al.
Published: (2025)
Contrastive Learning for Multimodal Human Activity Recognition with Limited Labeled Data
by: Jing, Long, et al.
Published: (2026)
by: Jing, Long, et al.
Published: (2026)
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
by: Zhang, Shuoshuo, et al.
Published: (2025)
by: Zhang, Shuoshuo, et al.
Published: (2025)
Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code Generation
by: Zhou, Jiawei, et al.
Published: (2026)
by: Zhou, Jiawei, et al.
Published: (2026)
TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution
by: Jiang, Deyang, et al.
Published: (2026)
by: Jiang, Deyang, et al.
Published: (2026)
DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
by: Yang, Longrong, et al.
Published: (2025)
by: Yang, Longrong, et al.
Published: (2025)
Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
by: Wang, Peiyu, et al.
Published: (2025)
by: Wang, Peiyu, et al.
Published: (2025)
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
by: Zhang, Haoji, et al.
Published: (2025)
by: Zhang, Haoji, et al.
Published: (2025)
Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models
by: Peng, Ruiying, et al.
Published: (2026)
by: Peng, Ruiying, et al.
Published: (2026)
InstructVEdit: A Holistic Approach for Instructional Video Editing
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
by: Ma, Xiaoxiao, et al.
Published: (2025)
by: Ma, Xiaoxiao, et al.
Published: (2025)
Seirênes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
by: Tian, Changyuan, et al.
Published: (2026)
by: Tian, Changyuan, et al.
Published: (2026)
Unleashing the Power of Generic Segmentation Models: A Simple Baseline for Infrared Small Target Detection
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)
by: He, Haibin, et al.
Published: (2026)
Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning
by: Xing, Ximing, et al.
Published: (2025)
by: Xing, Ximing, et al.
Published: (2025)
SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning
by: Yang, Xiao, et al.
Published: (2026)
by: Yang, Xiao, et al.
Published: (2026)
LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models
by: Yao, Ruilin, et al.
Published: (2025)
by: Yao, Ruilin, et al.
Published: (2025)
MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
by: Meng, Fanqing, et al.
Published: (2025)
by: Meng, Fanqing, et al.
Published: (2025)
Synthesizing Images on Perceptual Boundaries of ANNs for Uncovering Human Perceptual Variability on Facial Expressions
by: Deng, Haotian, et al.
Published: (2025)
by: Deng, Haotian, et al.
Published: (2025)
Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner
by: Zhang, Wenchuan, et al.
Published: (2025)
by: Zhang, Wenchuan, et al.
Published: (2025)
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
by: Zhang, Gengwei, et al.
Published: (2026)
by: Zhang, Gengwei, et al.
Published: (2026)
Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
by: Yang, Ruichao, et al.
Published: (2026)
by: Yang, Ruichao, et al.
Published: (2026)
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
by: Ma, Zehong, et al.
Published: (2026)
by: Ma, Zehong, et al.
Published: (2026)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
by: Yan, Zhonghao, et al.
Published: (2025)
by: Yan, Zhonghao, et al.
Published: (2025)
Na-IRSTD: Enhancing Infrared Small Target Detection via Native-Resolution Feature Selection and Fusion
by: Xu, Qian, et al.
Published: (2026)
by: Xu, Qian, et al.
Published: (2026)
MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving
by: Zhang, Lingjun, et al.
Published: (2026)
by: Zhang, Lingjun, et al.
Published: (2026)
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
by: Li, Zongzhao, et al.
Published: (2025)
by: Li, Zongzhao, et al.
Published: (2025)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
by: Lee, Jonathan, et al.
Published: (2025)
by: Lee, Jonathan, et al.
Published: (2025)
Similar Items
-
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
by: Zhang, Chi, et al.
Published: (2025) -
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
by: Yang, Siqi, et al.
Published: (2025) -
Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
by: Chen, Kun, et al.
Published: (2025) -
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025) -
STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2025)