Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Feiding, Zhang, Yongkang, Liao, Yuhao, Zeng, Zijian, Zhu, Chunzheng, Zheng, Yaozong, Liu, Yafei, Peng, Yeling, Wang, Youwei, Wang, Sibo, Yang, Huiming, Liao, Linglin, Yang, Shunzhi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
PoseStreamer: A Multi-modal Framework for 3D Tracking of Unseen Moving Objects
by: Yang, Huiming, et al.
Published: (2025)
by: Yang, Huiming, et al.
Published: (2025)
DexSim2Real: Foundation Model-Guided Sim-to-Real Transfer for Generalizable Dexterous Manipulation
by: Zeng, Zijian, et al.
Published: (2026)
by: Zeng, Zijian, et al.
Published: (2026)
Reducing Credit Assignment Variance via Counterfactual Reasoning Paths
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
Spatial-aware Symmetric Alignment for Text-guided Medical Image Segmentation
by: Liao, Linglin, et al.
Published: (2025)
by: Liao, Linglin, et al.
Published: (2025)
Robust Accelerated Adaptive Search: High-Probability Complexity Bounds under Bounded-Moment Stochastic Oracles
by: Zhang, Shunzhi, et al.
Published: (2026)
by: Zhang, Shunzhi, et al.
Published: (2026)
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
by: Ding, Fei, et al.
Published: (2025)
by: Ding, Fei, et al.
Published: (2025)
HELM: Harness-Enhanced Long-horizon Memory for Vision-Language-Action Manipulation
by: Zeng, Zijian, et al.
Published: (2026)
by: Zeng, Zijian, et al.
Published: (2026)
Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
Robust Insurance Pricing and Liquidity Management
by: Pang, Shunzhi
Published: (2025)
by: Pang, Shunzhi
Published: (2025)
Robust Investment-Driven Insurance Pricing under Correlation Ambiguity
by: Pang, Shunzhi
Published: (2026)
by: Pang, Shunzhi
Published: (2026)
Robust insurance pricing and liquidity management
by: Shunzhi Pang
Published: (2026)
by: Shunzhi Pang
Published: (2026)
CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection
by: Yang, Huiming, et al.
Published: (2025)
by: Yang, Huiming, et al.
Published: (2025)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
Data-Incremental Continual Offline Reinforcement Learning
by: Gai, Sibo, et al.
Published: (2024)
by: Gai, Sibo, et al.
Published: (2024)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
AlphaMath Almost Zero: Process Supervision without Process
by: Chen, Guoxin, et al.
Published: (2024)
by: Chen, Guoxin, et al.
Published: (2024)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
by: Wang, Yufei, et al.
Published: (2024)
by: Wang, Yufei, et al.
Published: (2024)
VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models
by: Kang, Shuhao, et al.
Published: (2026)
by: Kang, Shuhao, et al.
Published: (2026)
PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning
by: Wang, Yunxiao, et al.
Published: (2025)
by: Wang, Yunxiao, et al.
Published: (2025)
The equivalent condition for GRL codes to be MDS, AMDS or self-dual
by: Liang, Zhonghao, et al.
Published: (2025)
by: Liang, Zhonghao, et al.
Published: (2025)
The asymptotic estimation for two classes of generalized Fibonacci sub-sequences
by: Wan, Yongkang, et al.
Published: (2025)
by: Wan, Yongkang, et al.
Published: (2025)
The inverse of the (alternating) infinite sum of the reciprocal of the weighted sum for generalized Fibonacci sub-sequences
by: Wan, Yongkang, et al.
Published: (2025)
by: Wan, Yongkang, et al.
Published: (2025)
Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning
by: Zheng, Yaozong, et al.
Published: (2026)
by: Zheng, Yaozong, et al.
Published: (2026)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution
by: Zhu, Chunzheng, et al.
Published: (2026)
by: Zhu, Chunzheng, et al.
Published: (2026)
NeuSO: Neural Optimizer for Subgraph Queries
by: Yang, Linglin, et al.
Published: (2025)
by: Yang, Linglin, et al.
Published: (2025)
Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation
by: Zhu, Chunzheng, et al.
Published: (2026)
by: Zhu, Chunzheng, et al.
Published: (2026)
LLM-FK: Multi-Agent LLM Reasoning for Foreign Key Detection in Large-Scale Complex Databases
by: Tang, Zijian, et al.
Published: (2026)
by: Tang, Zijian, et al.
Published: (2026)
Hierarchical Cross-modal Prompt Learning for Vision-Language Models
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Annotation-Free Pathological Image Classification
by: Zhong, Lanfeng, et al.
Published: (2024)
by: Zhong, Lanfeng, et al.
Published: (2024)
Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
by: Luo, Zizhang, et al.
Published: (2026)
by: Luo, Zizhang, et al.
Published: (2026)
Mobile Recording Device Recognition Based Cross-Scale and Multi-Level Representation Learning
by: Zeng, Chunyan, et al.
Published: (2024)
by: Zeng, Chunyan, et al.
Published: (2024)
GrabS: Generative Embodied Agent for 3D Object Segmentation without Scene Supervision
by: Zhang, Zihui, et al.
Published: (2025)
by: Zhang, Zihui, et al.
Published: (2025)
A new class of self-orthogonal linear codes and their applications
by: Zhang, Yaozong, et al.
Published: (2023)
by: Zhang, Yaozong, et al.
Published: (2023)
Effect of Hydrogen Bonding Interaction on Kinetics of Cyclopentanol Reaction With Hydroperoxyl Radical at Atmospheric and Combustion Temperatures: A Theoretical Study
by: Yaozong Duan, et al.
Published: (2025)
by: Yaozong Duan, et al.
Published: (2025)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
Effects and Mechanisms of Extracellular Vesicles in Different Models of Acute Kidney Injury
by: Weidong Wang, et al.
Published: (2025)
by: Weidong Wang, et al.
Published: (2025)
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
Similar Items
-
Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction
by: Ding, Fei, et al.
Published: (2026) -
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning
by: Ding, Fei, et al.
Published: (2026) -
PoseStreamer: A Multi-modal Framework for 3D Tracking of Unseen Moving Objects
by: Yang, Huiming, et al.
Published: (2025) -
DexSim2Real: Foundation Model-Guided Sim-to-Real Transfer for Generalizable Dexterous Manipulation
by: Zeng, Zijian, et al.
Published: (2026) -
Reducing Credit Assignment Variance via Counterfactual Reasoning Paths
by: Ding, Fei, et al.
Published: (2026)