When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jeddi, Ahmadreza, Shaban, Kimia, Baghbanzadeh, Negin, Sharan, Natasha, Moturu, Abhishek, Dolatabadi, Elham, Taati, Babak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Similarity-Aware Token Pruning: Your VLM but Faster
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025)
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025)
Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning
von: Baghbanzadeh, Negin, et al.
Veröffentlicht: (2025)
von: Baghbanzadeh, Negin, et al.
Veröffentlicht: (2025)
LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2026)
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2026)
Pain in 3D: Generating Controllable Synthetic Faces for Automated Pain Assessment
von: Lin, Xin Lei, et al.
Veröffentlicht: (2025)
von: Lin, Xin Lei, et al.
Veröffentlicht: (2025)
LiBaGS: Lightweight Boundary Gap Synthesis for Targeted Synthetic Data Selection
von: Moturu, Abhishek, et al.
Veröffentlicht: (2026)
von: Moturu, Abhishek, et al.
Veröffentlicht: (2026)
SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers
von: Rajabi, Javad, et al.
Veröffentlicht: (2026)
von: Rajabi, Javad, et al.
Veröffentlicht: (2026)
LiLAW: Lightweight Learnable Adaptive Weighting to Learn Sample Difficulty & Improve Noisy Training
von: Moturu, Abhishek, et al.
Veröffentlicht: (2025)
von: Moturu, Abhishek, et al.
Veröffentlicht: (2025)
PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025)
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025)
SynPAIN: A Synthetic Dataset of Pain and Non-Pain Facial Expressions
von: Taati, Babak, et al.
Veröffentlicht: (2025)
von: Taati, Babak, et al.
Veröffentlicht: (2025)
Advancing Medical Representation Learning Through High-Quality Data
von: Baghbanzadeh, Negin, et al.
Veröffentlicht: (2025)
von: Baghbanzadeh, Negin, et al.
Veröffentlicht: (2025)
Automated Capability Evaluation of Foundation Models
von: Afkanpour, Arash, et al.
Veröffentlicht: (2025)
von: Afkanpour, Arash, et al.
Veröffentlicht: (2025)
RL makes MLLMs see better than SFT
von: Song, Junha, et al.
Veröffentlicht: (2025)
von: Song, Junha, et al.
Veröffentlicht: (2025)
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
von: Wang, Sudong, et al.
Veröffentlicht: (2026)
von: Wang, Sudong, et al.
Veröffentlicht: (2026)
GEAR: Genetic AutoResearch for Agentic Code Evolution
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2026)
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2026)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
von: Lu, Aojun, et al.
Veröffentlicht: (2026)
von: Lu, Aojun, et al.
Veröffentlicht: (2026)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
von: Wang, Junke, et al.
Veröffentlicht: (2025)
von: Wang, Junke, et al.
Veröffentlicht: (2025)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
von: Qiu, Haibo, et al.
Veröffentlicht: (2025)
von: Qiu, Haibo, et al.
Veröffentlicht: (2025)
FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding
von: Mehraban, Soroush, et al.
Veröffentlicht: (2025)
von: Mehraban, Soroush, et al.
Veröffentlicht: (2025)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
von: Chu, Tianzhe, et al.
Veröffentlicht: (2025)
von: Chu, Tianzhe, et al.
Veröffentlicht: (2025)
PyVision-RL: Forging Open Agentic Vision Models via RL
von: Zhao, Shitian, et al.
Veröffentlicht: (2026)
von: Zhao, Shitian, et al.
Veröffentlicht: (2026)
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
von: Chen, Jierun, et al.
Veröffentlicht: (2025)
von: Chen, Jierun, et al.
Veröffentlicht: (2025)
Benchmarking Vision-Language Contrastive Methods for Medical Representation Learning
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
von: Xia, Canming, et al.
Veröffentlicht: (2026)
von: Xia, Canming, et al.
Veröffentlicht: (2026)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models
von: Islam, Mohammed Saidul, et al.
Veröffentlicht: (2026)
von: Islam, Mohammed Saidul, et al.
Veröffentlicht: (2026)
SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
von: Shu, Fangxun, et al.
Veröffentlicht: (2025)
von: Shu, Fangxun, et al.
Veröffentlicht: (2025)
Scaling Vision-and-Language Navigation With Offline RL
von: Bundele, Valay, et al.
Veröffentlicht: (2024)
von: Bundele, Valay, et al.
Veröffentlicht: (2024)
Consistency-Guided Asynchronous Contrastive Tuning for Few-Shot Class-Incremental Tuning of Foundation Models
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
LIFT: Latent Implicit Functions for Task- and Data-Agnostic Encoding
von: Kazerouni, Amirhossein, et al.
Veröffentlicht: (2025)
von: Kazerouni, Amirhossein, et al.
Veröffentlicht: (2025)
Leveraging Clinical Text and Class Conditioning for 3D Prostate MRI Generation
von: Grabke, Emerson P., et al.
Veröffentlicht: (2025)
von: Grabke, Emerson P., et al.
Veröffentlicht: (2025)
Mitigating 3D Prostate Biparametric MRI Data Scarcity through Domain Adaptation using Locally-Trained Latent Diffusion Models for Prostate Cancer Detection
von: Grabke, Emerson P., et al.
Veröffentlicht: (2025)
von: Grabke, Emerson P., et al.
Veröffentlicht: (2025)
STARS: Self-supervised Tuning for 3D Action Recognition in Skeleton Sequences
von: Mehraban, Soroush, et al.
Veröffentlicht: (2024)
von: Mehraban, Soroush, et al.
Veröffentlicht: (2024)
GAITGen: Disentangled Motion-Pathology Impaired Gait Generative Model -- Bringing Motion Generation to the Clinical Domain
von: Adeli, Vida, et al.
Veröffentlicht: (2025)
von: Adeli, Vida, et al.
Veröffentlicht: (2025)
Exploring Bias and Prediction Metrics to Characterise the Fairness of Machine Learning for Equity-Centered Public Health Decision-Making: A Narrative Review
von: Raza, Shaina, et al.
Veröffentlicht: (2024)
von: Raza, Shaina, et al.
Veröffentlicht: (2024)
When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing
von: Sun, Libo, et al.
Veröffentlicht: (2026)
von: Sun, Libo, et al.
Veröffentlicht: (2026)
RL Fine-Tuning Heals OOD Forgetting in SFT
von: Jin, Hangzhan, et al.
Veröffentlicht: (2025)
von: Jin, Hangzhan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Similarity-Aware Token Pruning: Your VLM but Faster
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025) -
Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning
von: Baghbanzadeh, Negin, et al.
Veröffentlicht: (2025) -
LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2026) -
Pain in 3D: Generating Controllable Synthetic Faces for Automated Pain Assessment
von: Lin, Xin Lei, et al.
Veröffentlicht: (2025) -
LiBaGS: Lightweight Boundary Gap Synthesis for Targeted Synthetic Data Selection
von: Moturu, Abhishek, et al.
Veröffentlicht: (2026)