RL makes MLLMs see better than SFT
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Junha, Yun, Sangdoo, Han, Dongyoon, Choo, Jaegul, Heo, Byeongho |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026)
by: Song, Junha, et al.
Published: (2026)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
Masking meets Supervision: A Strong Learning Alliance
by: Heo, Byeongho, et al.
Published: (2023)
by: Heo, Byeongho, et al.
Published: (2023)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023)
by: Kim, Jiwon, et al.
Published: (2023)
DNNs May Determine Major Properties of Their Outputs Early, with Timing Possibly Driven by Bias
by: Park, Song, et al.
Published: (2025)
by: Park, Song, et al.
Published: (2025)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
MagiCapture: High-Resolution Multi-Concept Portrait Customization
by: Hyung, Junha, et al.
Published: (2023)
by: Hyung, Junha, et al.
Published: (2023)
Token Bottleneck: One Token to Remember Dynamics
by: Kim, Taekyung, et al.
Published: (2025)
by: Kim, Taekyung, et al.
Published: (2025)
Model Stock: All we need is just a few fine-tuned models
by: Jang, Dong-Hwan, et al.
Published: (2024)
by: Jang, Dong-Hwan, et al.
Published: (2024)
DaWin: Training-free Dynamic Weight Interpolation for Robust Adaptation
by: Oh, Changdae, et al.
Published: (2024)
by: Oh, Changdae, et al.
Published: (2024)
Similarity of Neural Architectures using Adversarial Attack Transferability
by: Hwang, Jaehui, et al.
Published: (2022)
by: Hwang, Jaehui, et al.
Published: (2022)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
by: Hwang, Dongyoon, et al.
Published: (2024)
by: Hwang, Dongyoon, et al.
Published: (2024)
Morphing Tokens Draw Strong Masked Image Models
by: Kim, Taekyung, et al.
Published: (2023)
by: Kim, Taekyung, et al.
Published: (2023)
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
by: Kim, Donghu, et al.
Published: (2024)
by: Kim, Donghu, et al.
Published: (2024)
Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
by: Kim, Kinam, et al.
Published: (2025)
by: Kim, Kinam, et al.
Published: (2025)
Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text Guidance
by: Nam, Giung, et al.
Published: (2024)
by: Nam, Giung, et al.
Published: (2024)
Learning with Unmasked Tokens Drives Stronger Vision Learners
by: Kim, Taekyung, et al.
Published: (2023)
by: Kim, Taekyung, et al.
Published: (2023)
SeiT++: Masked Token Modeling Improves Storage-efficient Training
by: Lee, Minhyun, et al.
Published: (2023)
by: Lee, Minhyun, et al.
Published: (2023)
Scratching Visual Transformer's Back with Uniform Attention
by: Hyeon-Woo, Nam, et al.
Published: (2022)
by: Hyeon-Woo, Nam, et al.
Published: (2022)
LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
by: Chun, Sanghyuk, et al.
Published: (2025)
by: Chun, Sanghyuk, et al.
Published: (2025)
Sparse autoencoders reveal selective remapping of visual concepts during adaptation
by: Lim, Hyesu, et al.
Published: (2024)
by: Lim, Hyesu, et al.
Published: (2024)
Is user feedback always informative? Retrieval Latent Defending for Semi-Supervised Domain Adaptation without Source Data
by: Song, Junha, et al.
Published: (2024)
by: Song, Junha, et al.
Published: (2024)
Probabilistic Language-Image Pre-Training
by: Chun, Sanghyuk, et al.
Published: (2024)
by: Chun, Sanghyuk, et al.
Published: (2024)
Towards Calibrated Robust Fine-Tuning of Vision-Language Models
by: Oh, Changdae, et al.
Published: (2023)
by: Oh, Changdae, et al.
Published: (2023)
MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation
by: Lee, Minhyun, et al.
Published: (2024)
by: Lee, Minhyun, et al.
Published: (2024)
Exploring Conditions for Diffusion models in Robotic Control
by: Shin, Heeseong, et al.
Published: (2025)
by: Shin, Heeseong, et al.
Published: (2025)
Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure
by: Yun, Jooyeol, et al.
Published: (2025)
by: Yun, Jooyeol, et al.
Published: (2025)
Scaling Up Personalized Image Aesthetic Assessment via Task Vector Customization
by: Yun, Jooyeol, et al.
Published: (2024)
by: Yun, Jooyeol, et al.
Published: (2024)
SelfSwapper: Self-Supervised Face Swapping via Shape Agnostic Masked AutoEncoder
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
by: Lu, Aojun, et al.
Published: (2026)
by: Lu, Aojun, et al.
Published: (2026)
VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
by: Park, Kyoungjun, et al.
Published: (2025)
by: Park, Kyoungjun, et al.
Published: (2025)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
by: Qiu, Haibo, et al.
Published: (2025)
by: Qiu, Haibo, et al.
Published: (2025)
When Test-Time Adaptation Meets Self-Supervised Models
by: Han, Jisu, et al.
Published: (2025)
by: Han, Jisu, et al.
Published: (2025)
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts
by: Kim, Wonjae, et al.
Published: (2024)
by: Kim, Wonjae, et al.
Published: (2024)
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
by: Hyung, Junha, et al.
Published: (2024)
by: Hyung, Junha, et al.
Published: (2024)
PromptRL: Prompt Matters in RL for Flow-Based Image Generation
by: Wang, Fu-Yun, et al.
Published: (2026)
by: Wang, Fu-Yun, et al.
Published: (2026)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)
by: Chu, Tianzhe, et al.
Published: (2025)
Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
by: Jo, Kyungmin, et al.
Published: (2025)
by: Jo, Kyungmin, et al.
Published: (2025)
EgoX: Egocentric Video Generation from a Single Exocentric Video
by: Kang, Taewoong, et al.
Published: (2025)
by: Kang, Taewoong, et al.
Published: (2025)
Similar Items
-
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026) -
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024) -
Masking meets Supervision: A Strong Learning Alliance
by: Heo, Byeongho, et al.
Published: (2023) -
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023) -
DNNs May Determine Major Properties of Their Outputs Early, with Timing Possibly Driven by Bias
by: Park, Song, et al.
Published: (2025)