Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Byung-Kwan, Wang, Yu-Chiang Frank, Hachiuma, Ryo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unified Reinforcement and Imitation Learning for Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025)
by: Lee, Byung-Kwan, et al.
Published: (2025)
VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
by: Lee, Byung-Kwan, et al.
Published: (2024)
by: Lee, Byung-Kwan, et al.
Published: (2024)
Learning from Synthetic Data via Provenance-Based Input Gradient Guidance
by: Nagano, Koshiro, et al.
Published: (2026)
by: Nagano, Koshiro, et al.
Published: (2026)
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
by: Lian, Chenyu, et al.
Published: (2025)
by: Lian, Chenyu, et al.
Published: (2025)
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models
by: Zhou, Andy, et al.
Published: (2023)
by: Zhou, Andy, et al.
Published: (2023)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
by: Yu, Seonghoon, et al.
Published: (2026)
by: Yu, Seonghoon, et al.
Published: (2026)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
Causal Unsupervised Semantic Segmentation
by: Kim, Junho, et al.
Published: (2023)
by: Kim, Junho, et al.
Published: (2023)
The Role of Teacher Calibration in Knowledge Distillation
by: Kim, Suyoung, et al.
Published: (2025)
by: Kim, Suyoung, et al.
Published: (2025)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
by: Li, YuQian, et al.
Published: (2025)
by: Li, YuQian, et al.
Published: (2025)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
Towards Long-window Anchoring in Vision-Language Model Distillation
by: Zhou, Haoyi, et al.
Published: (2025)
by: Zhou, Haoyi, et al.
Published: (2025)
PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting
by: Maqbool, Danyal, et al.
Published: (2025)
by: Maqbool, Danyal, et al.
Published: (2025)
Di$\mathtt{[M]}$O: Distilling Masked Diffusion Models into One-step Generator
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
Look Through Masks: Towards Masked Face Recognition with De-Occlusion Distillation
by: Li, Chenyu, et al.
Published: (2024)
by: Li, Chenyu, et al.
Published: (2024)
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
by: Kang, Seongjae, et al.
Published: (2025)
by: Kang, Seongjae, et al.
Published: (2025)
VIOLA: Towards Video In-Context Learning with Minimal Annotations
by: Fujii, Ryo, et al.
Published: (2026)
by: Fujii, Ryo, et al.
Published: (2026)
Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
MMM: Generative Masked Motion Model
by: Pinyoanuntapong, Ekkasit, et al.
Published: (2023)
by: Pinyoanuntapong, Ekkasit, et al.
Published: (2023)
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
by: Zhao, Yunhan, et al.
Published: (2024)
by: Zhao, Yunhan, et al.
Published: (2024)
Vision-Language Models Provide Promptable Representations for Reinforcement Learning
by: Chen, William, et al.
Published: (2024)
by: Chen, William, et al.
Published: (2024)
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2026)
by: Huang, Chi-Pin, et al.
Published: (2026)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
by: Yang, Senqiao, et al.
Published: (2025)
by: Yang, Senqiao, et al.
Published: (2025)
TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors
by: Cheng, Wei-Yuan, et al.
Published: (2026)
by: Cheng, Wei-Yuan, et al.
Published: (2026)
DINORANKCLIP: DINOv3 Distillation and Injection for Vision-Language Pretraining with High-Order Ranking Consistency
by: Jiang, Shuyang, et al.
Published: (2026)
by: Jiang, Shuyang, et al.
Published: (2026)
LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
by: Hirota, Yusuke, et al.
Published: (2025)
by: Hirota, Yusuke, et al.
Published: (2025)
DARK: Diagonal-Anchored Repulsive Knowledge Distillation for Vision-Language Models under Extreme Compression
by: Saeed, Numan, et al.
Published: (2026)
by: Saeed, Numan, et al.
Published: (2026)
MINR: Implicit Neural Representations with Masked Image Modelling
by: Lee, Sua, et al.
Published: (2025)
by: Lee, Sua, et al.
Published: (2025)
Weakly Semi-supervised Tool Detection in Minimally Invasive Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
by: Shin, Hyunjune, et al.
Published: (2024)
by: Shin, Hyunjune, et al.
Published: (2024)
Toward an Artificial General Teacher: Procedural Geometry Data Generation and Visual Grounding with Vision-Language Models
by: Nguyen-Truong, Hai, et al.
Published: (2026)
by: Nguyen-Truong, Hai, et al.
Published: (2026)
Good Teachers Explain: Explanation-Enhanced Knowledge Distillation
by: Parchami-Araghi, Amin, et al.
Published: (2024)
by: Parchami-Araghi, Amin, et al.
Published: (2024)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
OISD: On-Policy Internal Self-Distillation of Language Models
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
MuM: Multi-View Masked Image Modeling for 3D Vision
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
by: Wu, Yen-Siang, et al.
Published: (2025)
by: Wu, Yen-Siang, et al.
Published: (2025)
Diffusion Reinforcement Learning via Centered Reward Distillation
by: Zhu, Yuanzhi, et al.
Published: (2026)
by: Zhu, Yuanzhi, et al.
Published: (2026)
Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
by: Kim, Gahyeon, et al.
Published: (2025)
by: Kim, Gahyeon, et al.
Published: (2025)
AAPL: Adding Attributes to Prompt Learning for Vision-Language Models
by: Kim, Gahyeon, et al.
Published: (2024)
by: Kim, Gahyeon, et al.
Published: (2024)
Similar Items
-
Unified Reinforcement and Imitation Learning for Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025) -
VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
by: Lee, Byung-Kwan, et al.
Published: (2024) -
Learning from Synthetic Data via Provenance-Based Input Gradient Guidance
by: Nagano, Koshiro, et al.
Published: (2026) -
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
by: Lian, Chenyu, et al.
Published: (2025) -
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024)