Visual-Advantage On-Policy Distillation for Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Ruiqi, Lv, Xiaolei, Li, Gengsheng, Zhu, Ximo, Wang, Zhiheng, Zhang, Zhengbo, Chen, Junkai, Li, Zhiheng, Li, Bo, Gao, Jun, Wu, Shu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning
by: Chen, Junkai, et al.
Published: (2026)
by: Chen, Junkai, et al.
Published: (2026)
Unsupervised Domain Adaption Harnessing Vision-Language Pre-training
by: Zhou, Wenlve, et al.
Published: (2024)
by: Zhou, Wenlve, et al.
Published: (2024)
4DRaL: Bridging 4D Radar with LiDAR for Place Recognition using Knowledge Distillation
by: Huang, Ningyuan, et al.
Published: (2026)
by: Huang, Ningyuan, et al.
Published: (2026)
Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation
by: Li, Zhiheng, et al.
Published: (2026)
by: Li, Zhiheng, et al.
Published: (2026)
Sketch-to-Architecture: Generative AI-aided Architectural Design
by: Li, Pengzhi, et al.
Published: (2024)
by: Li, Pengzhi, et al.
Published: (2024)
How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings
by: Li, Zhiheng, et al.
Published: (2026)
by: Li, Zhiheng, et al.
Published: (2026)
Visual Bridge: Universal Visual Perception Representations Generating
by: Gao, Yilin, et al.
Published: (2025)
by: Gao, Yilin, et al.
Published: (2025)
LayerDiffusion: Layered Controlled Image Editing with Diffusion Models
by: Li, Pengzhi, et al.
Published: (2023)
by: Li, Pengzhi, et al.
Published: (2023)
FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training
by: Cao, Anjia, et al.
Published: (2024)
by: Cao, Anjia, et al.
Published: (2024)
DAIT: Distillation from Vision-Language Models to Lightweight Classifiers with Adaptive Intermediate Teacher Transfer
by: He, Zhengxu, et al.
Published: (2026)
by: He, Zhengxu, et al.
Published: (2026)
LOMA: Language-assisted Semantic Occupancy Network via Triplane Mamba
by: Cui, Yubo, et al.
Published: (2024)
by: Cui, Yubo, et al.
Published: (2024)
Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
by: Chen, Xuezheng, et al.
Published: (2025)
by: Chen, Xuezheng, et al.
Published: (2025)
Curriculum Dataset Distillation
by: Ma, Zhiheng, et al.
Published: (2024)
by: Ma, Zhiheng, et al.
Published: (2024)
Trajectory-Diversity-Driven Robust Vision-and-Language Navigation
by: Li, Jiangyang, et al.
Published: (2026)
by: Li, Jiangyang, et al.
Published: (2026)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
by: Dai, Tingjun, et al.
Published: (2026)
by: Dai, Tingjun, et al.
Published: (2026)
HERO: Rethinking Visual Token Early Dropping in High-Resolution Large Vision-Language Models
by: Li, Xu, et al.
Published: (2025)
by: Li, Xu, et al.
Published: (2025)
FlowTrack: Point-level Flow Network for 3D Single Object Tracking
by: Li, Shuo, et al.
Published: (2024)
by: Li, Shuo, et al.
Published: (2024)
ART3D: 3D Gaussian Splatting for Text-Guided Artistic Scenes Generation
by: Li, Pengzhi, et al.
Published: (2024)
by: Li, Pengzhi, et al.
Published: (2024)
Masked Diffusion Vision-Language Models for Temporal Action Localization
by: Wang, Fengshun, et al.
Published: (2026)
by: Wang, Fengshun, et al.
Published: (2026)
Pano360: Perspective to Panoramic Vision with Geometric Consistency
by: Zhu, Zhengdong, et al.
Published: (2026)
by: Zhu, Zhengdong, et al.
Published: (2026)
ReGLA: Efficient Receptive-Field Modeling with Gated Linear Attention Network
by: Li, Junzhou, et al.
Published: (2026)
by: Li, Junzhou, et al.
Published: (2026)
One-Shot Multilingual Font Generation Via ViT
by: Wang, Zhiheng, et al.
Published: (2024)
by: Wang, Zhiheng, et al.
Published: (2024)
Nested-TNT: Hierarchical Vision Transformers with Multi-Scale Feature Processing
by: Liu, Yuang, et al.
Published: (2024)
by: Liu, Yuang, et al.
Published: (2024)
DARK: Denoising, Amplification, Restoration Kit
by: Li, Zhuoheng, et al.
Published: (2024)
by: Li, Zhuoheng, et al.
Published: (2024)
SeqTrack3D: Exploring Sequence Information for Robust 3D Point Cloud Tracking
by: Lin, Yu, et al.
Published: (2024)
by: Lin, Yu, et al.
Published: (2024)
4D-CS: Exploiting Cluster Prior for 4D Spatio-Temporal LiDAR Semantic Segmentation
by: Zhong, Jiexi, et al.
Published: (2025)
by: Zhong, Jiexi, et al.
Published: (2025)
Beyond Weight Adaptation: Feature-Space Domain Injection for Cross-Modal Ship Re-Identification
by: Xian, Tingfeng, et al.
Published: (2025)
by: Xian, Tingfeng, et al.
Published: (2025)
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
by: Han, Yuhang, et al.
Published: (2026)
by: Han, Yuhang, et al.
Published: (2026)
The Devil is in the Edges: Monocular Depth Estimation with Edge-aware Consistency Fusion
by: Li, Pengzhi, et al.
Published: (2024)
by: Li, Pengzhi, et al.
Published: (2024)
PromptKD: Unsupervised Prompt Distillation for Vision-Language Models
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
by: Bai, Detao, et al.
Published: (2025)
by: Bai, Detao, et al.
Published: (2025)
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models
by: Li, Xu, et al.
Published: (2024)
by: Li, Xu, et al.
Published: (2024)
Dataset Distillation via Vision-Language Category Prototype
by: Zou, Yawen, et al.
Published: (2025)
by: Zou, Yawen, et al.
Published: (2025)
VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
by: Sargent, Kyle, et al.
Published: (2025)
by: Sargent, Kyle, et al.
Published: (2025)
StreamMOS: Streaming Moving Object Segmentation with Multi-View Perception and Dual-Span Memory
by: Li, Zhiheng, et al.
Published: (2024)
by: Li, Zhiheng, et al.
Published: (2024)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models
by: Sun, Haoyi, et al.
Published: (2026)
by: Sun, Haoyi, et al.
Published: (2026)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
by: Yoon, Hee Suk, et al.
Published: (2026)
by: Yoon, Hee Suk, et al.
Published: (2026)
Towards Long-window Anchoring in Vision-Language Model Distillation
by: Zhou, Haoyi, et al.
Published: (2025)
by: Zhou, Haoyi, et al.
Published: (2025)
HybridMIM: A Hybrid Masked Image Modeling Framework for 3D Medical Image Segmentation
by: Xing, Zhaohu, et al.
Published: (2023)
by: Xing, Zhaohu, et al.
Published: (2023)
Similar Items
-
Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning
by: Chen, Junkai, et al.
Published: (2026) -
Unsupervised Domain Adaption Harnessing Vision-Language Pre-training
by: Zhou, Wenlve, et al.
Published: (2024) -
4DRaL: Bridging 4D Radar with LiDAR for Place Recognition using Knowledge Distillation
by: Huang, Ningyuan, et al.
Published: (2026) -
Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation
by: Li, Zhiheng, et al.
Published: (2026) -
Sketch-to-Architecture: Generative AI-aided Architectural Design
by: Li, Pengzhi, et al.
Published: (2024)