Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yanlong, Habibian, Amirhossein, Benini, Luca, Li, Yawei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CamSAM2: Segment Anything Accurately in Camouflaged Videos
by: Zhou, Yuli, et al.
Published: (2025)
by: Zhou, Yuli, et al.
Published: (2025)
Ultra-Efficient On-Device Object Detection on AI-Integrated Smart Glasses with TinyissimoYOLO
by: Moosmann, Julian, et al.
Published: (2023)
by: Moosmann, Julian, et al.
Published: (2023)
Mobile Video Diffusion
by: Yahia, Haitam Ben, et al.
Published: (2024)
by: Yahia, Haitam Ben, et al.
Published: (2024)
When SAM2 Meets Video Camouflaged Object Segmentation: A Comprehensive Evaluation and Adaptation
by: Zhou, Yuli, et al.
Published: (2024)
by: Zhou, Yuli, et al.
Published: (2024)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)
by: Li, Shuo, et al.
Published: (2024)
Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving
by: Lian, Weitong, et al.
Published: (2026)
by: Lian, Weitong, et al.
Published: (2026)
WaveFormer: A Lightweight Transformer Model for sEMG-based Gesture Recognition
by: Chen, Yanlong, et al.
Published: (2025)
by: Chen, Yanlong, et al.
Published: (2025)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
by: Khan, Anas Anwarul Haq, et al.
Published: (2025)
by: Khan, Anas Anwarul Haq, et al.
Published: (2025)
Forecasting When to Forecast: Accelerating Diffusion Models with Confidence-Gated Taylor
by: Guan, Xiaoliu, et al.
Published: (2025)
by: Guan, Xiaoliu, et al.
Published: (2025)
PLA4D: Pixel-Level Alignments for Text-to-4D Gaussian Splatting
by: Miao, Qiaowei, et al.
Published: (2024)
by: Miao, Qiaowei, et al.
Published: (2024)
Graph Relation Distillation for Efficient Biomedical Instance Segmentation
by: Liu, Xiaoyu, et al.
Published: (2024)
by: Liu, Xiaoyu, et al.
Published: (2024)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
by: Ghafoorian, Mohsen, et al.
Published: (2026)
by: Ghafoorian, Mohsen, et al.
Published: (2026)
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
by: Ghafoorian, Mohsen, et al.
Published: (2025)
by: Ghafoorian, Mohsen, et al.
Published: (2025)
Beyond Dataset Distillation: Lossless Dataset Concentration via Diffusion-Assisted Distribution Alignment
by: Liu, Tongfei, et al.
Published: (2026)
by: Liu, Tongfei, et al.
Published: (2026)
Enhancing Medical Large Vision-Language Models via Alignment Distillation
by: Chang, Aofei, et al.
Published: (2025)
by: Chang, Aofei, et al.
Published: (2025)
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs
by: Zhang, Song, et al.
Published: (2026)
by: Zhang, Song, et al.
Published: (2026)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
by: Feng, Qianhan, et al.
Published: (2024)
by: Feng, Qianhan, et al.
Published: (2024)
VACoT: Rethinking Visual Data Augmentation with VLMs
by: Xu, Zhengzhuo, et al.
Published: (2025)
by: Xu, Zhengzhuo, et al.
Published: (2025)
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval
by: Lei, Youbo, et al.
Published: (2023)
by: Lei, Youbo, et al.
Published: (2023)
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
by: Shabtay, Nimrod, et al.
Published: (2026)
by: Shabtay, Nimrod, et al.
Published: (2026)
Clockwork Diffusion: Efficient Generation With Model-Step Distillation
by: Habibian, Amirhossein, et al.
Published: (2023)
by: Habibian, Amirhossein, et al.
Published: (2023)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
by: Berman, Shmuel, et al.
Published: (2025)
by: Berman, Shmuel, et al.
Published: (2025)
Efficient Low-Resolution Face Recognition via Bridge Distillation
by: Ge, Shiming, et al.
Published: (2024)
by: Ge, Shiming, et al.
Published: (2024)
HoloEv-Net: Efficient Event-based Action Recognition via Holographic Spatial Embedding and Global Spectral Gating
by: Hao, Weidong
Published: (2026)
by: Hao, Weidong
Published: (2026)
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
by: Pan, Jiazhen, et al.
Published: (2025)
by: Pan, Jiazhen, et al.
Published: (2025)
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks
by: Li, Chenjun
Published: (2026)
by: Li, Chenjun
Published: (2026)
Enhancing Novel View Synthesis via Geometry Grounded Set Diffusion
by: Zanjani, Farhad G., et al.
Published: (2026)
by: Zanjani, Farhad G., et al.
Published: (2026)
Listener-Rewarded Thinking in VLMs for Image Preferences
by: Gambashidze, Alexander, et al.
Published: (2025)
by: Gambashidze, Alexander, et al.
Published: (2025)
Towards Lossless Ultimate Vision Token Compression for VLMs
by: Zheng, Dehua, et al.
Published: (2025)
by: Zheng, Dehua, et al.
Published: (2025)
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2025)
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2025)
Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization
by: Lan, Yuqin, et al.
Published: (2026)
by: Lan, Yuqin, et al.
Published: (2026)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
by: Kang, Inha, et al.
Published: (2025)
by: Kang, Inha, et al.
Published: (2025)
PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association
by: Hannan, Abdul, et al.
Published: (2025)
by: Hannan, Abdul, et al.
Published: (2025)
Treble Counterfactual VLMs: A Causal Approach to Hallucination
by: Li, Shawn, et al.
Published: (2025)
by: Li, Shawn, et al.
Published: (2025)
Caption This, Reason That: VLMs Caught in the Middle
by: Weng, Zihan, et al.
Published: (2025)
by: Weng, Zihan, et al.
Published: (2025)
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
by: Wang, Shengao, et al.
Published: (2025)
by: Wang, Shengao, et al.
Published: (2025)
Similar Items
-
CamSAM2: Segment Anything Accurately in Camouflaged Videos
by: Zhou, Yuli, et al.
Published: (2025) -
Ultra-Efficient On-Device Object Detection on AI-Integrated Smart Glasses with TinyissimoYOLO
by: Moosmann, Julian, et al.
Published: (2023) -
Mobile Video Diffusion
by: Yahia, Haitam Ben, et al.
Published: (2024) -
When SAM2 Meets Video Camouflaged Object Segmentation: A Comprehensive Evaluation and Adaptation
by: Zhou, Yuli, et al.
Published: (2024) -
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)