Agglomerating Large Vision Encoders via Distillation for VFSS Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zeng, Chengxi, Jiang, Yuxuan, Zhang, Fan, Gambaruto, Alberto, Burghardt, Tilo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tuning Vision Foundation Model via Test-Time Prompt-Guided Training for VFSS Segmentations
von: Zeng, Chengxi, et al.
Veröffentlicht: (2025)
von: Zeng, Chengxi, et al.
Veröffentlicht: (2025)
Video-SwinUNet: Spatio-temporal Deep Learning Framework for VFSS Instance Segmentation
von: Zeng, Chengxi, et al.
Veröffentlicht: (2023)
von: Zeng, Chengxi, et al.
Veröffentlicht: (2023)
EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
von: Zeng, Chengxi, et al.
Veröffentlicht: (2025)
von: Zeng, Chengxi, et al.
Veröffentlicht: (2025)
Feature Mapping in Physics-Informed Neural Networks (PINNs)
von: Zeng, Chengxi, et al.
Veröffentlicht: (2024)
von: Zeng, Chengxi, et al.
Veröffentlicht: (2024)
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models
von: Heinrich, Greg, et al.
Veröffentlicht: (2024)
von: Heinrich, Greg, et al.
Veröffentlicht: (2024)
ChimpVLM: Ethogram-Enhanced Chimpanzee Behaviour Recognition
von: Brookes, Otto, et al.
Veröffentlicht: (2024)
von: Brookes, Otto, et al.
Veröffentlicht: (2024)
Enhancing Medical Large Vision-Language Models via Alignment Distillation
von: Chang, Aofei, et al.
Veröffentlicht: (2025)
von: Chang, Aofei, et al.
Veröffentlicht: (2025)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
Enhancing Polyp Segmentation via Encoder Attention and Dynamic Kernel Update
von: Chashmi, Fatemeh Salahi, et al.
Veröffentlicht: (2025)
von: Chashmi, Fatemeh Salahi, et al.
Veröffentlicht: (2025)
Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation
von: Luo, Jie, et al.
Veröffentlicht: (2025)
von: Luo, Jie, et al.
Veröffentlicht: (2025)
Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation
von: Cho, Yubin, et al.
Veröffentlicht: (2024)
von: Cho, Yubin, et al.
Veröffentlicht: (2024)
X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning
von: Shao, Maanping, et al.
Veröffentlicht: (2026)
von: Shao, Maanping, et al.
Veröffentlicht: (2026)
Adversarial Prompt Distillation for Vision-Language Models
von: Luo, Lin, et al.
Veröffentlicht: (2024)
von: Luo, Lin, et al.
Veröffentlicht: (2024)
Large Vision-Language Models as Emotion Recognizers in Context Awareness
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute Alignment
von: Jiang, Yankai, et al.
Veröffentlicht: (2024)
von: Jiang, Yankai, et al.
Veröffentlicht: (2024)
KDMOS:Knowledge Distillation for Motion Segmentation
von: Cao, Chunyu, et al.
Veröffentlicht: (2025)
von: Cao, Chunyu, et al.
Veröffentlicht: (2025)
SDCD: Structure-Disrupted Contrastive Decoding for Mitigating Hallucinations in Large Vision-Language Models
von: Xia, Yuxuan, et al.
Veröffentlicht: (2026)
von: Xia, Yuxuan, et al.
Veröffentlicht: (2026)
Oil Spill Segmentation using Deep Encoder-Decoder models
von: Satyanarayana, Abhishek Ramanathapura, et al.
Veröffentlicht: (2023)
von: Satyanarayana, Abhishek Ramanathapura, et al.
Veröffentlicht: (2023)
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
von: Kabra, Rishabh, et al.
Veröffentlicht: (2026)
von: Kabra, Rishabh, et al.
Veröffentlicht: (2026)
MT-CYP-Net: Multi-Task Network for Pixel-Level Crop Yield Prediction Under Very Few Samples
von: Liu, Shenzhou, et al.
Veröffentlicht: (2025)
von: Liu, Shenzhou, et al.
Veröffentlicht: (2025)
VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
von: Jia, Ziheng, et al.
Veröffentlicht: (2025)
von: Jia, Ziheng, et al.
Veröffentlicht: (2025)
Graph Relation Distillation for Efficient Biomedical Instance Segmentation
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
Evaluating Large and Lightweight Vision Models for Irregular Component Segmentation in E-Waste Disassembly
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
Vision-Enhanced Time Series Forecasting via Latent Diffusion Models
von: Ruan, Weilin, et al.
Veröffentlicht: (2025)
von: Ruan, Weilin, et al.
Veröffentlicht: (2025)
UMSPU: Universal Multi-Size Phase Unwrapping via Mutual Self-Distillation and Adaptive Boosting Ensemble Segmenters
von: Du, Lintong, et al.
Veröffentlicht: (2024)
von: Du, Lintong, et al.
Veröffentlicht: (2024)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
von: Diao, Haiwen, et al.
Veröffentlicht: (2025)
von: Diao, Haiwen, et al.
Veröffentlicht: (2025)
FoPru: Focal Pruning for Efficient Large Vision-Language Models
von: Jiang, Lei, et al.
Veröffentlicht: (2024)
von: Jiang, Lei, et al.
Veröffentlicht: (2024)
Teacher Encoder-Student Decoder Denoising Guided Segmentation Network for Anomaly Detection
von: Song, Shixuan, et al.
Veröffentlicht: (2025)
von: Song, Shixuan, et al.
Veröffentlicht: (2025)
WeatherSeg: Weather-Robust Image Segmentation using Teacher-Student Dual Learning and Classifier-Updating Attention
von: Zhang, Zhang, et al.
Veröffentlicht: (2026)
von: Zhang, Zhang, et al.
Veröffentlicht: (2026)
Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
von: Shi, Jin, et al.
Veröffentlicht: (2026)
von: Shi, Jin, et al.
Veröffentlicht: (2026)
Jewelry Recognition via Encoder-Decoder Models
von: Alcalde-Llergo, José M., et al.
Veröffentlicht: (2024)
von: Alcalde-Llergo, José M., et al.
Veröffentlicht: (2024)
YO-CSA-T: A Real-time Badminton Tracking System Utilizing YOLO Based on Contextual and Spatial Attention
von: Lai, Yuan, et al.
Veröffentlicht: (2025)
von: Lai, Yuan, et al.
Veröffentlicht: (2025)
LKM-UNet: Large Kernel Vision Mamba UNet for Medical Image Segmentation
von: Wang, Jinhong, et al.
Veröffentlicht: (2024)
von: Wang, Jinhong, et al.
Veröffentlicht: (2024)
WildLive: Near Real-time Visual Wildlife Tracking onboard UAVs
von: Dat, Nguyen Ngoc, et al.
Veröffentlicht: (2025)
von: Dat, Nguyen Ngoc, et al.
Veröffentlicht: (2025)
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
von: Huang, Yiyang, et al.
Veröffentlicht: (2025)
von: Huang, Yiyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Tuning Vision Foundation Model via Test-Time Prompt-Guided Training for VFSS Segmentations
von: Zeng, Chengxi, et al.
Veröffentlicht: (2025) -
Video-SwinUNet: Spatio-temporal Deep Learning Framework for VFSS Instance Segmentation
von: Zeng, Chengxi, et al.
Veröffentlicht: (2023) -
EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
von: Zeng, Chengxi, et al.
Veröffentlicht: (2025) -
Feature Mapping in Physics-Informed Neural Networks (PINNs)
von: Zeng, Chengxi, et al.
Veröffentlicht: (2024) -
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models
von: Heinrich, Greg, et al.
Veröffentlicht: (2024)