On Good Practices for Task-Specific Distillation of Large Pretrained Visual Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Marrie, Juliette, Arbel, Michael, Mairal, Julien, Larlus, Diane |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LUDVIG: Learning-Free Uplifting of 2D Visual Features to Gaussian Splatting Scenes
von: Marrie, Juliette, et al.
Veröffentlicht: (2024)
von: Marrie, Juliette, et al.
Veröffentlicht: (2024)
MicroFlow: Domain-Specific Optical Flow for Ground Deformation Estimation in Seismic Events
von: Bertrand, Juliette, et al.
Veröffentlicht: (2025)
von: Bertrand, Juliette, et al.
Veröffentlicht: (2025)
UNIC: Universal Classification Models via Multi-teacher Distillation
von: Sariyildiz, Mert Bulent, et al.
Veröffentlicht: (2024)
von: Sariyildiz, Mert Bulent, et al.
Veröffentlicht: (2024)
Unsupervised Imaging Inverse Problems with Diffusion Distribution Matching
von: Meanti, Giacomo, et al.
Veröffentlicht: (2025)
von: Meanti, Giacomo, et al.
Veröffentlicht: (2025)
Beyond MMSE: Enhancing PnP Restoration with ProxiMAP
von: Vert, Kenta, et al.
Veröffentlicht: (2026)
von: Vert, Kenta, et al.
Veröffentlicht: (2026)
Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos
von: Tschernezki, Vadim, et al.
Veröffentlicht: (2025)
von: Tschernezki, Vadim, et al.
Veröffentlicht: (2025)
What could go wrong? Discovering and describing failure modes in computer vision
von: Csurka, Gabriela, et al.
Veröffentlicht: (2024)
von: Csurka, Gabriela, et al.
Veröffentlicht: (2024)
Visual Instruction Pretraining for Domain-Specific Foundation Models
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
von: Jiang, Zeren, et al.
Veröffentlicht: (2026)
von: Jiang, Zeren, et al.
Veröffentlicht: (2026)
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
von: Böhle, Moritz, et al.
Veröffentlicht: (2025)
von: Böhle, Moritz, et al.
Veröffentlicht: (2025)
DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers
von: Sariyildiz, Mert Bulent, et al.
Veröffentlicht: (2025)
von: Sariyildiz, Mert Bulent, et al.
Veröffentlicht: (2025)
PanSt3R: Multi-view Consistent Panoptic Segmentation
von: Zust, Lojze, et al.
Veröffentlicht: (2025)
von: Zust, Lojze, et al.
Veröffentlicht: (2025)
Task Alignment: A simple and effective proxy for model merging in computer vision
von: de Jorge, Pau, et al.
Veröffentlicht: (2026)
von: de Jorge, Pau, et al.
Veröffentlicht: (2026)
Weatherproofing Retrieval for Localization with Generative AI and Geometric Consistency
von: Kalantidis, Yannis, et al.
Veröffentlicht: (2024)
von: Kalantidis, Yannis, et al.
Veröffentlicht: (2024)
Task Specific Pretraining with Noisy Labels for Remote Sensing Image Segmentation
von: Liu, Chenying, et al.
Veröffentlicht: (2024)
von: Liu, Chenying, et al.
Veröffentlicht: (2024)
Cross-Modal Knowledge Distillation from Spatial Transcriptomics to Histology
von: Hizmi, Arbel, et al.
Veröffentlicht: (2026)
von: Hizmi, Arbel, et al.
Veröffentlicht: (2026)
Vision Transformers Need Registers
von: Darcet, Timothée, et al.
Veröffentlicht: (2023)
von: Darcet, Timothée, et al.
Veröffentlicht: (2023)
SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining
von: Braham, Nassim Ait Ali, et al.
Veröffentlicht: (2026)
von: Braham, Nassim Ait Ali, et al.
Veröffentlicht: (2026)
Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
von: Jiang, Zeren, et al.
Veröffentlicht: (2025)
von: Jiang, Zeren, et al.
Veröffentlicht: (2025)
SLAD : Shared LoRA Adapters for Task Specific Distillation
von: Bensaid, Reda, et al.
Veröffentlicht: (2026)
von: Bensaid, Reda, et al.
Veröffentlicht: (2026)
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
Task-Specific Knowledge Distillation from the Vision Foundation Model for Enhanced Medical Image Segmentation
von: Liang, Pengchen, et al.
Veröffentlicht: (2025)
von: Liang, Pengchen, et al.
Veröffentlicht: (2025)
EPIC Fields: Marrying 3D Geometry and Video Understanding
von: Tschernezki, Vadim, et al.
Veröffentlicht: (2023)
von: Tschernezki, Vadim, et al.
Veröffentlicht: (2023)
Cluster and Predict Latent Patches for Improved Masked Image Modeling
von: Darcet, Timothée, et al.
Veröffentlicht: (2025)
von: Darcet, Timothée, et al.
Veröffentlicht: (2025)
Pretrained Visual Uncertainties
von: Kirchhof, Michael, et al.
Veröffentlicht: (2024)
von: Kirchhof, Michael, et al.
Veröffentlicht: (2024)
Are Pretrained Image Matchers Good Enough for SAR-Optical Satellite Registration?
von: Corley, Isaac, et al.
Veröffentlicht: (2026)
von: Corley, Isaac, et al.
Veröffentlicht: (2026)
CogVLM: Visual Expert for Pretrained Language Models
von: Wang, Weihan, et al.
Veröffentlicht: (2023)
von: Wang, Weihan, et al.
Veröffentlicht: (2023)
Optimal transport unlocks end-to-end learning for single-molecule localization
von: Seailles, Romain, et al.
Veröffentlicht: (2025)
von: Seailles, Romain, et al.
Veröffentlicht: (2025)
What Makes a Good Dataset for Knowledge Distillation?
von: Frank, Logan, et al.
Veröffentlicht: (2024)
von: Frank, Logan, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning of Large Pretrained Models for Instance Segmentation Tasks
von: Baker, Nermeen Abou, et al.
Veröffentlicht: (2026)
von: Baker, Nermeen Abou, et al.
Veröffentlicht: (2026)
PANDAS: Prototype-based Novel Class Discovery and Detection
von: Hayes, Tyler L., et al.
Veröffentlicht: (2024)
von: Hayes, Tyler L., et al.
Veröffentlicht: (2024)
From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction Tuning
von: Bai, Yang, et al.
Veröffentlicht: (2024)
von: Bai, Yang, et al.
Veröffentlicht: (2024)
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks
von: Lee, Jusung, et al.
Veröffentlicht: (2024)
von: Lee, Jusung, et al.
Veröffentlicht: (2024)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Task-Specific Adaptation with Restricted Model Access
von: Levy, Matan, et al.
Veröffentlicht: (2025)
von: Levy, Matan, et al.
Veröffentlicht: (2025)
Do Satellite Tasks Need Special Pretraining?
von: Vanyan, Ani, et al.
Veröffentlicht: (2025)
von: Vanyan, Ani, et al.
Veröffentlicht: (2025)
LocCa: Visual Pretraining with Location-aware Captioners
von: Wan, Bo, et al.
Veröffentlicht: (2024)
von: Wan, Bo, et al.
Veröffentlicht: (2024)
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining
von: Wang, Di, et al.
Veröffentlicht: (2024)
von: Wang, Di, et al.
Veröffentlicht: (2024)
SSL-AD: Spatiotemporal Self-Supervised Learning for Generalizability and Adaptability Across Alzheimer's Prediction Tasks and Datasets
von: Kaczmarek, Emily, et al.
Veröffentlicht: (2025)
von: Kaczmarek, Emily, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LUDVIG: Learning-Free Uplifting of 2D Visual Features to Gaussian Splatting Scenes
von: Marrie, Juliette, et al.
Veröffentlicht: (2024) -
MicroFlow: Domain-Specific Optical Flow for Ground Deformation Estimation in Seismic Events
von: Bertrand, Juliette, et al.
Veröffentlicht: (2025) -
UNIC: Universal Classification Models via Multi-teacher Distillation
von: Sariyildiz, Mert Bulent, et al.
Veröffentlicht: (2024) -
Unsupervised Imaging Inverse Problems with Diffusion Distribution Matching
von: Meanti, Giacomo, et al.
Veröffentlicht: (2025) -
Beyond MMSE: Enhancing PnP Restoration with ProxiMAP
von: Vert, Kenta, et al.
Veröffentlicht: (2026)