On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Baur, Simon, Benova, Alexandra, Cantú, Emilio Dolgener, Ma, Jackie |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SalNAS: Efficient Saliency-prediction Neural Architecture Search with self-knowledge distillation
par: Termritthikun, Chakkrit, et autres
Publié: (2024)
par: Termritthikun, Chakkrit, et autres
Publié: (2024)
What do vision-language models see in the context? Investigating multimodal in-context learning
par: Santos, Gabriel O. dos, et autres
Publié: (2025)
par: Santos, Gabriel O. dos, et autres
Publié: (2025)
Improved implicit diffusion model with knowledge distillation to estimate the spatial distribution density of carbon stock in remote sensing imagery
par: Yu, Zhenyu, et autres
Publié: (2024)
par: Yu, Zhenyu, et autres
Publié: (2024)
Efficient training for compact compression models via sequential distillation
par: Rodrigues, Caroline Mazini, et autres
Publié: (2026)
par: Rodrigues, Caroline Mazini, et autres
Publié: (2026)
Steering CLIP's vision transformer with sparse autoencoders
par: Joseph, Sonia, et autres
Publié: (2025)
par: Joseph, Sonia, et autres
Publié: (2025)
Learning few-step posterior samplers by unfolding and distillation of diffusion models
par: Mbakam, Charlesquin Kemajou, et autres
Publié: (2025)
par: Mbakam, Charlesquin Kemajou, et autres
Publié: (2025)
MULTIAQUA: A multimodal maritime dataset and robust training strategies for multimodal semantic segmentation
par: Muhovič, Jon, et autres
Publié: (2025)
par: Muhovič, Jon, et autres
Publié: (2025)
Linking heterogeneous microstructure informatics with expert characterization knowledge through customized and hybrid vision-language representations for industrial qualification
par: Safdar, Mutahar, et autres
Publié: (2025)
par: Safdar, Mutahar, et autres
Publié: (2025)
How to build a consistency model: Learning flow maps via self-distillation
par: Boffi, Nicholas M., et autres
Publié: (2025)
par: Boffi, Nicholas M., et autres
Publié: (2025)
Closing the gap in multimodal medical representation alignment
par: Grassucci, Eleonora, et autres
Publié: (2026)
par: Grassucci, Eleonora, et autres
Publié: (2026)
SPARE: Self-distillation for PARameter-Efficient Removal
par: Mola, Natnael, et autres
Publié: (2026)
par: Mola, Natnael, et autres
Publié: (2026)
Understanding vision transformer robustness through the lens of out-of-distribution detection
par: Kuang, Joey, et autres
Publié: (2026)
par: Kuang, Joey, et autres
Publié: (2026)
Fast and accurate neural reflectance transformation imaging through knowledge distillation
par: Dulecha, Tinsae G., et autres
Publié: (2025)
par: Dulecha, Tinsae G., et autres
Publié: (2025)
Millimeter-wave Imaging for Anthropometric Body Measurement
par: Senne, Miriam, et autres
Publié: (2026)
par: Senne, Miriam, et autres
Publié: (2026)
Knowledge distillation to effectively attain both region-of-interest and global semantics from an image where multiple objects appear
par: Jin, Seonwhee
Publié: (2024)
par: Jin, Seonwhee
Publié: (2024)
Towards reliable respiratory disease diagnosis based on cough sounds and vision transformers
par: Wang, Qian, et autres
Publié: (2024)
par: Wang, Qian, et autres
Publié: (2024)
An explainable vision transformer with transfer learning based efficient drought stress identification
par: Patra, Aswini Kumar, et autres
Publié: (2024)
par: Patra, Aswini Kumar, et autres
Publié: (2024)
Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
par: Chen, Kun, et autres
Publié: (2025)
par: Chen, Kun, et autres
Publié: (2025)
Classification of freshwater snails of the genus Radomaniola with multimodal triplet networks
par: Vetter, Dennis, et autres
Publié: (2024)
par: Vetter, Dennis, et autres
Publié: (2024)
Generalizable automated ischaemic stroke lesion segmentation with vision transformers
par: Foulon, Chris, et autres
Publié: (2025)
par: Foulon, Chris, et autres
Publié: (2025)
Buffer replay enhances the robustness of multimodal learning under missing-modality
par: Zhu, Hongye, et autres
Publié: (2025)
par: Zhu, Hongye, et autres
Publié: (2025)
Can multimodal representation learning by alignment preserve modality-specific information?
par: Thoreau, Romain, et autres
Publié: (2025)
par: Thoreau, Romain, et autres
Publié: (2025)
Visual concept ranking uncovers medical shortcuts used by large multimodal models
par: Janizek, Joseph D., et autres
Publié: (2026)
par: Janizek, Joseph D., et autres
Publié: (2026)
Towards a multimodal framework for remote sensing image change retrieval and captioning
par: Ferrod, Roger, et autres
Publié: (2024)
par: Ferrod, Roger, et autres
Publié: (2024)
General surgery vision transformer: A video pre-trained foundation model for general surgery
par: Schmidgall, Samuel, et autres
Publié: (2024)
par: Schmidgall, Samuel, et autres
Publié: (2024)
Deep Learning-based Multi Project InP Wafer Simulation for Unsupervised Surface Defect Detection
par: Cantú, Emílio Dolgener, et autres
Publié: (2025)
par: Cantú, Emílio Dolgener, et autres
Publié: (2025)
Adopting a human developmental visual diet yields robust, shape-based AI vision
par: Lu, Zejin, et autres
Publié: (2025)
par: Lu, Zejin, et autres
Publié: (2025)
A Laplace diffusion-based transformer model for heart rate forecasting within daily activity context
par: Mateescu, Andrei, et autres
Publié: (2025)
par: Mateescu, Andrei, et autres
Publié: (2025)
A multimodal slice discovery framework for systematic failure detection and explanation in medical image classification
par: Liu, Yixuan, et autres
Publié: (2026)
par: Liu, Yixuan, et autres
Publié: (2026)
AG-Fusion: adaptive gated multimodal fusion for 3d object detection in complex scenes
par: Liu, Sixian, et autres
Publié: (2025)
par: Liu, Sixian, et autres
Publié: (2025)
Hi-ALPS -- An Experimental Robustness Quantification of Six LiDAR-based Object Detection Systems for Autonomous Driving
par: Arzberger, Alexandra, et autres
Publié: (2025)
par: Arzberger, Alexandra, et autres
Publié: (2025)
Self-supervised cost of transport estimation for multimodal path planning
par: Gherold, Vincent, et autres
Publié: (2024)
par: Gherold, Vincent, et autres
Publié: (2024)
Copula-based mixture model identification for subgroup clustering with imaging applications
par: Zheng, Fei, et autres
Publié: (2025)
par: Zheng, Fei, et autres
Publié: (2025)
Enhancing compact convolutional transformers with super attention
par: Leandre, Simpenzwe Honore, et autres
Publié: (2025)
par: Leandre, Simpenzwe Honore, et autres
Publié: (2025)
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
par: Grzywaczewski, Jakub, et autres
Publié: (2026)
par: Grzywaczewski, Jakub, et autres
Publié: (2026)
Spiking monocular event based 6D pose estimation for space application
par: Courtois, Jonathan, et autres
Publié: (2025)
par: Courtois, Jonathan, et autres
Publié: (2025)
The effects of Hessian eigenvalue spectral density type on the applicability of Hessian analysis to generalization capability assessment of neural networks
par: Gabdullin, Nikita
Publié: (2025)
par: Gabdullin, Nikita
Publié: (2025)
Review of multimodal machine learning approaches in healthcare
par: Krones, Felix, et autres
Publié: (2024)
par: Krones, Felix, et autres
Publié: (2024)
Do multimodal models imagine electric sheep?
par: Ramakrishnan, Santhosh Kumar, et autres
Publié: (2026)
par: Ramakrishnan, Santhosh Kumar, et autres
Publié: (2026)
Contrastive vision-language learning with paraphrasing and negation
par: Ngan, Kwun Ho, et autres
Publié: (2025)
par: Ngan, Kwun Ho, et autres
Publié: (2025)
Documents similaires
-
SalNAS: Efficient Saliency-prediction Neural Architecture Search with self-knowledge distillation
par: Termritthikun, Chakkrit, et autres
Publié: (2024) -
What do vision-language models see in the context? Investigating multimodal in-context learning
par: Santos, Gabriel O. dos, et autres
Publié: (2025) -
Improved implicit diffusion model with knowledge distillation to estimate the spatial distribution density of carbon stock in remote sensing imagery
par: Yu, Zhenyu, et autres
Publié: (2024) -
Efficient training for compact compression models via sequential distillation
par: Rodrigues, Caroline Mazini, et autres
Publié: (2026) -
Steering CLIP's vision transformer with sparse autoencoders
par: Joseph, Sonia, et autres
Publié: (2025)