ReMem: Mutual Information-Aware Fine-tuning of Pretrained Vision Transformers for Effective Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Chengyu, Gui, Huan, Sachdeva, Noveen, Jin, Long, Yin, Ke, Shang, Jingbo, Hong, Lichan, Chi, Ed H., Zhao, Zhe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Magic Correlations: Understanding Knowledge Transfer from Pretraining to Supervised Fine-Tuning
by: Fan, Simin, et al.
Published: (2026)
by: Fan, Simin, et al.
Published: (2026)
ReMem-VLA: Empowering Vision-Language-Action Model with Memory via Dual-Level Recurrent Queries
by: Li, Hang, et al.
Published: (2026)
by: Li, Hang, et al.
Published: (2026)
LEVI: Generalizable Fine-tuning via Layer-wise Ensemble of Different Views
by: Roh, Yuji, et al.
Published: (2024)
by: Roh, Yuji, et al.
Published: (2024)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022)
by: Dong, Chengyu, et al.
Published: (2022)
How to Train Data-Efficient LLMs
by: Sachdeva, Noveen, et al.
Published: (2024)
by: Sachdeva, Noveen, et al.
Published: (2024)
Bridging the Gap: Unpacking the Hidden Challenges in Knowledge Distillation for Online Ranking Systems
by: Khani, Nikhil, et al.
Published: (2024)
by: Khani, Nikhil, et al.
Published: (2024)
Balancing Fine-tuning and RAG: A Hybrid Strategy for Dynamic LLM Recommendation Updates
by: Meng, Changping, et al.
Published: (2025)
by: Meng, Changping, et al.
Published: (2025)
Wisdom of Committee: Distilling from Foundation Model to Specialized Application Model
by: Liu, Zichang, et al.
Published: (2024)
by: Liu, Zichang, et al.
Published: (2024)
GC4NC: A Benchmark Framework for Graph Condensation on Node Classification with New Insights
by: Gong, Shengbo, et al.
Published: (2024)
by: Gong, Shengbo, et al.
Published: (2024)
KcMF: A Knowledge-compliant Framework for Schema and Entity Matching with Fine-tuning-free LLMs
by: Xu, Yongqin, et al.
Published: (2024)
by: Xu, Yongqin, et al.
Published: (2024)
Relational Knowledge Distillation Using Fine-tuned Function Vectors
by: Kang, Andrea, et al.
Published: (2026)
by: Kang, Andrea, et al.
Published: (2026)
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
by: Wang, Chengyu, et al.
Published: (2025)
by: Wang, Chengyu, et al.
Published: (2025)
Mutual Distillation Learning For Person Re-Identification
by: Fu, Huiyuan, et al.
Published: (2024)
by: Fu, Huiyuan, et al.
Published: (2024)
TuneShift-KD: Knowledge Distillation and Transfer for Fine-tuned Models
by: Guan, Yushi, et al.
Published: (2026)
by: Guan, Yushi, et al.
Published: (2026)
ActionPiece: Contextually Tokenizing Action Sequences for Generative Recommendation
by: Hou, Yupeng, et al.
Published: (2025)
by: Hou, Yupeng, et al.
Published: (2025)
Exploring the Effectiveness of Multi-stage Fine-tuning for Cross-encoder Re-rankers
by: Pezzuti, Francesca, et al.
Published: (2025)
by: Pezzuti, Francesca, et al.
Published: (2025)
BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
by: Wang, Shanmin, et al.
Published: (2025)
by: Wang, Shanmin, et al.
Published: (2025)
Can Muon Fine-tune Adam-Pretrained Models?
by: Qu, Xingyu, et al.
Published: (2026)
by: Qu, Xingyu, et al.
Published: (2026)
Aligning Large Language Models with Recommendation Knowledge
by: Cao, Yuwei, et al.
Published: (2024)
by: Cao, Yuwei, et al.
Published: (2024)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
by: Zhao, Lu, et al.
Published: (2025)
by: Zhao, Lu, et al.
Published: (2025)
When is the consistent prediction likely to be a correct prediction?
by: Nguyen, Alex, et al.
Published: (2024)
by: Nguyen, Alex, et al.
Published: (2024)
Evaluating the Smooth Control of Attribute Intensity in Text Generation with LLMs
by: Zhou, Shang, et al.
Published: (2024)
by: Zhou, Shang, et al.
Published: (2024)
CoMem: Context Management with A Decoupled Long-Context Model
by: Zhang, Yuwei, et al.
Published: (2026)
by: Zhang, Yuwei, et al.
Published: (2026)
IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuning
by: Wang, Qibin, et al.
Published: (2024)
by: Wang, Qibin, et al.
Published: (2024)
PACEvolve: Enabling Long-Horizon Progress-Aware Consistent Evolution
by: Yan, Minghao, et al.
Published: (2026)
by: Yan, Minghao, et al.
Published: (2026)
Confounder-Aware Medical Data Selection for Fine-Tuning Pretrained Vision Models
by: Ji, Anyang, et al.
Published: (2025)
by: Ji, Anyang, et al.
Published: (2025)
MemKD: Memory-Discrepancy Knowledge Distillation for Efficient Time Series Classification
by: Udayangani, Nilushika, et al.
Published: (2026)
by: Udayangani, Nilushika, et al.
Published: (2026)
InFiConD: Interactive No-code Fine-tuning with Concept-based Knowledge Distillation
by: Huang, Jinbin, et al.
Published: (2024)
by: Huang, Jinbin, et al.
Published: (2024)
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU?
by: Hościłowicz, Jakub, et al.
Published: (2023)
by: Hościłowicz, Jakub, et al.
Published: (2023)
Object-level Self-Distillation for Vision Pretraining
by: Hızlı, Çağlar, et al.
Published: (2025)
by: Hızlı, Çağlar, et al.
Published: (2025)
Going Beyond Feature Similarity: Effective Dataset Distillation based on Class-Aware Conditional Mutual Information
by: Zhong, Xinhao, et al.
Published: (2024)
by: Zhong, Xinhao, et al.
Published: (2024)
Linear Correlation in LM's Compositional Generalization and Hallucination
by: Peng, Letian, et al.
Published: (2025)
by: Peng, Letian, et al.
Published: (2025)
Beyond Fine-tuning: Unleashing the Potential of Continuous Pretraining for Clinical LLMs
by: Christophe, Clément, et al.
Published: (2024)
by: Christophe, Clément, et al.
Published: (2024)
Partial Knowledge Distillation for Alleviating the Inherent Inter-Class Discrepancy in Federated Learning
by: Gan, Xiaoyu, et al.
Published: (2024)
by: Gan, Xiaoyu, et al.
Published: (2024)
Fine-tuning a Multiple Instance Learning Feature Extractor with Masked Context Modelling and Knowledge Distillation
by: Pisula, Juan I., et al.
Published: (2024)
by: Pisula, Juan I., et al.
Published: (2024)
Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud
by: Yue, Yuanhao, et al.
Published: (2024)
by: Yue, Yuanhao, et al.
Published: (2024)
Fuse Before Transfer: Knowledge Fusion for Heterogeneous Distillation
by: Li, Guopeng, et al.
Published: (2024)
by: Li, Guopeng, et al.
Published: (2024)
Systematic Analysis for Pretrained Language Model Priming for Parameter-Efficient Fine-tuning
by: Huang, Shih-Cheng, et al.
Published: (2022)
by: Huang, Shih-Cheng, et al.
Published: (2022)
Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text Classification
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
Similar Items
-
The Magic Correlations: Understanding Knowledge Transfer from Pretraining to Supervised Fine-Tuning
by: Fan, Simin, et al.
Published: (2026) -
ReMem-VLA: Empowering Vision-Language-Action Model with Memory via Dual-Level Recurrent Queries
by: Li, Hang, et al.
Published: (2026) -
LEVI: Generalizable Fine-tuning via Layer-wise Ensemble of Different Views
by: Roh, Yuji, et al.
Published: (2024) -
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022) -
How to Train Data-Efficient LLMs
by: Sachdeva, Noveen, et al.
Published: (2024)