Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Jiabo, Chen, Chen, Lyu, Lingjuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
by: Wang, Yimu, et al.
Published: (2025)
by: Wang, Yimu, et al.
Published: (2025)
Empirical Recipes for Efficient and Compact Vision-Language Models
by: Huang, Jiabo, et al.
Published: (2026)
by: Huang, Jiabo, et al.
Published: (2026)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
See Further for Parameter Efficient Fine-tuning by Standing on the Shoulders of Decomposition
by: Si, Chongjie, et al.
Published: (2024)
by: Si, Chongjie, et al.
Published: (2024)
Closer to Reality: Practical Semi-Supervised Federated Learning for Foundation Model Adaptation
by: Sun, Guangyu, et al.
Published: (2025)
by: Sun, Guangyu, et al.
Published: (2025)
Detecting, Explaining, and Mitigating Memorization in Diffusion Models
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
by: Liao, Zijun, et al.
Published: (2025)
by: Liao, Zijun, et al.
Published: (2025)
See Further When Clear: Curriculum Consistency Model
by: Liu, Yunpeng, et al.
Published: (2024)
by: Liu, Yunpeng, et al.
Published: (2024)
COALA: A Practical and Vision-Centric Federated Learning Platform
by: Zhuang, Weiming, et al.
Published: (2024)
by: Zhuang, Weiming, et al.
Published: (2024)
Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection
by: Lin, Kaiqing, et al.
Published: (2024)
by: Lin, Kaiqing, et al.
Published: (2024)
Evaluating and Mitigating IP Infringement in Visual Generative AI
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
Task-Specific Knowledge Distillation from the Vision Foundation Model for Enhanced Medical Image Segmentation
by: Liang, Pengchen, et al.
Published: (2025)
by: Liang, Pengchen, et al.
Published: (2025)
A Simple Background Augmentation Method for Object Detection with Diffusion Model
by: Li, Yuhang, et al.
Published: (2024)
by: Li, Yuhang, et al.
Published: (2024)
Towards Fundamentally Scalable Model Selection: Asymptotically Fast Update and Selection
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
Training-Free Layout-to-Image Generation with Marginal Attention Constraints
by: Chen, Huancheng, et al.
Published: (2024)
by: Chen, Huancheng, et al.
Published: (2024)
CoCAViT: Compact Vision Transformer with Robust Global Coordination
by: Wang, Xuyang, et al.
Published: (2025)
by: Wang, Xuyang, et al.
Published: (2025)
Replay-Free Continual Low-Rank Adaptation with Dynamic Memory
by: Chen, Huancheng, et al.
Published: (2024)
by: Chen, Huancheng, et al.
Published: (2024)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
by: Yan, Yuping, et al.
Published: (2025)
by: Yan, Yuping, et al.
Published: (2025)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
Vision-Language Models Can't See the Obvious
by: Dahou, Yasser, et al.
Published: (2025)
by: Dahou, Yasser, et al.
Published: (2025)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
by: Chen, Siyi, et al.
Published: (2026)
by: Chen, Siyi, et al.
Published: (2026)
NocPlace: Nocturnal Visual Place Recognition via Generative and Inherited Knowledge Transfer
by: Liu, Bingxi, et al.
Published: (2024)
by: Liu, Bingxi, et al.
Published: (2024)
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
by: Guo, Jianyuan, et al.
Published: (2024)
by: Guo, Jianyuan, et al.
Published: (2024)
Fully Exploiting Vision Foundation Model's Profound Prior Knowledge for Generalizable RGB-Depth Driving Scene Parsing
by: Guo, Sicen, et al.
Published: (2025)
by: Guo, Sicen, et al.
Published: (2025)
Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models
by: Cai, Zhejia, et al.
Published: (2025)
by: Cai, Zhejia, et al.
Published: (2025)
DINOReg: Strong Point Cloud Registration with Vision Foundation Model
by: Chen, Congjia, et al.
Published: (2025)
by: Chen, Congjia, et al.
Published: (2025)
Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation
by: Lv, Chonghua, et al.
Published: (2026)
by: Lv, Chonghua, et al.
Published: (2026)
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
by: Zhang, Songyan, et al.
Published: (2024)
by: Zhang, Songyan, et al.
Published: (2024)
Rendering-Refined Stable Diffusion for Privacy Compliant Synthetic Data
by: Patwari, Kartik, et al.
Published: (2024)
by: Patwari, Kartik, et al.
Published: (2024)
A Unified Low-level Foundation Model for Enhancing Pathology Image Quality
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
by: Wang, Zhenting, et al.
Published: (2023)
by: Wang, Zhenting, et al.
Published: (2023)
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
by: Ortu, Francesco, et al.
Published: (2025)
by: Ortu, Francesco, et al.
Published: (2025)
Implicit Modeling for Transferability Estimation of Vision Foundation Models
by: Zheng, Yaoyan, et al.
Published: (2025)
by: Zheng, Yaoyan, et al.
Published: (2025)
All-in-One: Transferring Vision Foundation Models into Stereo Matching
by: Zhou, Jingyi, et al.
Published: (2024)
by: Zhou, Jingyi, et al.
Published: (2024)
Adapting Vision Foundation Models for Real-time Ultrasound Image Segmentation
by: Zhang, Xiaoran, et al.
Published: (2025)
by: Zhang, Xiaoran, et al.
Published: (2025)
A Breast Vision Pathology Foundation Model for Real-world Clinical Utility
by: Xu, Yingxue, et al.
Published: (2026)
by: Xu, Yingxue, et al.
Published: (2026)
Efficient Transfer Learning for Video-language Foundation Models
by: Chen, Haoxing, et al.
Published: (2024)
by: Chen, Haoxing, et al.
Published: (2024)
Bootstrapping SparseFormers from Vision Foundation Models
by: Gao, Ziteng, et al.
Published: (2023)
by: Gao, Ziteng, et al.
Published: (2023)
Seeing the Unseen: Towards Zero-Shot Inspection for Wind Turbine Blades using Knowledge-Augmented Vision Language Models
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AI
by: Cheng, Siyuan, et al.
Published: (2025)
by: Cheng, Siyuan, et al.
Published: (2025)
Similar Items
-
UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
by: Wang, Yimu, et al.
Published: (2025) -
Empirical Recipes for Efficient and Compact Vision-Language Models
by: Huang, Jiabo, et al.
Published: (2026) -
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026) -
See Further for Parameter Efficient Fine-tuning by Standing on the Shoulders of Decomposition
by: Si, Chongjie, et al.
Published: (2024) -
Closer to Reality: Practical Semi-Supervised Federated Learning for Foundation Model Adaptation
by: Sun, Guangyu, et al.
Published: (2025)