AMD: Automatic Multi-step Distillation of Large-scale Vision Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Cheng, Wang, Qifan, Dianat, Sohail A., Rabbani, Majid, Rao, Raghuveer M., Fang, Yi, Guan, Qiang, Huang, Lifu, Liu, Dongfang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Image Translation as Diffusion Visual Programmers
von: Han, Cheng, et al.
Veröffentlicht: (2024)
von: Han, Cheng, et al.
Veröffentlicht: (2024)
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
von: Wang, Jiamian, et al.
Veröffentlicht: (2024)
von: Wang, Jiamian, et al.
Veröffentlicht: (2024)
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
Effective Dual-Region Augmentation for Reduced Reliance on Large Amounts of Labeled Data
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
Re-Imagining Multimodal Instruction Tuning: A Representation View
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)
Shuffle PatchMix Augmentation with Confidence-Margin Weighted Pseudo-Labels for Enhanced Source-Free Domain Adaptation
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
Visual Self-Refinement for Autoregressive Models
von: Wang, Jiamian, et al.
Veröffentlicht: (2025)
von: Wang, Jiamian, et al.
Veröffentlicht: (2025)
Prototypical Transformer as Unified Motion Learners
von: Han, Cheng, et al.
Veröffentlicht: (2024)
von: Han, Cheng, et al.
Veröffentlicht: (2024)
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
Latent Chain-of-Thought for Visual Reasoning
von: Sun, Guohao, et al.
Veröffentlicht: (2025)
von: Sun, Guohao, et al.
Veröffentlicht: (2025)
Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?
von: Han, Cheng, et al.
Veröffentlicht: (2024)
von: Han, Cheng, et al.
Veröffentlicht: (2024)
Visual Fourier Prompt Tuning
von: Zeng, Runjia, et al.
Veröffentlicht: (2024)
von: Zeng, Runjia, et al.
Veröffentlicht: (2024)
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
von: Qi, Jingyuan, et al.
Veröffentlicht: (2025)
von: Qi, Jingyuan, et al.
Veröffentlicht: (2025)
Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
Self-supervised Adversarial Training of Monocular Depth Estimation against Physical-World Attacks
von: Cheng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Cheng, Zhiyuan, et al.
Veröffentlicht: (2024)
Radiance Field Learners As UAV First-Person Viewers
von: Yan, Liqi, et al.
Veröffentlicht: (2024)
von: Yan, Liqi, et al.
Veröffentlicht: (2024)
Multimodal Instruction Tuning with Conditional Mixture of LoRA
von: Shen, Ying, et al.
Veröffentlicht: (2024)
von: Shen, Ying, et al.
Veröffentlicht: (2024)
TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token Ditching
von: Zeng, Runjia, et al.
Veröffentlicht: (2026)
von: Zeng, Runjia, et al.
Veröffentlicht: (2026)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
A-SelecT: Automatic Timestep Selection for Diffusion Transformer Representation Learning
von: Liu, Changyu, et al.
Veröffentlicht: (2026)
von: Liu, Changyu, et al.
Veröffentlicht: (2026)
Prompt-based Adaptation in Large-scale Vision Models: A Survey
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
Inference Compute-Optimal Video Vision Language Models
von: Wang, Peiqi, et al.
Veröffentlicht: (2025)
von: Wang, Peiqi, et al.
Veröffentlicht: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
AMD: Adaptive Momentum and Decoupled Contrastive Learning Framework for Robust Long-Tail Trajectory Prediction
von: Rao, Bin, et al.
Veröffentlicht: (2025)
von: Rao, Bin, et al.
Veröffentlicht: (2025)
FDCT: Frequency-Aware Decomposition and Cross-Modal Token-Alignment for Multi-Sensor Target Classification
von: Sami, Shoaib Meraj, et al.
Veröffentlicht: (2025)
von: Sami, Shoaib Meraj, et al.
Veröffentlicht: (2025)
TAMISeg: Text-Aligned Multi-scale Medical Image Segmentation with Semantic Encoder Distillation
von: Gao, Qiang, et al.
Veröffentlicht: (2026)
von: Gao, Qiang, et al.
Veröffentlicht: (2026)
Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models
von: Sun, Haoyi, et al.
Veröffentlicht: (2026)
von: Sun, Haoyi, et al.
Veröffentlicht: (2026)
Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning
von: Zhang, Qifan, et al.
Veröffentlicht: (2024)
von: Zhang, Qifan, et al.
Veröffentlicht: (2024)
SSGA-Net: Stepwise Spatial Global-local Aggregation Networks for for Autonomous Driving
von: Cui, Yiming, et al.
Veröffentlicht: (2024)
von: Cui, Yiming, et al.
Veröffentlicht: (2024)
AMD-Mamba: A Phenotype-Aware Multi-Modal Framework for Robust AMD Prognosis
von: Wu, Puzhen, et al.
Veröffentlicht: (2025)
von: Wu, Puzhen, et al.
Veröffentlicht: (2025)
Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
von: Zhou, Zanwei, et al.
Veröffentlicht: (2025)
von: Zhou, Zanwei, et al.
Veröffentlicht: (2025)
TokenMotion: Motion-Guided Vision Transformer for Video Camouflaged Object Detection Via Learnable Token Selection
von: Yu, Zifan, et al.
Veröffentlicht: (2023)
von: Yu, Zifan, et al.
Veröffentlicht: (2023)
Error-driven Data-efficient Large Multimodal Model Tuning
von: Yao, Barry Menglong, et al.
Veröffentlicht: (2024)
von: Yao, Barry Menglong, et al.
Veröffentlicht: (2024)
MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models
von: Rahman, Amirul, et al.
Veröffentlicht: (2025)
von: Rahman, Amirul, et al.
Veröffentlicht: (2025)
Are Large-scale Soft Labels Necessary for Large-scale Dataset Distillation?
von: Xiao, Lingao, et al.
Veröffentlicht: (2024)
von: Xiao, Lingao, et al.
Veröffentlicht: (2024)
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
von: Wang, Haibo, et al.
Veröffentlicht: (2026)
von: Wang, Haibo, et al.
Veröffentlicht: (2026)
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
von: Tao, Haoyi, et al.
Veröffentlicht: (2026)
von: Tao, Haoyi, et al.
Veröffentlicht: (2026)
Multi-student Diffusion Distillation for Better One-step Generators
von: Song, Yanke, et al.
Veröffentlicht: (2024)
von: Song, Yanke, et al.
Veröffentlicht: (2024)
CollabOD: Collaborative Multi-Backbone with Cross-scale Vision for UAV Small Object Detection
von: Bai, Xuecheng, et al.
Veröffentlicht: (2026)
von: Bai, Xuecheng, et al.
Veröffentlicht: (2026)
ProMotion: Prototypes As Motion Learners
von: Lu, Yawen, et al.
Veröffentlicht: (2024)
von: Lu, Yawen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Image Translation as Diffusion Visual Programmers
von: Han, Cheng, et al.
Veröffentlicht: (2024) -
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
von: Wang, Jiamian, et al.
Veröffentlicht: (2024) -
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025) -
Effective Dual-Region Augmentation for Reduced Reliance on Large Amounts of Labeled Data
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025) -
Re-Imagining Multimodal Instruction Tuning: A Representation View
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)