Concept-skill Transferability-based Data Selection for Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Jaewoo, Li, Boyang, Hwang, Sung Ju |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
von: Lee, Jaewoo, et al.
Veröffentlicht: (2023)
von: Lee, Jaewoo, et al.
Veröffentlicht: (2023)
Visualizing the loss landscape of Self-supervised Vision Transformer
von: Lee, Youngwan, et al.
Veröffentlicht: (2024)
von: Lee, Youngwan, et al.
Veröffentlicht: (2024)
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
von: Kang, Seongjae, et al.
Veröffentlicht: (2025)
von: Kang, Seongjae, et al.
Veröffentlicht: (2025)
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
von: Hwang, Sunil, et al.
Veröffentlicht: (2022)
von: Hwang, Sunil, et al.
Veröffentlicht: (2022)
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2024)
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2024)
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2024)
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2024)
Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models
von: Bravo-Sánchez, Laura, et al.
Veröffentlicht: (2024)
von: Bravo-Sánchez, Laura, et al.
Veröffentlicht: (2024)
Understanding Task Transfer in Vision-Language Models
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025)
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025)
CASL: Concept-Aligned Sparse Latents for Interpreting Diffusion Models
von: He, Zhenghao, et al.
Veröffentlicht: (2026)
von: He, Zhenghao, et al.
Veröffentlicht: (2026)
Transferable Adversarial Attacks on Black-Box Vision-Language Models
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
von: Srivastava, Divyansh, et al.
Veröffentlicht: (2024)
von: Srivastava, Divyansh, et al.
Veröffentlicht: (2024)
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding
von: Le, Anjie, et al.
Veröffentlicht: (2025)
von: Le, Anjie, et al.
Veröffentlicht: (2025)
D-TPT: Dimensional Entropy Maximization for Calibrating Test-Time Prompt Tuning in Vision-Language Models
von: Han, Jisu, et al.
Veröffentlicht: (2025)
von: Han, Jisu, et al.
Veröffentlicht: (2025)
Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
von: Wang, Yuchen, et al.
Veröffentlicht: (2025)
von: Wang, Yuchen, et al.
Veröffentlicht: (2025)
Cross-Modal Adapter: Parameter-Efficient Transfer Learning Approach for Vision-Language Models
von: Yang, Juncheng, et al.
Veröffentlicht: (2024)
von: Yang, Juncheng, et al.
Veröffentlicht: (2024)
Vision-Language Models Encode Clinical Guidelines for Concept-Based Medical Reasoning
von: Harmanani, Mohamed, et al.
Veröffentlicht: (2026)
von: Harmanani, Mohamed, et al.
Veröffentlicht: (2026)
Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
von: Naharas, Nilay, et al.
Veröffentlicht: (2025)
von: Naharas, Nilay, et al.
Veröffentlicht: (2025)
On the Difficulty of Learning a Meta-network for Training Data Selection
von: Du, Zilin, et al.
Veröffentlicht: (2026)
von: Du, Zilin, et al.
Veröffentlicht: (2026)
An Adaptive Method Stabilizing Activations for Enhanced Generalization
von: Seung, Hyunseok, et al.
Veröffentlicht: (2025)
von: Seung, Hyunseok, et al.
Veröffentlicht: (2025)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
von: Du, Zilin, et al.
Veröffentlicht: (2024)
von: Du, Zilin, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Image Recognition in Vision-Language Models through Human-like Concept Guidance
von: Liu, Hui, et al.
Veröffentlicht: (2025)
von: Liu, Hui, et al.
Veröffentlicht: (2025)
Anomaly Score: Evaluating Generative Models and Individual Generated Images based on Complexity and Vulnerability
von: Hwang, Jaehui, et al.
Veröffentlicht: (2023)
von: Hwang, Jaehui, et al.
Veröffentlicht: (2023)
ICONS: Influence Consensus for Vision-Language Data Selection
von: Wu, Xindi, et al.
Veröffentlicht: (2024)
von: Wu, Xindi, et al.
Veröffentlicht: (2024)
Bridge the Modality and Capability Gaps in Vision-Language Model Selection
von: Yi, Chao, et al.
Veröffentlicht: (2024)
von: Yi, Chao, et al.
Veröffentlicht: (2024)
When are Lemons Purple? The Concept Association Bias of Vision-Language Models
von: Yamada, Yutaro, et al.
Veröffentlicht: (2022)
von: Yamada, Yutaro, et al.
Veröffentlicht: (2022)
Teaching Metric Distance to Discrete Autoregressive Language Models
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
CFM: Language-aligned Concept Foundation Model for Vision
von: Wittenmayer, Kai, et al.
Veröffentlicht: (2026)
von: Wittenmayer, Kai, et al.
Veröffentlicht: (2026)
Improved Alignment of Modalities in Large Vision Language Models
von: Jangra, Kartik, et al.
Veröffentlicht: (2025)
von: Jangra, Kartik, et al.
Veröffentlicht: (2025)
Detecting and Preventing Hallucinations in Large Vision Language Models
von: Gunjal, Anisha, et al.
Veröffentlicht: (2023)
von: Gunjal, Anisha, et al.
Veröffentlicht: (2023)
Self-Evolving Visual Concept Library using Vision-Language Critics
von: Sehgal, Atharva, et al.
Veröffentlicht: (2025)
von: Sehgal, Atharva, et al.
Veröffentlicht: (2025)
Towards Understanding How Knowledge Evolves in Large Vision-Language Models
von: Wang, Sudong, et al.
Veröffentlicht: (2025)
von: Wang, Sudong, et al.
Veröffentlicht: (2025)
TroL: Traversal of Layers for Large Language and Vision Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2024)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
Transferring Textual Preferences to Vision-Language Understanding through Model Merging
von: Li, Chen-An, et al.
Veröffentlicht: (2025)
von: Li, Chen-An, et al.
Veröffentlicht: (2025)
Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain
von: Lee, Seulbi, et al.
Veröffentlicht: (2026)
von: Lee, Seulbi, et al.
Veröffentlicht: (2026)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
von: Park, Jihwan, et al.
Veröffentlicht: (2025)
von: Park, Jihwan, et al.
Veröffentlicht: (2025)
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing
von: Kim, Jaihoon, et al.
Veröffentlicht: (2025)
von: Kim, Jaihoon, et al.
Veröffentlicht: (2025)
Self-Refining Video Sampling
von: Jang, Sangwon, et al.
Veröffentlicht: (2026)
von: Jang, Sangwon, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
von: Lee, Jaewoo, et al.
Veröffentlicht: (2023) -
Visualizing the loss landscape of Self-supervised Vision Transformer
von: Lee, Youngwan, et al.
Veröffentlicht: (2024) -
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
von: Kang, Seongjae, et al.
Veröffentlicht: (2025) -
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
von: Hwang, Sunil, et al.
Veröffentlicht: (2022) -
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
von: Lee, Daeun, et al.
Veröffentlicht: (2024)