Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Jihwan, Song, Taehoon, Lee, Sanghyeok, Choi, Miso, Kim, Hyunwoo J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
by: Park, Jihwan, et al.
Published: (2026)
by: Park, Jihwan, et al.
Published: (2026)
EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
by: Yang, Chanhyeong, et al.
Published: (2025)
by: Yang, Chanhyeong, et al.
Published: (2025)
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
by: Choi, Dasol, et al.
Published: (2026)
by: Choi, Dasol, et al.
Published: (2026)
MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2026)
by: Ko, Dohwan, et al.
Published: (2026)
Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
by: Park, Dogyun, et al.
Published: (2025)
by: Park, Dogyun, et al.
Published: (2025)
Constant Acceleration Flow
by: Park, Dogyun, et al.
Published: (2024)
by: Park, Dogyun, et al.
Published: (2024)
ReCo: Reminder Composition Mitigates Hallucinations in Vision-Language Models
by: Chytas, Sotirios Panagiotis, et al.
Published: (2025)
by: Chytas, Sotirios Panagiotis, et al.
Published: (2025)
Super-class guided Transformer for Zero-Shot Attribute Classification
by: Kim, Sehyung, et al.
Published: (2025)
by: Kim, Sehyung, et al.
Published: (2025)
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
by: Choi, Dasol, et al.
Published: (2025)
by: Choi, Dasol, et al.
Published: (2025)
Exploiting Style Latent Flows for Generalizing Deepfake Video Detection
by: Choi, Jongwook, et al.
Published: (2024)
by: Choi, Jongwook, et al.
Published: (2024)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
by: Kwon, Mincheol, et al.
Published: (2026)
by: Kwon, Mincheol, et al.
Published: (2026)
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
by: Kim, Jisoo, et al.
Published: (2026)
by: Kim, Jisoo, et al.
Published: (2026)
FIFO-Diffusion: Generating Infinite Videos from Text without Training
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Robust Multimodal 3D Object Detection via Modality-Agnostic Decoding and Proximity-based Modality Ensemble
by: Cha, Juhan, et al.
Published: (2024)
by: Cha, Juhan, et al.
Published: (2024)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2025)
by: Lee, Youngwan, et al.
Published: (2025)
Future-Proof Yourself: An AI Era Survival Guide
by: Kim, Taehoon
Published: (2025)
by: Kim, Taehoon
Published: (2025)
Towards Difficulty-Agnostic Efficient Transfer Learning for Vision-Language Models
by: Yang, Yongjin, et al.
Published: (2023)
by: Yang, Yongjin, et al.
Published: (2023)
Collaborative Edge-to-Server Inference for Vision-Language Models
by: Song, Soochang, et al.
Published: (2025)
by: Song, Soochang, et al.
Published: (2025)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
by: Kang, Beomseok, et al.
Published: (2026)
by: Kang, Beomseok, et al.
Published: (2026)
Evolving Prompt Adaptation for Vision-Language Models
by: Zhang, Enming, et al.
Published: (2026)
by: Zhang, Enming, et al.
Published: (2026)
AMRG: Extend Vision Language Models for Automatic Mammography Report Generation
by: Sung, Nak-Jun, et al.
Published: (2025)
by: Sung, Nak-Jun, et al.
Published: (2025)
Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
by: Back, Kyungryul, et al.
Published: (2025)
by: Back, Kyungryul, et al.
Published: (2025)
COTTA: Context-Aware Transfer Adaptation for Trajectory Prediction in Autonomous Driving
by: Park, Seohyoung, et al.
Published: (2026)
by: Park, Seohyoung, et al.
Published: (2026)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
by: Kang, Donggoo, et al.
Published: (2024)
by: Kang, Donggoo, et al.
Published: (2024)
Revealing Multi-View Hallucination in Large Vision-Language Models
by: Park, Wooje, et al.
Published: (2026)
by: Park, Wooje, et al.
Published: (2026)
Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers
by: Heo, Seongsoo, et al.
Published: (2025)
by: Heo, Seongsoo, et al.
Published: (2025)
Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
by: Choi, Dasol, et al.
Published: (2025)
by: Choi, Dasol, et al.
Published: (2025)
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
by: Kim, Kibum, et al.
Published: (2023)
by: Kim, Kibum, et al.
Published: (2023)
Compositional Image Synthesis with Inference-Time Scaling
by: Ji, Minsuk, et al.
Published: (2025)
by: Ji, Minsuk, et al.
Published: (2025)
Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation
by: Kim, Taehoon, et al.
Published: (2025)
by: Kim, Taehoon, et al.
Published: (2025)
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
by: Park, Yohan, et al.
Published: (2025)
by: Park, Yohan, et al.
Published: (2025)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
by: Kim, Sohee, et al.
Published: (2025)
by: Kim, Sohee, et al.
Published: (2025)
Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
by: Jeon, Yerim, et al.
Published: (2025)
by: Jeon, Yerim, et al.
Published: (2025)
Activating Self-Attention for Multi-Scene Absolute Pose Regression
by: Lee, Miso, et al.
Published: (2024)
by: Lee, Miso, et al.
Published: (2024)
Long-term Pre-training for Temporal Action Detection with Transformers
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
AH-OCDA: Amplitude-based Curriculum Learning and Hopfield Segmentation Model for Open Compound Domain Adaptation
by: Choi, Jaehyun, et al.
Published: (2024)
by: Choi, Jaehyun, et al.
Published: (2024)
Similar Items
-
RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
by: Park, Jihwan, et al.
Published: (2026) -
EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality
by: Lee, Sanghyeok, et al.
Published: (2024) -
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024) -
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
by: Yang, Chanhyeong, et al.
Published: (2025) -
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
by: Choi, Dasol, et al.
Published: (2026)