AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Yuhan, Ji, Yuyang, Zhao, Zhiyu, Wu, Gangshan, Wang, Limin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Test-Time Prompt Tuning for Vision-Language Models
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
Dual DETRs for Multi-Label Temporal Action Detection
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
Asymmetric Masked Distillation for Pre-Training Small Foundation Models
von: Zhao, Zhiyu, et al.
Veröffentlicht: (2023)
von: Zhao, Zhiyu, et al.
Veröffentlicht: (2023)
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
MixFormerV2: Efficient Fully Transformer Tracking
von: Cui, Yutao, et al.
Veröffentlicht: (2023)
von: Cui, Yutao, et al.
Veröffentlicht: (2023)
GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
von: Luo, Xingyu, et al.
Veröffentlicht: (2026)
von: Luo, Xingyu, et al.
Veröffentlicht: (2026)
STMixer: A One-Stage Sparse Action Detector
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
Open-Vocabulary Spatio-Temporal Action Detection
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
ZeroI2V: Zero-Cost Adaptation of Pre-trained Transformers from Image to Video
von: Li, Xinhao, et al.
Veröffentlicht: (2023)
von: Li, Xinhao, et al.
Veröffentlicht: (2023)
Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction
von: Li, Yuanbo, et al.
Veröffentlicht: (2026)
von: Li, Yuanbo, et al.
Veröffentlicht: (2026)
From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching
von: Ji, Yuyang, et al.
Veröffentlicht: (2026)
von: Ji, Yuyang, et al.
Veröffentlicht: (2026)
Attention! Your Vision Language Model Could Be Maliciously Manipulated
von: Wang, Xiaosen, et al.
Veröffentlicht: (2025)
von: Wang, Xiaosen, et al.
Veröffentlicht: (2025)
VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking
von: Xu, Boyue, et al.
Veröffentlicht: (2026)
von: Xu, Boyue, et al.
Veröffentlicht: (2026)
RS-OOD: A Vision-Language Augmented Framework for Out-of-Distribution Detection in Remote Sensing
von: Wang, Chenhao, et al.
Veröffentlicht: (2025)
von: Wang, Chenhao, et al.
Veröffentlicht: (2025)
RGB-D Video Object Segmentation via Enhanced Multi-store Feature Memory
von: Xu, Boyue, et al.
Veröffentlicht: (2025)
von: Xu, Boyue, et al.
Veröffentlicht: (2025)
Motion-Aware Generative Frame Interpolation
von: Zhang, Guozhen, et al.
Veröffentlicht: (2025)
von: Zhang, Guozhen, et al.
Veröffentlicht: (2025)
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
von: Zhan, Yufei, et al.
Veröffentlicht: (2024)
von: Zhan, Yufei, et al.
Veröffentlicht: (2024)
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
von: Yang, Jiange, et al.
Veröffentlicht: (2025)
von: Yang, Jiange, et al.
Veröffentlicht: (2025)
MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object Tracking
von: Gao, Ruopeng, et al.
Veröffentlicht: (2023)
von: Gao, Ruopeng, et al.
Veröffentlicht: (2023)
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
von: Liu, Chunxu, et al.
Veröffentlicht: (2025)
von: Liu, Chunxu, et al.
Veröffentlicht: (2025)
Enhancing Continual Learning of Vision-Language Models via Dynamic Prefix Weighting
von: Jang, Hyeonseo, et al.
Veröffentlicht: (2026)
von: Jang, Hyeonseo, et al.
Veröffentlicht: (2026)
Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization
von: Zhao, Minyi, et al.
Veröffentlicht: (2024)
von: Zhao, Minyi, et al.
Veröffentlicht: (2024)
ArchMap: Arch-Flattening and Knowledge-Guided Vision Language Model for Tooth Counting and Structured Dental Understanding
von: Zhang, Bohan, et al.
Veröffentlicht: (2025)
von: Zhang, Bohan, et al.
Veröffentlicht: (2025)
ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision Experts
von: Wu, Yuanchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuanchen, et al.
Veröffentlicht: (2025)
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
von: Chi, Haohan, et al.
Veröffentlicht: (2025)
von: Chi, Haohan, et al.
Veröffentlicht: (2025)
Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
AIGCs Confuse AI Too: Investigating and Explaining Synthetic Image-induced Hallucinations in Large Vision-Language Models
von: Gao, Yifei, et al.
Veröffentlicht: (2024)
von: Gao, Yifei, et al.
Veröffentlicht: (2024)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
von: Liang, Cheng, et al.
Veröffentlicht: (2026)
von: Liang, Cheng, et al.
Veröffentlicht: (2026)
Unifying Visual and Vision-Language Tracking via Contrastive Learning
von: Ma, Yinchao, et al.
Veröffentlicht: (2024)
von: Ma, Yinchao, et al.
Veröffentlicht: (2024)
Pure-Pass: Fine-Grained, Adaptive Masking for Dynamic Token-Mixing Routing in Lightweight Image Super-Resolution
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
Bridging Vision and Language: Optimal Transport-Driven Radiology Report Generation via LLMs
von: Zhao, Haifeng, et al.
Veröffentlicht: (2025)
von: Zhao, Haifeng, et al.
Veröffentlicht: (2025)
Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
VPTracker: Global Vision-Language Tracking via Visual Prompt
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding
von: Zhu, Lanyun, et al.
Veröffentlicht: (2024)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2024)
Do Vision Models Develop Human-Like Progressive Difficulty Understanding?
von: Huang, Zeyi, et al.
Veröffentlicht: (2025)
von: Huang, Zeyi, et al.
Veröffentlicht: (2025)
KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object Detection
von: Li, Xingyuan, et al.
Veröffentlicht: (2025)
von: Li, Xingyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient Test-Time Prompt Tuning for Vision-Language Models
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024) -
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025) -
Dual DETRs for Multi-Label Temporal Action Detection
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024) -
Asymmetric Masked Distillation for Pre-Training Small Foundation Models
von: Zhao, Zhiyu, et al.
Veröffentlicht: (2023) -
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)