Pre-Trained Vision-Language Models as Partial Annotators
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Qian-Wei, Xie, Yuqiu, Zhang, Letian, Liu, Zimo, Xia, Shu-Tao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fine-tuning Pre-trained Vision-Language Models in a Human-Annotation-Free Manner
di: Wang, Qian-Wei, et al.
Pubblicazione: (2026)
di: Wang, Qian-Wei, et al.
Pubblicazione: (2026)
Bridging Weakly-Supervised Learning and VLM Distillation: Noisy Partial Label Learning for Efficient Downstream Adaptation
di: Wang, Qian-Wei, et al.
Pubblicazione: (2025)
di: Wang, Qian-Wei, et al.
Pubblicazione: (2025)
Explicit Uncertainty Modeling for Active CLIP Adaptation with Dual Prompt Tuning
di: Wang, Qian-Wei, et al.
Pubblicazione: (2026)
di: Wang, Qian-Wei, et al.
Pubblicazione: (2026)
Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models
di: Fan, Jiawei, et al.
Pubblicazione: (2026)
di: Fan, Jiawei, et al.
Pubblicazione: (2026)
HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained Models
di: Liu, Han, et al.
Pubblicazione: (2026)
di: Liu, Han, et al.
Pubblicazione: (2026)
Avoid Wasted Annotation Costs in Open-set Active Learning with Pre-trained Vision-Language Model
di: Heo, Jaehyuk, et al.
Pubblicazione: (2024)
di: Heo, Jaehyuk, et al.
Pubblicazione: (2024)
3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation
di: Zhang, Jusheng, et al.
Pubblicazione: (2026)
di: Zhang, Jusheng, et al.
Pubblicazione: (2026)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
di: Zhang, Xinsong, et al.
Pubblicazione: (2025)
di: Zhang, Xinsong, et al.
Pubblicazione: (2025)
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
di: Zhao, Bingchen, et al.
Pubblicazione: (2024)
di: Zhao, Bingchen, et al.
Pubblicazione: (2024)
Holmes: Towards Effective and Harmless Model Ownership Verification to Personalized Large Vision Models via Decoupling Common Features
di: Zhu, Linghui, et al.
Pubblicazione: (2025)
di: Zhu, Linghui, et al.
Pubblicazione: (2025)
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
di: Kumar, Ashutosh, et al.
Pubblicazione: (2026)
di: Kumar, Ashutosh, et al.
Pubblicazione: (2026)
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
di: Liu, Jinlong, et al.
Pubblicazione: (2026)
di: Liu, Jinlong, et al.
Pubblicazione: (2026)
ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving
di: Sha, Lin, et al.
Pubblicazione: (2026)
di: Sha, Lin, et al.
Pubblicazione: (2026)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
di: Tao, Hongyuan, et al.
Pubblicazione: (2025)
di: Tao, Hongyuan, et al.
Pubblicazione: (2025)
Training-Free Unsupervised Prompt for Vision-Language Models
di: Long, Sifan, et al.
Pubblicazione: (2024)
di: Long, Sifan, et al.
Pubblicazione: (2024)
PPU-Bench:Real World Benchmark for Personalized Partial Unlearning in Vision Language Models
di: Guang, Jiahui, et al.
Pubblicazione: (2026)
di: Guang, Jiahui, et al.
Pubblicazione: (2026)
Multimodal Fusion with Pre-Trained Model Features in Affective Behaviour Analysis In-the-wild
di: Wen, Zhuofan, et al.
Pubblicazione: (2024)
di: Wen, Zhuofan, et al.
Pubblicazione: (2024)
Do Pre-trained Vision-Language Models Encode Object States?
di: Newman, Kaleb, et al.
Pubblicazione: (2024)
di: Newman, Kaleb, et al.
Pubblicazione: (2024)
DeVAn: Dense Video Annotation for Video-Language Models
di: Liu, Tingkai, et al.
Pubblicazione: (2023)
di: Liu, Tingkai, et al.
Pubblicazione: (2023)
VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
di: Dong, Haotian, et al.
Pubblicazione: (2025)
di: Dong, Haotian, et al.
Pubblicazione: (2025)
Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias
di: Wan, Zhongwei, et al.
Pubblicazione: (2023)
di: Wan, Zhongwei, et al.
Pubblicazione: (2023)
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
di: Chen, Yi-Chia, et al.
Pubblicazione: (2024)
di: Chen, Yi-Chia, et al.
Pubblicazione: (2024)
Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization
di: Kawaharazuka, Kento, et al.
Pubblicazione: (2024)
di: Kawaharazuka, Kento, et al.
Pubblicazione: (2024)
Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute Alignment
di: Jiang, Yankai, et al.
Pubblicazione: (2024)
di: Jiang, Yankai, et al.
Pubblicazione: (2024)
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
di: Liu, Ming, et al.
Pubblicazione: (2025)
di: Liu, Ming, et al.
Pubblicazione: (2025)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
di: Wu, Junfei, et al.
Pubblicazione: (2025)
di: Wu, Junfei, et al.
Pubblicazione: (2025)
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
di: Huang, Han, et al.
Pubblicazione: (2024)
di: Huang, Han, et al.
Pubblicazione: (2024)
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
di: Zhou, Yuhao, et al.
Pubblicazione: (2026)
di: Zhou, Yuhao, et al.
Pubblicazione: (2026)
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
di: Wang, Lu, et al.
Pubblicazione: (2025)
di: Wang, Lu, et al.
Pubblicazione: (2025)
VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events
di: Bhat, Mohammad Qazim, et al.
Pubblicazione: (2026)
di: Bhat, Mohammad Qazim, et al.
Pubblicazione: (2026)
Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
di: Zhang, Naifu, et al.
Pubblicazione: (2025)
di: Zhang, Naifu, et al.
Pubblicazione: (2025)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
di: Ye, Wei, et al.
Pubblicazione: (2024)
di: Ye, Wei, et al.
Pubblicazione: (2024)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
di: Jiang, Chaoya, et al.
Pubblicazione: (2023)
di: Jiang, Chaoya, et al.
Pubblicazione: (2023)
Synthesizing Efficient Data with Diffusion Models for Person Re-Identification Pre-Training
di: Niu, Ke, et al.
Pubblicazione: (2024)
di: Niu, Ke, et al.
Pubblicazione: (2024)
EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
di: Yuan, Botai, et al.
Pubblicazione: (2025)
di: Yuan, Botai, et al.
Pubblicazione: (2025)
Grounded Knowledge-Enhanced Medical Vision-Language Pre-training for Chest X-Ray
di: Deng, Qiao, et al.
Pubblicazione: (2024)
di: Deng, Qiao, et al.
Pubblicazione: (2024)
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval
di: Wang, Yuting, et al.
Pubblicazione: (2023)
di: Wang, Yuting, et al.
Pubblicazione: (2023)
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
di: Zhang, Shu-Hao, et al.
Pubblicazione: (2025)
di: Zhang, Shu-Hao, et al.
Pubblicazione: (2025)
Can Medical Vision-Language Pre-training Succeed with Purely Synthetic Data?
di: Liu, Che, et al.
Pubblicazione: (2024)
di: Liu, Che, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Fine-tuning Pre-trained Vision-Language Models in a Human-Annotation-Free Manner
di: Wang, Qian-Wei, et al.
Pubblicazione: (2026) -
Bridging Weakly-Supervised Learning and VLM Distillation: Noisy Partial Label Learning for Efficient Downstream Adaptation
di: Wang, Qian-Wei, et al.
Pubblicazione: (2025) -
Explicit Uncertainty Modeling for Active CLIP Adaptation with Dual Prompt Tuning
di: Wang, Qian-Wei, et al.
Pubblicazione: (2026) -
Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models
di: Fan, Jiawei, et al.
Pubblicazione: (2026) -
HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained Models
di: Liu, Han, et al.
Pubblicazione: (2026)