Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Yaxin, Shen, Zhiqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025)
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025)
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
von: Shen, Huawen, et al.
Veröffentlicht: (2024)
von: Shen, Huawen, et al.
Veröffentlicht: (2024)
Vision-Language Models Create Cross-Modal Task Representations
von: Luo, Grace, et al.
Veröffentlicht: (2024)
von: Luo, Grace, et al.
Veröffentlicht: (2024)
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026)
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
von: Qi, Yayun, et al.
Veröffentlicht: (2024)
von: Qi, Yayun, et al.
Veröffentlicht: (2024)
When are Lemons Purple? The Concept Association Bias of Vision-Language Models
von: Yamada, Yutaro, et al.
Veröffentlicht: (2022)
von: Yamada, Yutaro, et al.
Veröffentlicht: (2022)
Renaissance: Investigating the Pretraining of Vision-Language Encoders
von: Fields, Clayton, et al.
Veröffentlicht: (2024)
von: Fields, Clayton, et al.
Veröffentlicht: (2024)
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
von: Ging, Simon, et al.
Veröffentlicht: (2026)
von: Ging, Simon, et al.
Veröffentlicht: (2026)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
Federated Learning from Vision-Language Foundation Models: Theoretical Analysis and Method
von: Pan, Bikang, et al.
Veröffentlicht: (2024)
von: Pan, Bikang, et al.
Veröffentlicht: (2024)
Mordal: Automated Pretrained Model Selection for Vision Language Models
von: He, Shiqi, et al.
Veröffentlicht: (2025)
von: He, Shiqi, et al.
Veröffentlicht: (2025)
A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models
von: Xiu, Lixin, et al.
Veröffentlicht: (2026)
von: Xiu, Lixin, et al.
Veröffentlicht: (2026)
BioCLIP: A Vision Foundation Model for the Tree of Life
von: Stevens, Samuel, et al.
Veröffentlicht: (2023)
von: Stevens, Samuel, et al.
Veröffentlicht: (2023)
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
von: Zhao, Dachuan, et al.
Veröffentlicht: (2025)
von: Zhao, Dachuan, et al.
Veröffentlicht: (2025)
Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
von: Park, Kwanyong, et al.
Veröffentlicht: (2024)
von: Park, Kwanyong, et al.
Veröffentlicht: (2024)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models
von: Luo, Jun, et al.
Veröffentlicht: (2024)
von: Luo, Jun, et al.
Veröffentlicht: (2024)
Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization
von: Luo, Richard, et al.
Veröffentlicht: (2024)
von: Luo, Richard, et al.
Veröffentlicht: (2024)
Distributionally Robust Alignment for Medical Federated Vision-Language Pre-training Under Data Heterogeneity
von: Shuai, Zitao, et al.
Veröffentlicht: (2024)
von: Shuai, Zitao, et al.
Veröffentlicht: (2024)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
von: Ma, Yan, et al.
Veröffentlicht: (2025)
von: Ma, Yan, et al.
Veröffentlicht: (2025)
A Practitioner's Guide to Continual Multimodal Pretraining
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
Unleashing the Potential of Model Bias for Generalized Category Discovery
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
Context-Aware Multimodal Pretraining
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
A Vision Check-up for Language Models
von: Sharma, Pratyusha, et al.
Veröffentlicht: (2024)
von: Sharma, Pratyusha, et al.
Veröffentlicht: (2024)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation
von: Jung, Hoin, et al.
Veröffentlicht: (2026)
von: Jung, Hoin, et al.
Veröffentlicht: (2026)
A Survey on Hallucination in Large Vision-Language Models
von: Liu, Hanchao, et al.
Veröffentlicht: (2024)
von: Liu, Hanchao, et al.
Veröffentlicht: (2024)
The Neglected Tails in Vision-Language Models
von: Parashar, Shubham, et al.
Veröffentlicht: (2024)
von: Parashar, Shubham, et al.
Veröffentlicht: (2024)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023)
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023)
BiasConnect: Investigating Bias Interactions in Text-to-Image Models
von: Shukla, Pushkar, et al.
Veröffentlicht: (2025)
von: Shukla, Pushkar, et al.
Veröffentlicht: (2025)
Differentially Private Bias-Term Fine-tuning of Foundation Models
von: Bu, Zhiqi, et al.
Veröffentlicht: (2022)
von: Bu, Zhiqi, et al.
Veröffentlicht: (2022)
SOLO: A Single Transformer for Scalable Vision-Language Modeling
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
TULIP: Towards Unified Language-Image Pretraining
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
Calibrated Self-Rewarding Vision Language Models
von: Zhou, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025) -
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
von: Shen, Huawen, et al.
Veröffentlicht: (2024) -
Vision-Language Models Create Cross-Modal Task Representations
von: Luo, Grace, et al.
Veröffentlicht: (2024) -
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2026) -
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)