Vision Language Models are Biased
Fuente:
arXiv
Saved in:
| Main Authors: | Vo, An, Nguyen, Khai-Nguyen, Taesiri, Mohammad Reza, Dang, Vy Tuong, Nguyen, Anh Totti, Kim, Daeyoung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023)
by: Giang, et al.
Published: (2023)
Vision language models are blind: Failing to translate detailed visual features into words
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
by: Collins, Brandon, et al.
Published: (2026)
by: Collins, Brandon, et al.
Published: (2026)
Generating Synthetic Satellite Imagery With Deep-Learning Text-to-Image Models -- Technical Challenges and Implications for Monitoring and Verification
by: Nguyen, Tuong Vy, et al.
Published: (2024)
by: Nguyen, Tuong Vy, et al.
Published: (2024)
LiteGPT: Large Vision-Language Model for Joint Chest X-ray Localization and Classification Task
by: Le-Duc, Khai, et al.
Published: (2024)
by: Le-Duc, Khai, et al.
Published: (2024)
IBMA: An Imputation-Based Mixup Augmentation Using Self-Supervised Learning for Time Series Data
by: Nguyen, Dang Nha, et al.
Published: (2025)
by: Nguyen, Dang Nha, et al.
Published: (2025)
Allowing humans to interactively guide machines where to look does not always improve human-AI team's classification accuracy
by: Nguyen, Giang, et al.
Published: (2024)
by: Nguyen, Giang, et al.
Published: (2024)
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization
by: Nguyen, Son, et al.
Published: (2025)
by: Nguyen, Son, et al.
Published: (2025)
WriteViT: Handwritten Text Generation with Vision Transformer
by: Nam, Dang Hoai, et al.
Published: (2025)
by: Nam, Dang Hoai, et al.
Published: (2025)
MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification
by: Nguyen, Anh-Tien, et al.
Published: (2025)
by: Nguyen, Anh-Tien, et al.
Published: (2025)
Bellman Optimal Stepsize Straightening of Flow-Matching Models
by: Nguyen, Bao, et al.
Published: (2023)
by: Nguyen, Bao, et al.
Published: (2023)
S-Chain: Structured Visual Chain-of-Thought For Medicine
by: Le-Duc, Khai, et al.
Published: (2025)
by: Le-Duc, Khai, et al.
Published: (2025)
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
Energy-Based Sliced Wasserstein Distance
by: Nguyen, Khai, et al.
Published: (2023)
by: Nguyen, Khai, et al.
Published: (2023)
Sliced Wasserstein Estimation with Control Variates
by: Nguyen, Khai, et al.
Published: (2023)
by: Nguyen, Khai, et al.
Published: (2023)
Generating Synthetic Satellite Imagery for Rare Objects: An Empirical Comparison of Models and Metrics
by: Nguyen, Tuong Vy, et al.
Published: (2024)
by: Nguyen, Tuong Vy, et al.
Published: (2024)
Leveraging Habitat Information for Fine-grained Bird Identification
by: Nguyen, Tin, et al.
Published: (2023)
by: Nguyen, Tin, et al.
Published: (2023)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
by: Le, Minh, et al.
Published: (2025)
by: Le, Minh, et al.
Published: (2025)
FW-GAN: Frequency-Driven Handwriting Synthesis with Wave-Modulated MLP Generator
by: Khoa, Huynh Tong Dang, et al.
Published: (2025)
by: Khoa, Huynh Tong Dang, et al.
Published: (2025)
Improving Diversity in Black-box Few-shot Knowledge Distillation
by: Vo, Tri-Nhan, et al.
Published: (2026)
by: Vo, Tri-Nhan, et al.
Published: (2026)
Hierarchical Hybrid Sliced Wasserstein: A Scalable Metric for Heterogeneous Joint Distributions
by: Nguyen, Khai, et al.
Published: (2024)
by: Nguyen, Khai, et al.
Published: (2024)
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
by: Naharas, Nilay, et al.
Published: (2025)
by: Naharas, Nilay, et al.
Published: (2025)
GlitchBench: Can large multimodal models detect video game glitches?
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
Diverse Image Priors for Black-box Data-free Knowledge Distillation
by: Vo, Tri-Nhan, et al.
Published: (2026)
by: Vo, Tri-Nhan, et al.
Published: (2026)
Diversity-Aware Agnostic Ensemble of Sharpness Minimizers
by: Bui, Anh, et al.
Published: (2024)
by: Bui, Anh, et al.
Published: (2024)
LoG-VMamba: Local-Global Vision Mamba for Medical Image Segmentation
by: Dang, Trung Dinh Quoc, et al.
Published: (2024)
by: Dang, Trung Dinh Quoc, et al.
Published: (2024)
HTR-ConvText: Leveraging Convolution and Textual Information for Handwritten Text Recognition
by: Truc, Pham Thach Thanh, et al.
Published: (2025)
by: Truc, Pham Thach Thanh, et al.
Published: (2025)
EVL-ECG: Efficient ECG Interpretation With Multi-Aspect Heterogeneous Knowledge Distillation
by: Hong, Dang Nguyen, et al.
Published: (2026)
by: Hong, Dang Nguyen, et al.
Published: (2026)
CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
by: Le, Tri, et al.
Published: (2025)
by: Le, Tri, et al.
Published: (2025)
Retrospective Feature Estimation for Continual Learning
by: Nguyen, Nghia D., et al.
Published: (2024)
by: Nguyen, Nghia D., et al.
Published: (2024)
Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Dual-Model Defense: Safeguarding Diffusion Models from Membership Inference Attacks through Disjoint Data Splitting
by: Tran, Bao Q., et al.
Published: (2024)
by: Tran, Bao Q., et al.
Published: (2024)
Sliced Wasserstein with Random-Path Projecting Directions
by: Nguyen, Khai, et al.
Published: (2024)
by: Nguyen, Khai, et al.
Published: (2024)
Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
by: Nguyen, Khoi Anh, et al.
Published: (2025)
by: Nguyen, Khoi Anh, et al.
Published: (2025)
Memory-efficient Continual Learning with Neural Collapse Contrastive
by: Dang, Trung-Anh, et al.
Published: (2024)
by: Dang, Trung-Anh, et al.
Published: (2024)
Using Game Engines and Machine Learning to Create Synthetic Satellite Imagery for a Tabletop Verification Exercise
by: Hoster, Johannes, et al.
Published: (2024)
by: Hoster, Johannes, et al.
Published: (2024)
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
by: Nguyen, Hung Huy, et al.
Published: (2025)
by: Nguyen, Hung Huy, et al.
Published: (2025)
N-EIoU-YOLOv9: A Signal-Aware Bounding Box Regression Loss for Lightweight Mobile Detection of Rice Leaf Diseases
by: Duc, Dung Ta Nguyen, et al.
Published: (2026)
by: Duc, Dung Ta Nguyen, et al.
Published: (2026)
Similar Items
-
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025) -
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023) -
Vision language models are blind: Failing to translate detailed visual features into words
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024) -
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
by: Collins, Brandon, et al.
Published: (2026) -
Generating Synthetic Satellite Imagery With Deep-Learning Text-to-Image Models -- Technical Challenges and Implications for Monitoring and Verification
by: Nguyen, Tuong Vy, et al.
Published: (2024)