How Culturally Aware are Vision-Language Models?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Burda-Lassen, Olena, Chadha, Aman, Goswami, Shashank, Jain, Vinija |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)
Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages
von: Burda-Lassen, Olena
Veröffentlicht: (2024)
von: Burda-Lassen, Olena
Veröffentlicht: (2024)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
The Evolution of Multimodal Model Architectures
von: Wadekar, Shakti N., et al.
Veröffentlicht: (2024)
von: Wadekar, Shakti N., et al.
Veröffentlicht: (2024)
How Do Vision-Language Models Process Conflicting Information Across Modalities?
von: Hua, Tianze, et al.
Veröffentlicht: (2025)
von: Hua, Tianze, et al.
Veröffentlicht: (2025)
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
von: Lakhanpal, Sanyam, et al.
Veröffentlicht: (2024)
von: Lakhanpal, Sanyam, et al.
Veröffentlicht: (2024)
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
von: Zhu, Kangyu, et al.
Veröffentlicht: (2024)
von: Zhu, Kangyu, et al.
Veröffentlicht: (2024)
Situational Awareness Matters in 3D Vision Language Reasoning
von: Man, Yunze, et al.
Veröffentlicht: (2024)
von: Man, Yunze, et al.
Veröffentlicht: (2024)
How Do Training Methods Influence the Utilization of Vision Models?
von: Gavrikov, Paul, et al.
Veröffentlicht: (2024)
von: Gavrikov, Paul, et al.
Veröffentlicht: (2024)
VisionZip: Longer is Better but Not Necessary in Vision Language Models
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
Decoding the Diversity: A Review of the Indic AI Research Landscape
von: KJ, Sankalp, et al.
Veröffentlicht: (2024)
von: KJ, Sankalp, et al.
Veröffentlicht: (2024)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
Imperfect Vision Encoders: Efficient and Robust Tuning for Vision-Language Models
von: Panos, Aristeidis, et al.
Veröffentlicht: (2024)
von: Panos, Aristeidis, et al.
Veröffentlicht: (2024)
Vision-Language Model Based Handwriting Verification
von: Chauhan, Mihir, et al.
Veröffentlicht: (2024)
von: Chauhan, Mihir, et al.
Veröffentlicht: (2024)
Differentiable Prompt Learning for Vision Language Models
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
von: Yang, Senqiao, et al.
Veröffentlicht: (2025)
von: Yang, Senqiao, et al.
Veröffentlicht: (2025)
AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
von: Lin, Zichuan, et al.
Veröffentlicht: (2025)
von: Lin, Zichuan, et al.
Veröffentlicht: (2025)
Energy-Based Transformers are Scalable Learners and Thinkers
von: Gladstone, Alexi, et al.
Veröffentlicht: (2025)
von: Gladstone, Alexi, et al.
Veröffentlicht: (2025)
Mamba in Vision: A Comprehensive Survey of Techniques and Applications
von: Rahman, Md Maklachur, et al.
Veröffentlicht: (2024)
von: Rahman, Md Maklachur, et al.
Veröffentlicht: (2024)
Mordal: Automated Pretrained Model Selection for Vision Language Models
von: He, Shiqi, et al.
Veröffentlicht: (2025)
von: He, Shiqi, et al.
Veröffentlicht: (2025)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
Efficient Architectures for High Resolution Vision-Language Models
von: Carvalho, Miguel, et al.
Veröffentlicht: (2025)
von: Carvalho, Miguel, et al.
Veröffentlicht: (2025)
Understanding the Effects of Distractors on Reasoning Vision-Language Models
von: Bae, Jiyun, et al.
Veröffentlicht: (2025)
von: Bae, Jiyun, et al.
Veröffentlicht: (2025)
Coordinated Robustness Evaluation Framework for Vision-Language Models
von: Babu, Ashwin Ramesh, et al.
Veröffentlicht: (2025)
von: Babu, Ashwin Ramesh, et al.
Veröffentlicht: (2025)
Benchmarking Vision Language Models for Cultural Understanding
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation
von: Padhan, Swagat, et al.
Veröffentlicht: (2026)
von: Padhan, Swagat, et al.
Veröffentlicht: (2026)
MouSi: Poly-Visual-Expert Vision-Language Models
von: Fan, Xiaoran, et al.
Veröffentlicht: (2024)
von: Fan, Xiaoran, et al.
Veröffentlicht: (2024)
Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models
von: Fu, Shuai, et al.
Veröffentlicht: (2024)
von: Fu, Shuai, et al.
Veröffentlicht: (2024)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
von: Trager, Matthew, et al.
Veröffentlicht: (2023)
von: Trager, Matthew, et al.
Veröffentlicht: (2023)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
von: Li, YuQian, et al.
Veröffentlicht: (2025)
von: Li, YuQian, et al.
Veröffentlicht: (2025)
Unified Vision-Language Modeling via Concept Space Alignment
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024) -
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
von: Ghosh, Akash, et al.
Veröffentlicht: (2024) -
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024) -
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024) -
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)