Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect
Fuente:
arXiv
Guardado en:
| Autores principales: | Kouwenhoven, Tom, Shahrasbi, Kiana, Verhoef, Tessa |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
What does Kiki look like? Cross-modal associations between speech sounds and visual shapes in vision-and-language models
por: Verhoef, Tessa, et al.
Publicado: (2024)
por: Verhoef, Tessa, et al.
Publicado: (2024)
Kiki or Bouba? Sound Symbolism in Vision-and-Language Models
por: Alper, Morris, et al.
Publicado: (2023)
por: Alper, Morris, et al.
Publicado: (2023)
Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls
por: Pitta, Elena, et al.
Publicado: (2025)
por: Pitta, Elena, et al.
Publicado: (2025)
Searching for Structure: Investigating Emergent Communication with Large Language Models
por: Kouwenhoven, Tom, et al.
Publicado: (2024)
por: Kouwenhoven, Tom, et al.
Publicado: (2024)
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
por: Dagan, Gautier, et al.
Publicado: (2024)
por: Dagan, Gautier, et al.
Publicado: (2024)
Shaping Shared Languages: Human and Large Language Models' Inductive Biases in Emergent Communication
por: Kouwenhoven, Tom, et al.
Publicado: (2025)
por: Kouwenhoven, Tom, et al.
Publicado: (2025)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
por: Liang, Qiao, et al.
Publicado: (2025)
por: Liang, Qiao, et al.
Publicado: (2025)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
por: Shang, Yuying, et al.
Publicado: (2024)
por: Shang, Yuying, et al.
Publicado: (2024)
Revisiting the Role of Language Priors in Vision-Language Models
por: Lin, Zhiqiu, et al.
Publicado: (2023)
por: Lin, Zhiqiu, et al.
Publicado: (2023)
Cross-modal Information Flow in Multimodal Large Language Models
por: Zhang, Zhi, et al.
Publicado: (2024)
por: Zhang, Zhi, et al.
Publicado: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
por: Zhang, Ming, et al.
Publicado: (2024)
por: Zhang, Ming, et al.
Publicado: (2024)
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
por: Miao, Yongzhu, et al.
Publicado: (2023)
por: Miao, Yongzhu, et al.
Publicado: (2023)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
por: Jiang, Lei, et al.
Publicado: (2025)
por: Jiang, Lei, et al.
Publicado: (2025)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
por: Zhu, Tinghui, et al.
Publicado: (2024)
por: Zhu, Tinghui, et al.
Publicado: (2024)
Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
por: Wang, Yabing, et al.
Publicado: (2024)
por: Wang, Yabing, et al.
Publicado: (2024)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
por: Li, Yanwei, et al.
Publicado: (2024)
por: Li, Yanwei, et al.
Publicado: (2024)
DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
por: Liu, Zixuan, et al.
Publicado: (2025)
por: Liu, Zixuan, et al.
Publicado: (2025)
Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval
por: Cai, Rui, et al.
Publicado: (2024)
por: Cai, Rui, et al.
Publicado: (2024)
CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers
por: Shi, Dachuan, et al.
Publicado: (2023)
por: Shi, Dachuan, et al.
Publicado: (2023)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
por: Miranda, Imanol, et al.
Publicado: (2026)
por: Miranda, Imanol, et al.
Publicado: (2026)
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
por: Huang, Qidong, et al.
Publicado: (2024)
por: Huang, Qidong, et al.
Publicado: (2024)
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
por: Nazi, Zabir Al, et al.
Publicado: (2025)
por: Nazi, Zabir Al, et al.
Publicado: (2025)
GroundingGPT:Language Enhanced Multi-modal Grounding Model
por: Li, Zhaowei, et al.
Publicado: (2024)
por: Li, Zhaowei, et al.
Publicado: (2024)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
por: Ohi, Masanari, et al.
Publicado: (2024)
por: Ohi, Masanari, et al.
Publicado: (2024)
Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
por: Gigant, Théo, et al.
Publicado: (2025)
por: Gigant, Théo, et al.
Publicado: (2025)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
por: Lin, Jingyang, et al.
Publicado: (2023)
por: Lin, Jingyang, et al.
Publicado: (2023)
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
por: Gao, Jianjun, et al.
Publicado: (2024)
por: Gao, Jianjun, et al.
Publicado: (2024)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
por: Irawan, Patrick Amadeus, et al.
Publicado: (2026)
por: Irawan, Patrick Amadeus, et al.
Publicado: (2026)
The Curious Case of Representational Alignment: Unravelling Visio-Linguistic Tasks in Emergent Communication
por: Kouwenhoven, Tom, et al.
Publicado: (2024)
por: Kouwenhoven, Tom, et al.
Publicado: (2024)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
por: Lee, Seongyun, et al.
Publicado: (2024)
por: Lee, Seongyun, et al.
Publicado: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
por: Xu, Xiaohao, et al.
Publicado: (2024)
por: Xu, Xiaohao, et al.
Publicado: (2024)
Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
por: Tian, Yuanhe, et al.
Publicado: (2025)
por: Tian, Yuanhe, et al.
Publicado: (2025)
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
por: Chen, Junzhe, et al.
Publicado: (2024)
por: Chen, Junzhe, et al.
Publicado: (2024)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
por: Zhang, Kaichen, et al.
Publicado: (2024)
por: Zhang, Kaichen, et al.
Publicado: (2024)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
por: Ye, Jiacheng, et al.
Publicado: (2025)
por: Ye, Jiacheng, et al.
Publicado: (2025)
Vision-Language Models Create Cross-Modal Task Representations
por: Luo, Grace, et al.
Publicado: (2024)
por: Luo, Grace, et al.
Publicado: (2024)
Cross-Cultural Value Awareness in Large Vision-Language Models
por: Howard, Phillip, et al.
Publicado: (2026)
por: Howard, Phillip, et al.
Publicado: (2026)
Hierarchical Cross-modal Prompt Learning for Vision-Language Models
por: Zheng, Hao, et al.
Publicado: (2025)
por: Zheng, Hao, et al.
Publicado: (2025)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
por: Yang, Rui, et al.
Publicado: (2025)
por: Yang, Rui, et al.
Publicado: (2025)
Ejemplares similares
-
What does Kiki look like? Cross-modal associations between speech sounds and visual shapes in vision-and-language models
por: Verhoef, Tessa, et al.
Publicado: (2024) -
Kiki or Bouba? Sound Symbolism in Vision-and-Language Models
por: Alper, Morris, et al.
Publicado: (2023) -
Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls
por: Pitta, Elena, et al.
Publicado: (2025) -
Searching for Structure: Investigating Emergent Communication with Large Language Models
por: Kouwenhoven, Tom, et al.
Publicado: (2024) -
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
por: Dagan, Gautier, et al.
Publicado: (2024)