Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
Fuente:
arXiv
Guardado en:
| Autores principales: | Rudman, William, Golovanevsky, Michal, Bar, Amir, Palit, Vedant, LeCun, Yann, Eickhoff, Carsten, Singh, Ritambhara |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
What Do VLMs NOTICE? A Mechanistic Interpretability Pipeline for Gaussian-Noise-free Text-Image Corruption and Evaluation
por: Golovanevsky, Michal, et al.
Publicado: (2024)
por: Golovanevsky, Michal, et al.
Publicado: (2024)
Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
por: Golovanevsky, Michal, et al.
Publicado: (2025)
por: Golovanevsky, Michal, et al.
Publicado: (2025)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
por: Rudman, William, et al.
Publicado: (2026)
por: Rudman, William, et al.
Publicado: (2026)
Is There Knowledge Left to Extract? Evidence of Fragility in Medically Fine-Tuned Vision-Language Models
por: McLaughlin, Oliver, et al.
Publicado: (2026)
por: McLaughlin, Oliver, et al.
Publicado: (2026)
Stable Anisotropic Regularization
por: Rudman, William, et al.
Publicado: (2023)
por: Rudman, William, et al.
Publicado: (2023)
Outlier Dimensions Encode Task-Specific Knowledge
por: Rudman, William, et al.
Publicado: (2023)
por: Rudman, William, et al.
Publicado: (2023)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
por: Nemitz, Jonathan, et al.
Publicado: (2026)
por: Nemitz, Jonathan, et al.
Publicado: (2026)
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
por: Huang, Hai, et al.
Publicado: (2025)
por: Huang, Hai, et al.
Publicado: (2025)
PiCME: Pipeline for Contrastive Modality Evaluation and Encoding in the MIMIC Dataset
por: Golovanevsky, Michal, et al.
Publicado: (2025)
por: Golovanevsky, Michal, et al.
Publicado: (2025)
TRIM: Achieving Extreme Sparsity with Targeted Row-wise Iterative Metric-driven Pruning
por: Beck, Florentin, et al.
Publicado: (2025)
por: Beck, Florentin, et al.
Publicado: (2025)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
por: Skean, Oscar, et al.
Publicado: (2024)
por: Skean, Oscar, et al.
Publicado: (2024)
Less is More: Label-Guided Summarization of Procedural and Instructional Videos
por: Rajpal, Shreya, et al.
Publicado: (2026)
por: Rajpal, Shreya, et al.
Publicado: (2026)
UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages
por: Abdullahi, Tassallah, et al.
Publicado: (2026)
por: Abdullahi, Tassallah, et al.
Publicado: (2026)
From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures
por: Rottach, Florian, et al.
Publicado: (2025)
por: Rottach, Florian, et al.
Publicado: (2025)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
por: Sun, Shangwen, et al.
Publicado: (2026)
por: Sun, Shangwen, et al.
Publicado: (2026)
One-Versus-Others Attention: Scalable Multimodal Integration for Biomedical Data
por: Golovanevsky, Michal, et al.
Publicado: (2023)
por: Golovanevsky, Michal, et al.
Publicado: (2023)
APP: Accelerated Path Patching with Task-Specific Pruning
por: Andersen, Frauke, et al.
Publicado: (2025)
por: Andersen, Frauke, et al.
Publicado: (2025)
The Persona Paradox: Medical Personas as Behavioral Priors in Clinical Language Models
por: Abdullahi, Tassallah, et al.
Publicado: (2026)
por: Abdullahi, Tassallah, et al.
Publicado: (2026)
Transformers without Normalization
por: Zhu, Jiachen, et al.
Publicado: (2025)
por: Zhu, Jiachen, et al.
Publicado: (2025)
Navigation World Models
por: Bar, Amir, et al.
Publicado: (2024)
por: Bar, Amir, et al.
Publicado: (2024)
K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
por: Abdullahi, Tassallah, et al.
Publicado: (2025)
por: Abdullahi, Tassallah, et al.
Publicado: (2025)
CroCoSum: A Benchmark Dataset for Cross-Lingual Code-Switched Summarization
por: Zhang, Ruochen, et al.
Publicado: (2023)
por: Zhang, Ruochen, et al.
Publicado: (2023)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
por: Balestriero, Randall, et al.
Publicado: (2025)
por: Balestriero, Randall, et al.
Publicado: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
por: Balestriero, Randall, et al.
Publicado: (2024)
por: Balestriero, Randall, et al.
Publicado: (2024)
Large Language Models for Mental Health Diagnostic Assessments: Exploring The Potential of Large Language Models for Assisting with Mental Health Diagnostic Assessments -- The Depression and Anxiety Case
por: Roy, Kaushik, et al.
Publicado: (2025)
por: Roy, Kaushik, et al.
Publicado: (2025)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
por: Shani, Chen, et al.
Publicado: (2025)
por: Shani, Chen, et al.
Publicado: (2025)
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
por: Mohammadi, Seyedali, et al.
Publicado: (2024)
por: Mohammadi, Seyedali, et al.
Publicado: (2024)
Circuit Component Reuse Across Tasks in Transformer Language Models
por: Merullo, Jack, et al.
Publicado: (2023)
por: Merullo, Jack, et al.
Publicado: (2023)
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
por: Merullo, Jack, et al.
Publicado: (2024)
por: Merullo, Jack, et al.
Publicado: (2024)
Language Models Implement Simple Word2Vec-style Vector Arithmetic
por: Merullo, Jack, et al.
Publicado: (2023)
por: Merullo, Jack, et al.
Publicado: (2023)
Layer by Layer: Uncovering Hidden Representations in Language Models
por: Skean, Oscar, et al.
Publicado: (2025)
por: Skean, Oscar, et al.
Publicado: (2025)
Whole-Body Conditioned Egocentric Video Prediction
por: Bai, Yutong, et al.
Publicado: (2025)
por: Bai, Yutong, et al.
Publicado: (2025)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
por: Zhai, Yuexiang, et al.
Publicado: (2024)
por: Zhai, Yuexiang, et al.
Publicado: (2024)
Multimodal QUD: Inquisitive Questions from Scientific Figures
por: Wu, Yating, et al.
Publicado: (2026)
por: Wu, Yating, et al.
Publicado: (2026)
A hierarchical loss and its problems when classifying non-hierarchically
por: Wu, Cinna, et al.
Publicado: (2017)
por: Wu, Cinna, et al.
Publicado: (2017)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
por: Arefin, Md Rifat, et al.
Publicado: (2024)
por: Arefin, Md Rifat, et al.
Publicado: (2024)
PubMed-Ophtha: An open resource for training ophthalmology vision-language models on scientific literature
por: Hallitschke, Verena Jasmin, et al.
Publicado: (2026)
por: Hallitschke, Verena Jasmin, et al.
Publicado: (2026)
EgoPet: Egomotion and Interaction Data from an Animal's Perspective
por: Bar, Amir, et al.
Publicado: (2024)
por: Bar, Amir, et al.
Publicado: (2024)
Unchecked and Overlooked: Addressing the Checkbox Blind Spot in Large Language Models with CheckboxQA
por: Turski, Michał, et al.
Publicado: (2025)
por: Turski, Michał, et al.
Publicado: (2025)
Paths Not Taken: Understanding and Mending the Multilingual Factual Recall Pipeline
por: Lu, Meng, et al.
Publicado: (2025)
por: Lu, Meng, et al.
Publicado: (2025)
Ejemplares similares
-
What Do VLMs NOTICE? A Mechanistic Interpretability Pipeline for Gaussian-Noise-free Text-Image Corruption and Evaluation
por: Golovanevsky, Michal, et al.
Publicado: (2024) -
Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
por: Golovanevsky, Michal, et al.
Publicado: (2025) -
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
por: Rudman, William, et al.
Publicado: (2026) -
Is There Knowledge Left to Extract? Evidence of Fragility in Medically Fine-Tuned Vision-Language Models
por: McLaughlin, Oliver, et al.
Publicado: (2026) -
Stable Anisotropic Regularization
por: Rudman, William, et al.
Publicado: (2023)