Visual cognition in multimodal large language models
Fuente:
arXiv
Saved in:
| Main Authors: | Buschoff, Luca M. Schulze, Akata, Elif, Bethge, Matthias, Schulz, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language Models
by: Buschoff, Luca M. Schulze, et al.
Published: (2025)
by: Buschoff, Luca M. Schulze, et al.
Published: (2025)
Can Vision Language Models Learn Intuitive Physics from Interaction?
by: Buschoff, Luca M. Schulze, et al.
Published: (2026)
by: Buschoff, Luca M. Schulze, et al.
Published: (2026)
metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
by: Kipnis, Alex, et al.
Published: (2024)
by: Kipnis, Alex, et al.
Published: (2024)
Next state prediction gives rise to entangled, yet compositional representations of objects
by: Saanum, Tankred, et al.
Published: (2024)
by: Saanum, Tankred, et al.
Published: (2024)
In-Context Function Learning in Large Language Models
by: Akata, Elif, et al.
Published: (2026)
by: Akata, Elif, et al.
Published: (2026)
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023)
by: Coda-Forno, Julian, et al.
Published: (2023)
Centaur: a foundation model of human cognition
by: Binz, Marcel, et al.
Published: (2024)
by: Binz, Marcel, et al.
Published: (2024)
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
by: Thede, Lukas, et al.
Published: (2025)
by: Thede, Lukas, et al.
Published: (2025)
Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models
by: Thede, Lukas, et al.
Published: (2024)
by: Thede, Lukas, et al.
Published: (2024)
Insights into a radiology-specialised multimodal large language model with sparse autoencoders
by: Bouzid, Kenza, et al.
Published: (2025)
by: Bouzid, Kenza, et al.
Published: (2025)
Discovering Chunks in Neural Embeddings for Interpretability
by: Wu, Shuchen, et al.
Published: (2025)
by: Wu, Shuchen, et al.
Published: (2025)
CogBench: a large language model walks into a psychology lab
by: Coda-Forno, Julian, et al.
Published: (2024)
by: Coda-Forno, Julian, et al.
Published: (2024)
Post-training makes large language models less human-like
by: Binz, Marcel, et al.
Published: (2026)
by: Binz, Marcel, et al.
Published: (2026)
Playing repeated games with Large Language Models
by: Akata, Elif, et al.
Published: (2023)
by: Akata, Elif, et al.
Published: (2023)
Equivariance by Contrast: Identifiable Equivariant Embeddings from Unlabeled Finite Group Actions
by: Schmidt, Tobias, et al.
Published: (2025)
by: Schmidt, Tobias, et al.
Published: (2025)
How to Merge Your Multimodal Models Over Time?
by: Dziadzio, Sebastian, et al.
Published: (2024)
by: Dziadzio, Sebastian, et al.
Published: (2024)
Amortizing intractable inference in large language models
by: Hu, Edward J., et al.
Published: (2023)
by: Hu, Edward J., et al.
Published: (2023)
Visual concept ranking uncovers medical shortcuts used by large multimodal models
by: Janizek, Joseph D., et al.
Published: (2026)
by: Janizek, Joseph D., et al.
Published: (2026)
Probing the limitations of multimodal language models for chemistry and materials research
by: Alampara, Nawaf, et al.
Published: (2024)
by: Alampara, Nawaf, et al.
Published: (2024)
Building, Reusing, and Generalizing Abstract Representations from Concrete Sequences
by: Wu, Shuchen, et al.
Published: (2024)
by: Wu, Shuchen, et al.
Published: (2024)
Hypothesis generation and updating in large language models
by: Xiong, Hua-Dong
Published: (2026)
by: Xiong, Hua-Dong
Published: (2026)
Towards an automated workflow in materials science for combining multi-modal simulative and experimental information using data mining and large language models
by: Katzer, Balduin, et al.
Published: (2025)
by: Katzer, Balduin, et al.
Published: (2025)
Visual representations in the human brain are aligned with large language models
by: Doerig, Adrien, et al.
Published: (2022)
by: Doerig, Adrien, et al.
Published: (2022)
Prompt reinforcing for long-term planning of large language models
by: Lin, Hsien-Chin, et al.
Published: (2025)
by: Lin, Hsien-Chin, et al.
Published: (2025)
Human-like object concept representations emerge naturally in multimodal large language models
by: Du, Changde, et al.
Published: (2024)
by: Du, Changde, et al.
Published: (2024)
RDumb: A simple approach that questions our progress in continual test-time adaptation
by: Press, Ori, et al.
Published: (2023)
by: Press, Ori, et al.
Published: (2023)
Modeling Saliency Dataset Bias
by: Kümmerer, Matthias, et al.
Published: (2025)
by: Kümmerer, Matthias, et al.
Published: (2025)
Representation in large language models
by: Yetman, Cameron
Published: (2025)
by: Yetman, Cameron
Published: (2025)
Visual hallucination detection in large vision-language models via evidential conflict
by: Huang, Tao, et al.
Published: (2025)
by: Huang, Tao, et al.
Published: (2025)
Layer-wise dynamic rank for compressing large language models
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
NIRVANA: Structured pruning reimagined for large language models compression
by: Ai, Mengting, et al.
Published: (2025)
by: Ai, Mengting, et al.
Published: (2025)
Training microrobots to swim by a large language model
by: Xu, Zhuoqun, et al.
Published: (2024)
by: Xu, Zhuoqun, et al.
Published: (2024)
Identifying latent state transition in non-linear dynamical systems
by: Hızlı, Çağlar, et al.
Published: (2024)
by: Hızlı, Çağlar, et al.
Published: (2024)
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
by: Wiedemer, Thaddäus, et al.
Published: (2025)
by: Wiedemer, Thaddäus, et al.
Published: (2025)
Are large language models superhuman chemists?
by: Mirza, Adrian, et al.
Published: (2024)
by: Mirza, Adrian, et al.
Published: (2024)
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
by: Oota, Subba Reddy, et al.
Published: (2025)
by: Oota, Subba Reddy, et al.
Published: (2025)
Alignment faking in large language models
by: Greenblatt, Ryan, et al.
Published: (2024)
by: Greenblatt, Ryan, et al.
Published: (2024)
Reference-Free Rating of LLM Responses via Latent Information
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
AI-AI Bias: large language models favor communications generated by large language models
by: Laurito, Walter, et al.
Published: (2024)
by: Laurito, Walter, et al.
Published: (2024)
Concept-Guided Interpretability via Neural Chunking
by: Wu, Shuchen, et al.
Published: (2025)
by: Wu, Shuchen, et al.
Published: (2025)
Similar Items
-
Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language Models
by: Buschoff, Luca M. Schulze, et al.
Published: (2025) -
Can Vision Language Models Learn Intuitive Physics from Interaction?
by: Buschoff, Luca M. Schulze, et al.
Published: (2026) -
metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
by: Kipnis, Alex, et al.
Published: (2024) -
Next state prediction gives rise to entangled, yet compositional representations of objects
by: Saanum, Tankred, et al.
Published: (2024) -
In-Context Function Learning in Large Language Models
by: Akata, Elif, et al.
Published: (2026)