GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
Fuente:
arXiv
Saved in:
| Main Authors: | Jassim, Serwan, Holubar, Mario, Richter, Annika, Wolff, Cornelius, Ohmer, Xenia, Bruni, Elia |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bidirectional Emergent Language in Situated Environments
by: Wolff, Cornelius, et al.
Published: (2024)
by: Wolff, Cornelius, et al.
Published: (2024)
Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models
by: Ballout, Mohamad, et al.
Published: (2025)
by: Ballout, Mohamad, et al.
Published: (2025)
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
by: Ohmer, Xenia, et al.
Published: (2024)
by: Ohmer, Xenia, et al.
Published: (2024)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
by: Mayer, Julius, et al.
Published: (2025)
by: Mayer, Julius, et al.
Published: (2025)
On the Relationship between Skill Neurons and Robustness in Prompt Tuning
by: Ackermann, Leon, et al.
Published: (2023)
by: Ackermann, Leon, et al.
Published: (2023)
DevBench: A multimodal developmental benchmark for language learning
by: Tan, Alvin Wei Ming, et al.
Published: (2024)
by: Tan, Alvin Wei Ming, et al.
Published: (2024)
\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
by: Min, Junghyun, et al.
Published: (2025)
by: Min, Junghyun, et al.
Published: (2025)
DocLLM: A layout-aware generative language model for multimodal document understanding
by: Wang, Dongsheng, et al.
Published: (2023)
by: Wang, Dongsheng, et al.
Published: (2023)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
by: Li, Kunning, et al.
Published: (2025)
by: Li, Kunning, et al.
Published: (2025)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
by: Wen, Zichen, et al.
Published: (2024)
by: Wen, Zichen, et al.
Published: (2024)
MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models
by: Yoshitake, Michiko, et al.
Published: (2026)
by: Yoshitake, Michiko, et al.
Published: (2026)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
by: Matos, João, et al.
Published: (2024)
by: Matos, João, et al.
Published: (2024)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
by: Rädsch, Tim, et al.
Published: (2025)
by: Rädsch, Tim, et al.
Published: (2025)
BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
by: Lavechin, Marvin, et al.
Published: (2023)
by: Lavechin, Marvin, et al.
Published: (2023)
Protecting multimodal large language models against misleading visualizations
by: Tonglet, Jonathan, et al.
Published: (2025)
by: Tonglet, Jonathan, et al.
Published: (2025)
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
by: Padlewski, Piotr, et al.
Published: (2024)
by: Padlewski, Piotr, et al.
Published: (2024)
Anthropocentric bias in language model evaluation
by: Millière, Raphaël, et al.
Published: (2024)
by: Millière, Raphaël, et al.
Published: (2024)
What does it mean to understand language?
by: Casto, Colton, et al.
Published: (2025)
by: Casto, Colton, et al.
Published: (2025)
Is your multimodal large language model a good science tutor?
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
Anchor function: a type of benchmark functions for studying language models
by: Zhang, Zhongwang, et al.
Published: (2024)
by: Zhang, Zhongwang, et al.
Published: (2024)
LongTail-Swap: benchmarking language models' abilities on rare words
by: Algayres, Robin, et al.
Published: (2025)
by: Algayres, Robin, et al.
Published: (2025)
Linguini: A benchmark for language-agnostic linguistic reasoning
by: Sánchez, Eduardo, et al.
Published: (2024)
by: Sánchez, Eduardo, et al.
Published: (2024)
FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models
by: Shamsfard, Mehrnoush, et al.
Published: (2025)
by: Shamsfard, Mehrnoush, et al.
Published: (2025)
As easy as PIE: understanding when pruning causes language models to disagree
by: Tropeano, Pietro, et al.
Published: (2025)
by: Tropeano, Pietro, et al.
Published: (2025)
Are LLM-generated plain language summaries truly understandable? A large-scale crowdsourced evaluation
by: Guo, Yue, et al.
Published: (2025)
by: Guo, Yue, et al.
Published: (2025)
The intersection of philosophy of language and artificial intelligence: Challenges in replicating human language understanding
by: Sooraj Kumar Maurya
Published: (2024)
by: Sooraj Kumar Maurya
Published: (2024)
The SMeL Test: A simple benchmark for media literacy in language models
by: Ahdritz, Gustaf, et al.
Published: (2025)
by: Ahdritz, Gustaf, et al.
Published: (2025)
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
by: Barboule, Camille, et al.
Published: (2024)
by: Barboule, Camille, et al.
Published: (2024)
Towards understanding evolution of science through language model series
by: Dong, Junjie, et al.
Published: (2024)
by: Dong, Junjie, et al.
Published: (2024)
Creativity Benchmark: A benchmark for marketing creativity for large language models
by: Bhat, Ninad, et al.
Published: (2025)
by: Bhat, Ninad, et al.
Published: (2025)
Lost without translation -- Can transformer (language models) understand mood states?
by: Shivaprakash, Prakrithi, et al.
Published: (2025)
by: Shivaprakash, Prakrithi, et al.
Published: (2025)
Can large language models understand uncommon meanings of common words?
by: Wu, Jinyang, et al.
Published: (2024)
by: Wu, Jinyang, et al.
Published: (2024)
Re-evaluating Theory of Mind evaluation in large language models
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
CMMLU: Measuring massive multitask language understanding in Chinese
by: Li, Haonan, et al.
Published: (2023)
by: Li, Haonan, et al.
Published: (2023)
NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
by: Jelodar, Hamed, et al.
Published: (2025)
by: Jelodar, Hamed, et al.
Published: (2025)
Logical forms complement probability in understanding language model (and human) performance
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Realtime, multimodal invasive ventilation risk monitoring using language models and BoXHED
by: Pakbin, Arash, et al.
Published: (2024)
by: Pakbin, Arash, et al.
Published: (2024)
A dataset and benchmark for hospital course summarization with adapted large language models
by: Aali, Asad, et al.
Published: (2024)
by: Aali, Asad, et al.
Published: (2024)
Clinical named entity recognition in the Portuguese language: a benchmark of modern BERT models and LLMs
by: de Almeida, Vinicius Anjos, et al.
Published: (2026)
by: de Almeida, Vinicius Anjos, et al.
Published: (2026)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
Similar Items
-
Bidirectional Emergent Language in Situated Environments
by: Wolff, Cornelius, et al.
Published: (2024) -
Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models
by: Ballout, Mohamad, et al.
Published: (2025) -
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
by: Ohmer, Xenia, et al.
Published: (2024) -
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
by: Mayer, Julius, et al.
Published: (2025) -
On the Relationship between Skill Neurons and Robustness in Prompt Tuning
by: Ackermann, Leon, et al.
Published: (2023)