GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Tang, Zhisheng, Kejriwal, Mayank |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
di: Tang, Zhisheng, et al.
Pubblicazione: (2024)
di: Tang, Zhisheng, et al.
Pubblicazione: (2024)
An Evaluation of Estimative Uncertainty in Large Language Models
di: Tang, Zhisheng, et al.
Pubblicazione: (2024)
di: Tang, Zhisheng, et al.
Pubblicazione: (2024)
Is persona enough for personality? Using ChatGPT to reconstruct an agent's latent personality from simple descriptions
di: Ji, Yongyi, et al.
Pubblicazione: (2024)
di: Ji, Yongyi, et al.
Pubblicazione: (2024)
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
di: Shen, Ke, et al.
Pubblicazione: (2024)
di: Shen, Ke, et al.
Pubblicazione: (2024)
A Compound AI Agent for Conversational Grant Discovery
di: Tang, Zhisheng, et al.
Pubblicazione: (2026)
di: Tang, Zhisheng, et al.
Pubblicazione: (2026)
Navigating Semantic Relations: Challenges for Language Models in Abstract Common-Sense Reasoning
di: Gawin, Cole, et al.
Pubblicazione: (2025)
di: Gawin, Cole, et al.
Pubblicazione: (2025)
Code-Driven Planning in Grid Worlds with Large Language Models
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2025)
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2025)
SelECT-SQL: Self-correcting ensemble Chain-of-Thought for Text-to-SQL
di: Shen, Ke, et al.
Pubblicazione: (2024)
di: Shen, Ke, et al.
Pubblicazione: (2024)
HALO: An Ontology for Representing and Categorizing Hallucinations in Large Language Models
di: Nananukul, Navapat, et al.
Pubblicazione: (2023)
di: Nananukul, Navapat, et al.
Pubblicazione: (2023)
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2026)
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2026)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2026)
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2026)
LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning
di: Junias, Obed, et al.
Pubblicazione: (2026)
di: Junias, Obed, et al.
Pubblicazione: (2026)
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
di: You, Wangjie, et al.
Pubblicazione: (2025)
di: You, Wangjie, et al.
Pubblicazione: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
di: Palta, Shramay, et al.
Pubblicazione: (2024)
di: Palta, Shramay, et al.
Pubblicazione: (2024)
BrainBench: Exposing the Commonsense Reasoning Gap in Large Language Models
di: Tang, Yuzhe
Pubblicazione: (2026)
di: Tang, Yuzhe
Pubblicazione: (2026)
Com$^2$: A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models
di: Xiong, Kai, et al.
Pubblicazione: (2025)
di: Xiong, Kai, et al.
Pubblicazione: (2025)
Detecting Emotional Incongruity of Sarcasm by Commonsense Reasoning
di: Qiu, Ziqi, et al.
Pubblicazione: (2024)
di: Qiu, Ziqi, et al.
Pubblicazione: (2024)
ConstraintChecker: A Plugin for Large Language Models to Reason on Commonsense Knowledge Bases
di: Do, Quyet V., et al.
Pubblicazione: (2024)
di: Do, Quyet V., et al.
Pubblicazione: (2024)
Zero-shot Commonsense Reasoning over Machine Imagination
di: Park, Hyuntae, et al.
Pubblicazione: (2024)
di: Park, Hyuntae, et al.
Pubblicazione: (2024)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
di: Yadav, Ankit, et al.
Pubblicazione: (2024)
di: Yadav, Ankit, et al.
Pubblicazione: (2024)
Meta-RTL: Reinforcement-Based Meta-Transfer Learning for Low-Resource Commonsense Reasoning
di: Fu, Yu, et al.
Pubblicazione: (2024)
di: Fu, Yu, et al.
Pubblicazione: (2024)
Complex Reasoning over Logical Queries on Commonsense Knowledge Graphs
di: Fang, Tianqing, et al.
Pubblicazione: (2024)
di: Fang, Tianqing, et al.
Pubblicazione: (2024)
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
di: Cheng, Qi, et al.
Pubblicazione: (2024)
di: Cheng, Qi, et al.
Pubblicazione: (2024)
Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not
di: Karakaş, Sercan
Pubblicazione: (2026)
di: Karakaş, Sercan
Pubblicazione: (2026)
Towards Quantifying Commonsense Reasoning with Mechanistic Insights
di: Joshi, Abhinav, et al.
Pubblicazione: (2025)
di: Joshi, Abhinav, et al.
Pubblicazione: (2025)
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
di: Li, Fangjun, et al.
Pubblicazione: (2024)
di: Li, Fangjun, et al.
Pubblicazione: (2024)
Zero-Shot Commonsense Validation and Reasoning with Large Language Models: An Evaluation on SemEval-2020 Task 4 Dataset
di: Alfugaha, Rawand, et al.
Pubblicazione: (2025)
di: Alfugaha, Rawand, et al.
Pubblicazione: (2025)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
di: Tang, Weizhi, et al.
Pubblicazione: (2024)
di: Tang, Weizhi, et al.
Pubblicazione: (2024)
Commonsense Knowledge Editing Based on Free-Text in LLMs
di: Huang, Xiusheng, et al.
Pubblicazione: (2024)
di: Huang, Xiusheng, et al.
Pubblicazione: (2024)
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
di: Li, Jiachun, et al.
Pubblicazione: (2024)
di: Li, Jiachun, et al.
Pubblicazione: (2024)
MIDGARD: Self-Consistency Using Minimum Description Length for Structured Commonsense Reasoning
di: Nair, Inderjeet, et al.
Pubblicazione: (2024)
di: Nair, Inderjeet, et al.
Pubblicazione: (2024)
LINKED: Eliciting, Filtering and Integrating Knowledge in Large Language Model for Commonsense Reasoning
di: Li, Jiachun, et al.
Pubblicazione: (2024)
di: Li, Jiachun, et al.
Pubblicazione: (2024)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
di: Almheiri, Saeed, et al.
Pubblicazione: (2025)
di: Almheiri, Saeed, et al.
Pubblicazione: (2025)
An Analysis of Artificial Intelligence Adoption in NIH-Funded Research
di: Nananukul, Navapat, et al.
Pubblicazione: (2026)
di: Nananukul, Navapat, et al.
Pubblicazione: (2026)
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
di: Ye, Hengwei, et al.
Pubblicazione: (2026)
di: Ye, Hengwei, et al.
Pubblicazione: (2026)
GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives
di: Prabhakaran, Vinodkumar, et al.
Pubblicazione: (2023)
di: Prabhakaran, Vinodkumar, et al.
Pubblicazione: (2023)
What Really is Commonsense Knowledge?
di: Do, Quyet V., et al.
Pubblicazione: (2024)
di: Do, Quyet V., et al.
Pubblicazione: (2024)
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
di: Chen, Kesheng, et al.
Pubblicazione: (2026)
di: Chen, Kesheng, et al.
Pubblicazione: (2026)
Evaluating Large Language Models for Financial Reasoning: A CFA-Based Benchmark Study
di: Yao, Xuan, et al.
Pubblicazione: (2025)
di: Yao, Xuan, et al.
Pubblicazione: (2025)
Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap
di: Srivastava, Saurabh, et al.
Pubblicazione: (2024)
di: Srivastava, Saurabh, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
di: Tang, Zhisheng, et al.
Pubblicazione: (2024) -
An Evaluation of Estimative Uncertainty in Large Language Models
di: Tang, Zhisheng, et al.
Pubblicazione: (2024) -
Is persona enough for personality? Using ChatGPT to reconstruct an agent's latent personality from simple descriptions
di: Ji, Yongyi, et al.
Pubblicazione: (2024) -
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
di: Shen, Ke, et al.
Pubblicazione: (2024) -
A Compound AI Agent for Conversational Grant Discovery
di: Tang, Zhisheng, et al.
Pubblicazione: (2026)