Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?
Fuente:
arXiv
Saved in:
| Main Authors: | Triebel, Maximilian, Menner, Marco, Helfenstein, Dominik |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Online library learning in human visual puzzle solving
by: Zhao, Pinzhe, et al.
Published: (2026)
by: Zhao, Pinzhe, et al.
Published: (2026)
Language models show human-like content effects on reasoning tasks
by: Dasgupta, Ishita, et al.
Published: (2022)
by: Dasgupta, Ishita, et al.
Published: (2022)
RECALL: Rehearsal-free Continual Learning for Object Classification
by: Knauer, Markus, et al.
Published: (2022)
by: Knauer, Markus, et al.
Published: (2022)
Learning Expressive Priors for Generalization and Uncertainty Estimation in Neural Networks
by: Schnaus, Dominik, et al.
Published: (2023)
by: Schnaus, Dominik, et al.
Published: (2023)
Large Language Models show both individual and collective creativity comparable to humans
by: Sun, Luning, et al.
Published: (2024)
by: Sun, Luning, et al.
Published: (2024)
Procedurally generating rules to adapt difficulty for narrative puzzle games
by: Volden, Thomas, et al.
Published: (2023)
by: Volden, Thomas, et al.
Published: (2023)
Towards Explaining Uncertainty Estimates in Point Cloud Registration
by: Qin, Ziyuan, et al.
Published: (2024)
by: Qin, Ziyuan, et al.
Published: (2024)
Evaluation of LLMs for mathematical problem solving
by: Wang, Ruonan, et al.
Published: (2025)
by: Wang, Ruonan, et al.
Published: (2025)
LLMs model how humans induce logically structured rules
by: Loo, Alyssa, et al.
Published: (2025)
by: Loo, Alyssa, et al.
Published: (2025)
Can Large Language Models generalize analogy solving like children can?
by: Stevenson, Claire E., et al.
Published: (2024)
by: Stevenson, Claire E., et al.
Published: (2024)
AlphaBeta is not as good as you think: a simple class of synthetic games for a better analysis of deterministic game-solving algorithms
by: Boige, Raphaël, et al.
Published: (2025)
by: Boige, Raphaël, et al.
Published: (2025)
Tracing the ongoing emergence of human-like reasoning in Large Language Models
by: Morosi, Paolo, et al.
Published: (2026)
by: Morosi, Paolo, et al.
Published: (2026)
The promise and limits of LLMs in constructing proofs and hints for logic problems in intelligent tutoring systems
by: Tithi, Sutapa Dey, et al.
Published: (2025)
by: Tithi, Sutapa Dey, et al.
Published: (2025)
Can generative AI and ChatGPT outperform humans on cognitive-demanding problem-solving tasks in science?
by: Zhai, Xiaoming, et al.
Published: (2024)
by: Zhai, Xiaoming, et al.
Published: (2024)
The receptron is a nonlinear threshold logic gate with intrinsic multi-dimensional selective capabilities for analog inputs
by: Paroli, B., et al.
Published: (2025)
by: Paroli, B., et al.
Published: (2025)
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
by: Hsieh, Wen-Han, et al.
Published: (2025)
by: Hsieh, Wen-Han, et al.
Published: (2025)
Logical recognition method for solving the problem of identification in the Internet of Things
by: Saymanov, Islambek
Published: (2024)
by: Saymanov, Islambek
Published: (2024)
A method for quantifying the generalization capabilities of generative models for solving Ising models
by: Ma, Qunlong, et al.
Published: (2024)
by: Ma, Qunlong, et al.
Published: (2024)
Performance Review on LLM for solving leetcode problems
by: Wang, Lun, et al.
Published: (2025)
by: Wang, Lun, et al.
Published: (2025)
Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?
by: Lu, Yi-Long, et al.
Published: (2025)
by: Lu, Yi-Long, et al.
Published: (2025)
Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?
by: Yang, Ruixin, et al.
Published: (2026)
by: Yang, Ruixin, et al.
Published: (2026)
CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
by: Li, Peiyu, et al.
Published: (2025)
by: Li, Peiyu, et al.
Published: (2025)
Assessing SPARQL capabilities of Large Language Models
by: Meyer, Lars-Peter, et al.
Published: (2024)
by: Meyer, Lars-Peter, et al.
Published: (2024)
Evaluating List Construction and Temporal Understanding capabilities of Large Language Models
by: Dumitru, Alexandru, et al.
Published: (2025)
by: Dumitru, Alexandru, et al.
Published: (2025)
Large language models show fragile cognitive reasoning about human emotions
by: Bhattacharyya, Sree, et al.
Published: (2025)
by: Bhattacharyya, Sree, et al.
Published: (2025)
Among Them: A game-based framework for assessing persuasion capabilities of LLMs
by: Idziejczak, Mateusz, et al.
Published: (2025)
by: Idziejczak, Mateusz, et al.
Published: (2025)
Is AI currently capable of identifying wild oysters? A comparison of human annotators against the AI model, ODYSSEE
by: Campbell, Brendan, et al.
Published: (2025)
by: Campbell, Brendan, et al.
Published: (2025)
Do All Individual Layers Help? An Empirical Study of Task-Interfering Layers in Vision-Language Models
by: Liu, Zhiming, et al.
Published: (2026)
by: Liu, Zhiming, et al.
Published: (2026)
Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark
by: Mushkani, Rashid
Published: (2025)
by: Mushkani, Rashid
Published: (2025)
Which symbol grounding problem should we try to solve?
by: Müller, Vincent C.
Published: (2025)
by: Müller, Vincent C.
Published: (2025)
Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
by: Ernhofer, Benjamin Raphael, et al.
Published: (2025)
by: Ernhofer, Benjamin Raphael, et al.
Published: (2025)
VideoGameBench: Can Vision-Language Models complete popular video games?
by: Zhang, Alex L., et al.
Published: (2025)
by: Zhang, Alex L., et al.
Published: (2025)
YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures
by: Schwarz, Dominik
Published: (2025)
by: Schwarz, Dominik
Published: (2025)
Why Do Vision Language Models Struggle To Recognize Human Emotions?
by: Agarwal, Madhav, et al.
Published: (2026)
by: Agarwal, Madhav, et al.
Published: (2026)
Do Pre-trained Vision-Language Models Encode Object States?
by: Newman, Kaleb, et al.
Published: (2024)
by: Newman, Kaleb, et al.
Published: (2024)
RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via Romanization
by: Husain, Jaavid Aktar, et al.
Published: (2024)
by: Husain, Jaavid Aktar, et al.
Published: (2024)
Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement?
by: Ilić, David, et al.
Published: (2023)
by: Ilić, David, et al.
Published: (2023)
Conditional score-based diffusion models for solving inverse problems in mechanics
by: Dasgupta, Agnimitra, et al.
Published: (2024)
by: Dasgupta, Agnimitra, et al.
Published: (2024)
RE-tune: Incremental Fine Tuning of Biomedical Vision-Language Models for Multi-label Chest X-ray Classification
by: Mistretta, Marco, et al.
Published: (2024)
by: Mistretta, Marco, et al.
Published: (2024)
Similar Items
-
Online library learning in human visual puzzle solving
by: Zhao, Pinzhe, et al.
Published: (2026) -
Language models show human-like content effects on reasoning tasks
by: Dasgupta, Ishita, et al.
Published: (2022) -
RECALL: Rehearsal-free Continual Learning for Object Classification
by: Knauer, Markus, et al.
Published: (2022) -
Learning Expressive Priors for Generalization and Uncertainty Estimation in Neural Networks
by: Schnaus, Dominik, et al.
Published: (2023) -
Large Language Models show both individual and collective creativity comparable to humans
by: Sun, Luning, et al.
Published: (2024)