Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Heekyung, Ge, Jiaxin, Wu, Tsung-Han, Kang, Minwoo, Darrell, Trevor, Chan, David M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
by: Wu, Tsung-Han, et al.
Published: (2025)
by: Wu, Tsung-Han, et al.
Published: (2025)
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023)
by: Ge, Jiaxin, et al.
Published: (2023)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
by: Kamath, Amita, et al.
Published: (2026)
by: Kamath, Amita, et al.
Published: (2026)
Vision-Language Models Create Cross-Modal Task Representations
by: Luo, Grace, et al.
Published: (2024)
by: Luo, Grace, et al.
Published: (2024)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
by: Leng, Jixuan, et al.
Published: (2025)
by: Leng, Jixuan, et al.
Published: (2025)
GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles
by: Shan, Mengyi, et al.
Published: (2025)
by: Shan, Mengyi, et al.
Published: (2025)
Constantly Improving Image Models Need Constantly Improving Benchmarks
by: Ge, Jiaxin, et al.
Published: (2025)
by: Ge, Jiaxin, et al.
Published: (2025)
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
by: Liu, Daixian, et al.
Published: (2026)
by: Liu, Daixian, et al.
Published: (2026)
$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
by: Das, Trishanu, et al.
Published: (2025)
by: Das, Trishanu, et al.
Published: (2025)
Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles
by: Chen, Qi, et al.
Published: (2024)
by: Chen, Qi, et al.
Published: (2024)
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
by: Li, Hengzhi, et al.
Published: (2025)
by: Li, Hengzhi, et al.
Published: (2025)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
by: Ghosal, Deepanway, et al.
Published: (2024)
by: Ghosal, Deepanway, et al.
Published: (2024)
Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
by: Wu, Tsung-Han, et al.
Published: (2024)
by: Wu, Tsung-Han, et al.
Published: (2024)
Can Vision-Language Models Solve the Shell Game?
by: Liu, Tiedong, et al.
Published: (2026)
by: Liu, Tiedong, et al.
Published: (2026)
Pose Priors from Language Models
by: Subramanian, Sanjay, et al.
Published: (2024)
by: Subramanian, Sanjay, et al.
Published: (2024)
VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models
by: Ren, Yufan, et al.
Published: (2025)
by: Ren, Yufan, et al.
Published: (2025)
When Do We Not Need Larger Vision Models?
by: Shi, Baifeng, et al.
Published: (2024)
by: Shi, Baifeng, et al.
Published: (2024)
A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
by: Ye, Weixin, et al.
Published: (2026)
by: Ye, Weixin, et al.
Published: (2026)
Vision-Language Models Can't See the Obvious
by: Dahou, Yasser, et al.
Published: (2025)
by: Dahou, Yasser, et al.
Published: (2025)
Analyzing The Language of Visual Tokens
by: Chan, David M., et al.
Published: (2024)
by: Chan, David M., et al.
Published: (2024)
PuzzlePoles: Cylindrical Fiducial Markers Based on the PuzzleBoard Pattern
by: Zach, Juri, et al.
Published: (2025)
by: Zach, Juri, et al.
Published: (2025)
Solving Masked Jigsaw Puzzles with Diffusion Vision Transformers
by: Liu, Jinyang, et al.
Published: (2024)
by: Liu, Jinyang, et al.
Published: (2024)
Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
by: Wang, Zifu, et al.
Published: (2025)
by: Wang, Zifu, et al.
Published: (2025)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
by: Mitra, Chancharik, et al.
Published: (2025)
by: Mitra, Chancharik, et al.
Published: (2025)
Reasoning or Pattern Matching? Probing Large Vision-Language Models with Visual Puzzles
by: Lymperaiou, Maria, et al.
Published: (2026)
by: Lymperaiou, Maria, et al.
Published: (2026)
PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
by: Jeddi, Ahmadreza, et al.
Published: (2025)
by: Jeddi, Ahmadreza, et al.
Published: (2025)
Benchmarking Content-Based Puzzle Solvers on Corrupted Jigsaw Puzzles
by: Dirauf, Richard, et al.
Published: (2025)
by: Dirauf, Richard, et al.
Published: (2025)
The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles
by: Toh, Vernon Y. H., et al.
Published: (2025)
by: Toh, Vernon Y. H., et al.
Published: (2025)
AutoPresent: Designing Structured Visuals from Scratch
by: Ge, Jiaxin, et al.
Published: (2025)
by: Ge, Jiaxin, et al.
Published: (2025)
ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?
by: Ling, Zijian, et al.
Published: (2025)
by: Ling, Zijian, et al.
Published: (2025)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
by: Mitra, Chancharik, et al.
Published: (2024)
by: Mitra, Chancharik, et al.
Published: (2024)
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
by: Kang, Jialiang, et al.
Published: (2025)
by: Kang, Jialiang, et al.
Published: (2025)
Dual-Process Image Generation
by: Luo, Grace, et al.
Published: (2025)
by: Luo, Grace, et al.
Published: (2025)
Can Vision-Language Models Evaluate Handwritten Math?
by: Nath, Oikantik, et al.
Published: (2025)
by: Nath, Oikantik, et al.
Published: (2025)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation
by: Tang, Siao, et al.
Published: (2025)
by: Tang, Siao, et al.
Published: (2025)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
by: Chia, Yew Ken, et al.
Published: (2024)
by: Chia, Yew Ken, et al.
Published: (2024)
TULIP: Towards Unified Language-Image Pretraining
by: Tang, Zineng, et al.
Published: (2025)
by: Tang, Zineng, et al.
Published: (2025)
Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
by: Góral, Gracjan, et al.
Published: (2024)
by: Góral, Gracjan, et al.
Published: (2024)
Similar Items
-
Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
by: Wu, Tsung-Han, et al.
Published: (2025) -
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023) -
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
by: Kamath, Amita, et al.
Published: (2026) -
Vision-Language Models Create Cross-Modal Task Representations
by: Luo, Grace, et al.
Published: (2024) -
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
by: Leng, Jixuan, et al.
Published: (2025)