TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Daixian, Kuang, Jiayi, Li, Yinghui, Li, Yangning, Yin, Di, Cao, Haoyu, Sun, Xing, Shen, Ying, Zheng, Hai-Tao, Lin, Liang, Yu, Philip S. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
by: Li, Yinghui, et al.
Published: (2026)
by: Li, Yinghui, et al.
Published: (2026)
Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
by: Kuang, Jiayi, et al.
Published: (2025)
by: Kuang, Jiayi, et al.
Published: (2025)
EvoConfig: Self-Evolving Multi-Agent Systems for Efficient Autonomous Environment Configuration
by: Guo, Xinshuai, et al.
Published: (2026)
by: Guo, Xinshuai, et al.
Published: (2026)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
by: Kuang, Jiayi, et al.
Published: (2025)
by: Kuang, Jiayi, et al.
Published: (2025)
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
by: Li, Yinghui, et al.
Published: (2025)
by: Li, Yinghui, et al.
Published: (2025)
Refine Knowledge of Large Language Models via Adaptive Contrastive Learning
by: Li, Yinghui, et al.
Published: (2025)
by: Li, Yinghui, et al.
Published: (2025)
Mitigating Catastrophic Forgetting in Multi-domain Chinese Spelling Correction by Multi-stage Knowledge Transfer Framework
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
Bidirectional End-to-End Learning of Retriever-Reader Paradigm for Entity Linking
by: Li, Yinghui, et al.
Published: (2023)
by: Li, Yinghui, et al.
Published: (2023)
MDIT: A Model-free Data Interpolation Method for Diverse Instruction Tuning
by: Li, Yangning, et al.
Published: (2025)
by: Li, Yangning, et al.
Published: (2025)
From Retrieval to Generation: Efficient and Effective Entity Set Expansion
by: Huang, Shulin, et al.
Published: (2023)
by: Huang, Shulin, et al.
Published: (2023)
Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion
by: Wang, Hebin, et al.
Published: (2024)
by: Wang, Hebin, et al.
Published: (2024)
Correct Like Humans: Progressive Learning Framework for Chinese Text Error Correction
by: Li, Yinghui, et al.
Published: (2023)
by: Li, Yinghui, et al.
Published: (2023)
Automatic Context Pattern Generation for Entity Set Expansion
by: Li, Yinghui, et al.
Published: (2022)
by: Li, Yinghui, et al.
Published: (2022)
Embracing Ambiguity: Improving Similarity-oriented Tasks with Contextual Synonym Knowledge
by: Li, Yangning, et al.
Published: (2022)
by: Li, Yangning, et al.
Published: (2022)
When LLMs Meet Cunning Texts: A Fallacy Understanding Benchmark for Large Language Models
by: Li, Yinghui, et al.
Published: (2024)
by: Li, Yinghui, et al.
Published: (2024)
Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
by: Li, Yangning, et al.
Published: (2024)
by: Li, Yangning, et al.
Published: (2024)
AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
by: Li, Yangning, et al.
Published: (2025)
by: Li, Yangning, et al.
Published: (2025)
Rethinking the Roles of Large Language Models in Chinese Grammatical Error Correction
by: Li, Yinghui, et al.
Published: (2024)
by: Li, Yinghui, et al.
Published: (2024)
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles
by: Huang, Shulin, et al.
Published: (2023)
by: Huang, Shulin, et al.
Published: (2023)
UltraWiki: Ultra-fine-grained Entity Set Expansion with Negative Seed Entities
by: Li, Yangning, et al.
Published: (2024)
by: Li, Yangning, et al.
Published: (2024)
On the (In)Effectiveness of Large Language Models for Chinese Text Correction
by: Li, Yinghui, et al.
Published: (2023)
by: Li, Yinghui, et al.
Published: (2023)
DAST: Context-Aware Compression in LLMs via Dynamic Allocation of Soft Tokens
by: Chen, Shaoshen, et al.
Published: (2025)
by: Chen, Shaoshen, et al.
Published: (2025)
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
by: Li, Hengzhi, et al.
Published: (2025)
by: Li, Hengzhi, et al.
Published: (2025)
Teaching According to Talents! Instruction Tuning LLMs with Competence-Aware Curriculum Learning
by: Li, Yangning, et al.
Published: (2025)
by: Li, Yangning, et al.
Published: (2025)
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
by: Kuang, Jiayi, et al.
Published: (2024)
by: Kuang, Jiayi, et al.
Published: (2024)
ProductAgent: Benchmarking Conversational Product Search Agent with Asking Clarification Questions
by: Ye, Jingheng, et al.
Published: (2024)
by: Ye, Jingheng, et al.
Published: (2024)
Tangram: Benchmark for Evaluating Geometric Element Recognition in Large Multimodal Models
by: Zhang, Chao, et al.
Published: (2024)
by: Zhang, Chao, et al.
Published: (2024)
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
by: Song, Yueqi, et al.
Published: (2025)
by: Song, Yueqi, et al.
Published: (2025)
From Token to Line: Enhancing Code Generation with a Long-Term Perspective
by: Lu, Tingwei, et al.
Published: (2025)
by: Lu, Tingwei, et al.
Published: (2025)
Loss-Aware Curriculum Learning for Chinese Grammatical Error Correction
by: Zhang, Ding, et al.
Published: (2024)
by: Zhang, Ding, et al.
Published: (2024)
RAISE: Reinforced Adaptive Instruction Selection For Large Language Models
by: Lv, Qingsong, et al.
Published: (2025)
by: Lv, Qingsong, et al.
Published: (2025)
SceneGram: Conceptualizing and Describing Tangrams in Scene Context
by: Junker, Simeon, et al.
Published: (2025)
by: Junker, Simeon, et al.
Published: (2025)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
by: Peng, Haosong, et al.
Published: (2024)
by: Peng, Haosong, et al.
Published: (2024)
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
by: Zhan, Shaoxiong, et al.
Published: (2026)
by: Zhan, Shaoxiong, et al.
Published: (2026)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
by: Long, Yitao, et al.
Published: (2025)
by: Long, Yitao, et al.
Published: (2025)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
by: Zhan, Xiaoyu, et al.
Published: (2025)
by: Zhan, Xiaoyu, et al.
Published: (2025)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
by: Ghosal, Deepanway, et al.
Published: (2024)
by: Ghosal, Deepanway, et al.
Published: (2024)
Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
by: Li, Yangning, et al.
Published: (2025)
by: Li, Yangning, et al.
Published: (2025)
ARL-Tangram: Unleash the Resource Efficiency in Agentic Reinforcement Learning
by: Xiao, Bangjun, et al.
Published: (2026)
by: Xiao, Bangjun, et al.
Published: (2026)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
by: Zhu, Wenbin, et al.
Published: (2025)
by: Zhu, Wenbin, et al.
Published: (2025)
Similar Items
-
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
by: Li, Yinghui, et al.
Published: (2026) -
Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
by: Kuang, Jiayi, et al.
Published: (2025) -
EvoConfig: Self-Evolving Multi-Agent Systems for Efficient Autonomous Environment Configuration
by: Guo, Xinshuai, et al.
Published: (2026) -
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
by: Kuang, Jiayi, et al.
Published: (2025) -
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
by: Li, Yinghui, et al.
Published: (2025)