Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Junjie, Yu, Mo, Liu, Lemao, Yeung, Dit-Yan, Zhou, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
by: Chung, Tsz Ting, et al.
Published: (2025)
by: Chung, Tsz Ting, et al.
Published: (2025)
The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding
by: Yu, Mo, et al.
Published: (2025)
by: Yu, Mo, et al.
Published: (2025)
Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability
by: Chung, Tsz Ting, et al.
Published: (2024)
by: Chung, Tsz Ting, et al.
Published: (2024)
Rethinking Targeted Adversarial Attacks For Neural Machine Translation
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
SitEmb-v1.5: Improved Context-Aware Dense Retrieval for Semantic Association and Long Story Comprehension
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
Towards Threshold-Free KV Cache Pruning
by: Ni, Xuanfan, et al.
Published: (2025)
by: Ni, Xuanfan, et al.
Published: (2025)
ARC-TGI: Human-Validated Task Generators with Reasoning Chain Templates for ARC-AGI
by: Lehmann, Jens, et al.
Published: (2026)
by: Lehmann, Jens, et al.
Published: (2026)
Graph-Based Exploration for ARC-AGI-3 Interactive Reasoning Tasks
by: Rudakov, Evgenii, et al.
Published: (2025)
by: Rudakov, Evgenii, et al.
Published: (2025)
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
by: Foundation, ARC Prize
Published: (2026)
by: Foundation, ARC Prize
Published: (2026)
G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
by: Tong, Tony Cheng, et al.
Published: (2024)
by: Tong, Tony Cheng, et al.
Published: (2024)
Large Language Models Can Self-Improve in Long-context Reasoning
by: Li, Siheng, et al.
Published: (2024)
by: Li, Siheng, et al.
Published: (2024)
Towards Privacy-Preserving Machine Translation at the Inference Stage: A New Task and Benchmark
by: Shao, Wei, et al.
Published: (2026)
by: Shao, Wei, et al.
Published: (2026)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
by: Lei, Chao, et al.
Published: (2025)
by: Lei, Chao, et al.
Published: (2025)
InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game Context
by: Liu, Ziyi, et al.
Published: (2024)
by: Liu, Ziyi, et al.
Published: (2024)
Disperse-Then-Merge: Pushing the Limits of Instruction Tuning via Alignment Tax Reduction
by: Fu, Tingchen, et al.
Published: (2024)
by: Fu, Tingchen, et al.
Published: (2024)
LLM-ARC: Enhancing LLMs with an Automated Reasoning Critic
by: Kalyanpur, Aditya, et al.
Published: (2024)
by: Kalyanpur, Aditya, et al.
Published: (2024)
Compositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
by: Mondorf, Philipp, et al.
Published: (2025)
by: Mondorf, Philipp, et al.
Published: (2025)
Impact of Noise on LLM-Models Performance in Abstraction and Reasoning Corpus (ARC) Tasks with Model Temperature Considerations
by: Khandalkar, Nikhil, et al.
Published: (2025)
by: Khandalkar, Nikhil, et al.
Published: (2025)
ARC Prize 2025: Technical Report
by: Chollet, François, et al.
Published: (2026)
by: Chollet, François, et al.
Published: (2026)
ARC Prize 2024: Technical Report
by: Chollet, Francois, et al.
Published: (2024)
by: Chollet, Francois, et al.
Published: (2024)
ARC-RL: A Reinforcement Learning Playground Inspired by ARC Raiders
by: Romeo, Carlo, et al.
Published: (2026)
by: Romeo, Carlo, et al.
Published: (2026)
Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective
by: Franzen, Daniel, et al.
Published: (2025)
by: Franzen, Daniel, et al.
Published: (2025)
Can Complexity and Uncomputability Explain Intelligence? SuperARC: A Test for Artificial Super Intelligence Based on Recursive Compression
by: Hernández-Espinosa, Alberto, et al.
Published: (2025)
by: Hernández-Espinosa, Alberto, et al.
Published: (2025)
Procedural Refinement by LLM-driven Algorithmic Debugging for ARC-AGI-2
by: Qiu, Yu-Ning, et al.
Published: (2026)
by: Qiu, Yu-Ning, et al.
Published: (2026)
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models
by: Zhang, Jiaming, et al.
Published: (2024)
by: Zhang, Jiaming, et al.
Published: (2024)
MF-CLIP: Leveraging CLIP as Surrogate Models for No-box Adversarial Attacks
by: Zhang, Jiaming, et al.
Published: (2023)
by: Zhang, Jiaming, et al.
Published: (2023)
MagicDrive: Street View Generation with Diverse 3D Geometry Control
by: Gao, Ruiyuan, et al.
Published: (2023)
by: Gao, Ruiyuan, et al.
Published: (2023)
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs
by: Qi, Yunjin, et al.
Published: (2026)
by: Qi, Yunjin, et al.
Published: (2026)
Multi-Perspective Transformers in ARC-AGI-2 Challenge
by: Talley, Caleb, et al.
Published: (2026)
by: Talley, Caleb, et al.
Published: (2026)
A Survey on the Honesty of Large Language Models
by: Li, Siheng, et al.
Published: (2024)
by: Li, Siheng, et al.
Published: (2024)
Can LLMs Understand Social Norms in Autonomous Driving Games?
by: Wang, Boxuan, et al.
Published: (2024)
by: Wang, Boxuan, et al.
Published: (2024)
A Decomposition Perspective to Long-context Reasoning for LLMs
by: Xiao, Yanling, et al.
Published: (2026)
by: Xiao, Yanling, et al.
Published: (2026)
The role of positional encodings in the ARC benchmark
by: Costa, Guilherme H. Bandeira, et al.
Published: (2025)
by: Costa, Guilherme H. Bandeira, et al.
Published: (2025)
NSA: Neuro-symbolic ARC Challenge
by: Batorski, Paweł, et al.
Published: (2025)
by: Batorski, Paweł, et al.
Published: (2025)
Truly Assessing Fluid Intelligence of Large Language Models through Dynamic Reasoning Evaluation
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
ARC-AGI-2 Technical Report
by: de Oliveira, Wallyson Lemes, et al.
Published: (2026)
by: de Oliveira, Wallyson Lemes, et al.
Published: (2026)
TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
by: Li, Pengxiang, et al.
Published: (2023)
by: Li, Pengxiang, et al.
Published: (2023)
Aligning Paralinguistic Understanding and Generation in Speech LLMs via Multi-Task Reinforcement Learning
by: Chen, Jingxiang, et al.
Published: (2026)
by: Chen, Jingxiang, et al.
Published: (2026)
Similar Items
-
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
by: Chung, Tsz Ting, et al.
Published: (2025) -
The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding
by: Yu, Mo, et al.
Published: (2025) -
Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability
by: Chung, Tsz Ting, et al.
Published: (2024) -
Rethinking Targeted Adversarial Attacks For Neural Machine Translation
by: Wu, Junjie, et al.
Published: (2024) -
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
by: Wu, Junjie, et al.
Published: (2024)