GraphARC: A Comprehensive Benchmark for Graph-Based Abstract Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peltonen, Saku, Rønberg, August Bøgh, Plesner, Andreas, Wattenhofer, Roger
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911731622608896
author Peltonen, Saku
Rønberg, August Bøgh
Plesner, Andreas
Wattenhofer, Roger
author_facet Peltonen, Saku
Rønberg, August Bøgh
Plesner, Andreas
Wattenhofer, Roger
contents Relational reasoning lies at the heart of intelligence, but existing benchmarks are typically confined to formats such as grids or text. We introduce GraphARC, a benchmark for abstract reasoning on graph-structured data. GraphARC generalizes the few-shot transformation learning paradigm of the Abstraction and Reasoning Corpus (ARC). Each task requires inferring a transformation rule from a few input-output pairs and applying it to a new test graph, covering local, global, and hierarchical graph transformations. Unlike grid-based ARC, GraphARC instances can be generated at scale across diverse graph families and sizes, enabling systematic evaluation of generalization abilities. We evaluate state-of-the-art language models on GraphARC and observe clear limitations. Models can answer questions about graph properties but often fail to solve the full graph transformation task, revealing a comprehension-execution gap. Performance further degrades on larger instances, exposing scaling barriers. More broadly, by combining aspects of node classification, link prediction, and graph generation within a single framework, GraphARC provides a promising testbed for future graph foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31031
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GraphARC: A Comprehensive Benchmark for Graph-Based Abstract Reasoning
Peltonen, Saku
Rønberg, August Bøgh
Plesner, Andreas
Wattenhofer, Roger
Artificial Intelligence
Relational reasoning lies at the heart of intelligence, but existing benchmarks are typically confined to formats such as grids or text. We introduce GraphARC, a benchmark for abstract reasoning on graph-structured data. GraphARC generalizes the few-shot transformation learning paradigm of the Abstraction and Reasoning Corpus (ARC). Each task requires inferring a transformation rule from a few input-output pairs and applying it to a new test graph, covering local, global, and hierarchical graph transformations. Unlike grid-based ARC, GraphARC instances can be generated at scale across diverse graph families and sizes, enabling systematic evaluation of generalization abilities. We evaluate state-of-the-art language models on GraphARC and observe clear limitations. Models can answer questions about graph properties but often fail to solve the full graph transformation task, revealing a comprehension-execution gap. Performance further degrades on larger instances, exposing scaling barriers. More broadly, by combining aspects of node classification, link prediction, and graph generation within a single framework, GraphARC provides a promising testbed for future graph foundation models.
title GraphARC: A Comprehensive Benchmark for Graph-Based Abstract Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2605.31031