Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Franzen, Daniel, Disselhoff, Jan, Hartmann, David
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909645626408960
author Franzen, Daniel
Disselhoff, Jan
Hartmann, David
author_facet Franzen, Daniel
Disselhoff, Jan
Hartmann, David
contents The Abstraction and Reasoning Corpus (ARC-AGI) poses a significant challenge for large language models (LLMs), exposing limitations in their abstract reasoning abilities. In this work, we leverage task-specific data augmentations throughout the training, generation, and scoring phases, and employ a depth-first search algorithm to generate diverse, high-probability candidate solutions. Furthermore, we utilize the LLM not only as a generator but also as a scorer, using its output probabilities to select the most promising solutions. Our method achieves a score of 71.6% (286.5/400 solved tasks) on the public ARC-AGI evaluation set, demonstrating state-of-the-art performance among publicly available approaches. While concurrent closed-source work has reported higher scores, our method distinguishes itself through its transparency, reproducibility, and remarkably low inference cost, averaging only around 2ct per task on readily available hardware (we assume a price of 36ct/hour for a Nvidia 4090 GPU).
format Preprint
id arxiv_https___arxiv_org_abs_2505_07859
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective
Franzen, Daniel
Disselhoff, Jan
Hartmann, David
Computation and Language
Artificial Intelligence
Machine Learning
The Abstraction and Reasoning Corpus (ARC-AGI) poses a significant challenge for large language models (LLMs), exposing limitations in their abstract reasoning abilities. In this work, we leverage task-specific data augmentations throughout the training, generation, and scoring phases, and employ a depth-first search algorithm to generate diverse, high-probability candidate solutions. Furthermore, we utilize the LLM not only as a generator but also as a scorer, using its output probabilities to select the most promising solutions. Our method achieves a score of 71.6% (286.5/400 solved tasks) on the public ARC-AGI evaluation set, demonstrating state-of-the-art performance among publicly available approaches. While concurrent closed-source work has reported higher scores, our method distinguishes itself through its transparency, reproducibility, and remarkably low inference cost, averaging only around 2ct per task on readily available hardware (we assume a price of 36ct/hour for a Nvidia 4090 GPU).
title Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.07859