COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Taoudi-Benchekroun, Yassine, Troyan, Klim, Sager, Pascal, Gerber, Stefan, Tuggener, Lukas, Grewe, Benjamin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910025243426816
author Taoudi-Benchekroun, Yassine
Troyan, Klim
Sager, Pascal
Gerber, Stefan
Tuggener, Lukas
Grewe, Benjamin
author_facet Taoudi-Benchekroun, Yassine
Troyan, Klim
Sager, Pascal
Gerber, Stefan
Tuggener, Lukas
Grewe, Benjamin
contents The ability to compose learned concepts and apply them in novel settings is key to human intelligence, but remains a persistent limitation in state-of-the-art machine learning models. To address this issue, we introduce COGITAO, a modular and extensible data generation framework and benchmark designed to systematically study compositionality and generalization in visual domains. Drawing inspiration from ARC-AGI's problem-setting, COGITAO constructs rule-based tasks which apply a set of transformations to objects in grid-like environments. It supports composition, at adjustable depth, over a set of 28 interoperable transformations, along with extensive control over grid parametrization and object properties. This flexibility enables the creation of millions of unique task rules -- surpassing concurrent datasets by several orders of magnitude -- across a wide range of difficulties, while allowing virtually unlimited sample generation per rule. We provide baseline experiments using state-of-the-art vision models, highlighting their consistent failures to generalize to novel combinations of familiar elements, despite strong in-domain performance. COGITAO is fully open-sourced, including all code and datasets, to support continued research in this field.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05249
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
Taoudi-Benchekroun, Yassine
Troyan, Klim
Sager, Pascal
Gerber, Stefan
Tuggener, Lukas
Grewe, Benjamin
Computer Vision and Pattern Recognition
Artificial Intelligence
The ability to compose learned concepts and apply them in novel settings is key to human intelligence, but remains a persistent limitation in state-of-the-art machine learning models. To address this issue, we introduce COGITAO, a modular and extensible data generation framework and benchmark designed to systematically study compositionality and generalization in visual domains. Drawing inspiration from ARC-AGI's problem-setting, COGITAO constructs rule-based tasks which apply a set of transformations to objects in grid-like environments. It supports composition, at adjustable depth, over a set of 28 interoperable transformations, along with extensive control over grid parametrization and object properties. This flexibility enables the creation of millions of unique task rules -- surpassing concurrent datasets by several orders of magnitude -- across a wide range of difficulties, while allowing virtually unlimited sample generation per rule. We provide baseline experiments using state-of-the-art vision models, highlighting their consistent failures to generalize to novel combinations of familiar elements, despite strong in-domain performance. COGITAO is fully open-sourced, including all code and datasets, to support continued research in this field.
title COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.05249