COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910025243426816 |
|---|---|
| author | Taoudi-Benchekroun, Yassine Troyan, Klim Sager, Pascal Gerber, Stefan Tuggener, Lukas Grewe, Benjamin |
| author_facet | Taoudi-Benchekroun, Yassine Troyan, Klim Sager, Pascal Gerber, Stefan Tuggener, Lukas Grewe, Benjamin |
| contents | The ability to compose learned concepts and apply them in novel settings is key to human intelligence, but remains a persistent limitation in state-of-the-art machine learning models. To address this issue, we introduce COGITAO, a modular and extensible data generation framework and benchmark designed to systematically study compositionality and generalization in visual domains. Drawing inspiration from ARC-AGI's problem-setting, COGITAO constructs rule-based tasks which apply a set of transformations to objects in grid-like environments. It supports composition, at adjustable depth, over a set of 28 interoperable transformations, along with extensive control over grid parametrization and object properties. This flexibility enables the creation of millions of unique task rules -- surpassing concurrent datasets by several orders of magnitude -- across a wide range of difficulties, while allowing virtually unlimited sample generation per rule. We provide baseline experiments using state-of-the-art vision models, highlighting their consistent failures to generalize to novel combinations of familiar elements, despite strong in-domain performance. COGITAO is fully open-sourced, including all code and datasets, to support continued research in this field. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_05249 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization Taoudi-Benchekroun, Yassine Troyan, Klim Sager, Pascal Gerber, Stefan Tuggener, Lukas Grewe, Benjamin Computer Vision and Pattern Recognition Artificial Intelligence The ability to compose learned concepts and apply them in novel settings is key to human intelligence, but remains a persistent limitation in state-of-the-art machine learning models. To address this issue, we introduce COGITAO, a modular and extensible data generation framework and benchmark designed to systematically study compositionality and generalization in visual domains. Drawing inspiration from ARC-AGI's problem-setting, COGITAO constructs rule-based tasks which apply a set of transformations to objects in grid-like environments. It supports composition, at adjustable depth, over a set of 28 interoperable transformations, along with extensive control over grid parametrization and object properties. This flexibility enables the creation of millions of unique task rules -- surpassing concurrent datasets by several orders of magnitude -- across a wide range of difficulties, while allowing virtually unlimited sample generation per rule. We provide baseline experiments using state-of-the-art vision models, highlighting their consistent failures to generalize to novel combinations of familiar elements, despite strong in-domain performance. COGITAO is fully open-sourced, including all code and datasets, to support continued research in this field. |
| title | COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2509.05249 |