CauSight: Learning to Supersense for Visual Causal Discovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yize, Chen, Meiqi, Chen, Sirui, Peng, Bo, Zhang, Yanxi, Li, Tianyu, Lu, Chaochao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915647183650816
author Zhang, Yize
Chen, Meiqi
Chen, Sirui
Peng, Bo
Zhang, Yanxi
Li, Tianyu
Lu, Chaochao
author_facet Zhang, Yize
Chen, Meiqi
Chen, Sirui
Peng, Bo
Zhang, Yanxi
Li, Tianyu
Lu, Chaochao
contents Causal thinking enables humans to understand not just what is seen, but why it happens. To replicate this capability in modern AI systems, we introduce the task of visual causal discovery. It requires models to infer cause-and-effect relations among visual entities across diverse scenarios instead of merely perceiving their presence. To this end, we first construct the Visual Causal Graph dataset (VCG-32K), a large-scale collection of over 32,000 images annotated with entity-level causal graphs, and further develop CauSight, a novel vision-language model to perform visual causal discovery through causally aware reasoning. Our training recipe integrates three components: (1) training data curation from VCG-32K, (2) Tree-of-Causal-Thought (ToCT) for synthesizing reasoning trajectories, and (3) reinforcement learning with a designed causal reward to refine the reasoning policy. Experiments show that CauSight outperforms GPT-4.1 on visual causal discovery, achieving over a threefold performance boost (21% absolute gain). Our code, model, and dataset are fully open-sourced at project page: https://github.com/OpenCausaLab/CauSight.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CauSight: Learning to Supersense for Visual Causal Discovery
Zhang, Yize
Chen, Meiqi
Chen, Sirui
Peng, Bo
Zhang, Yanxi
Li, Tianyu
Lu, Chaochao
Computer Vision and Pattern Recognition
Causal thinking enables humans to understand not just what is seen, but why it happens. To replicate this capability in modern AI systems, we introduce the task of visual causal discovery. It requires models to infer cause-and-effect relations among visual entities across diverse scenarios instead of merely perceiving their presence. To this end, we first construct the Visual Causal Graph dataset (VCG-32K), a large-scale collection of over 32,000 images annotated with entity-level causal graphs, and further develop CauSight, a novel vision-language model to perform visual causal discovery through causally aware reasoning. Our training recipe integrates three components: (1) training data curation from VCG-32K, (2) Tree-of-Causal-Thought (ToCT) for synthesizing reasoning trajectories, and (3) reinforcement learning with a designed causal reward to refine the reasoning policy. Experiments show that CauSight outperforms GPT-4.1 on visual causal discovery, achieving over a threefold performance boost (21% absolute gain). Our code, model, and dataset are fully open-sourced at project page: https://github.com/OpenCausaLab/CauSight.
title CauSight: Learning to Supersense for Visual Causal Discovery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.01827