CauSight: Learning to Supersense for Visual Causal Discovery
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yize, Chen, Meiqi, Chen, Sirui, Peng, Bo, Zhang, Yanxi, Li, Tianyu, Lu, Chaochao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CELLO: Causal Evaluation of Large Vision-Language Models
by: Chen, Meiqi, et al.
Published: (2024)
by: Chen, Meiqi, et al.
Published: (2024)
Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective
by: Chen, Meiqi, et al.
Published: (2024)
by: Chen, Meiqi, et al.
Published: (2024)
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
Solving Spatial Supersensing Without Spatial Supersensing
by: Udandarao, Vishaal, et al.
Published: (2025)
by: Udandarao, Vishaal, et al.
Published: (2025)
Cambrian-S: Towards Spatial Supersensing in Video
by: Yang, Shusheng, et al.
Published: (2025)
by: Yang, Shusheng, et al.
Published: (2025)
CauScale: Neural Causal Discovery at Scale
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
CauSkelNet: Causal Representation Learning for Human Behaviour Analysis
by: Gu, Xingrui, et al.
Published: (2024)
by: Gu, Xingrui, et al.
Published: (2024)
Toward Cognitive Supersensing in Multimodal Large Language Model
by: Li, Boyi, et al.
Published: (2026)
by: Li, Boyi, et al.
Published: (2026)
CauCLIP: Bridging the Sim-to-Real Gap in Surgical Video Understanding via Causality-Inspired Vision-Language Modeling
by: He, Yuxin, et al.
Published: (2026)
by: He, Yuxin, et al.
Published: (2026)
PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World
by: Wang, Changpeng, et al.
Published: (2026)
by: Wang, Changpeng, et al.
Published: (2026)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
by: Peng, Bo, et al.
Published: (2023)
by: Peng, Bo, et al.
Published: (2023)
ADAM: An Embodied Causal Agent in Open-World Environments
by: Yu, Shu, et al.
Published: (2024)
by: Yu, Shu, et al.
Published: (2024)
Understanding Robustness of Visual State Space Models for Image Classification
by: Du, Chengbin, et al.
Published: (2024)
by: Du, Chengbin, et al.
Published: (2024)
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
by: Li, Honglin, et al.
Published: (2024)
by: Li, Honglin, et al.
Published: (2024)
From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain
by: Golbari, Yuval, et al.
Published: (2026)
by: Golbari, Yuval, et al.
Published: (2026)
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
by: Li, Kaican, et al.
Published: (2025)
by: Li, Kaican, et al.
Published: (2025)
MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
The Shape of Sight: A Homological Framework for Unifying Visual Perception
by: Li, Xin
Published: (2018)
by: Li, Xin
Published: (2018)
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
by: Lu, Chaochao, et al.
Published: (2024)
by: Lu, Chaochao, et al.
Published: (2024)
Semantically Guided Dynamic Visual Prototype Refinement for Compositional Zero-Shot Learning
by: Peng, Zhong, et al.
Published: (2025)
by: Peng, Zhong, et al.
Published: (2025)
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
by: Chen, Tieyuan, et al.
Published: (2024)
by: Chen, Tieyuan, et al.
Published: (2024)
Novel Class Discovery for Point Cloud Segmentation via Joint Learning of Causal Representation and Reasoning
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
DTLLM-VLT: Diverse Text Generation for Visual Language Tracking Based on LLM
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Interpreting Low-level Vision Models with Causal Effect Maps
by: Hu, Jinfan, et al.
Published: (2024)
by: Hu, Jinfan, et al.
Published: (2024)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
by: Hu, Yucheng, et al.
Published: (2024)
by: Hu, Yucheng, et al.
Published: (2024)
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
by: Gong, Meiqi, et al.
Published: (2025)
by: Gong, Meiqi, et al.
Published: (2025)
Learning Triangular Distribution in Visual World
by: Chen, Ping, et al.
Published: (2023)
by: Chen, Ping, et al.
Published: (2023)
3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum
by: Zhang, Yuliang, et al.
Published: (2026)
by: Zhang, Yuliang, et al.
Published: (2026)
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
by: Xu, Tianyu, et al.
Published: (2025)
by: Xu, Tianyu, et al.
Published: (2025)
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
by: Duan, Zicheng, et al.
Published: (2026)
by: Duan, Zicheng, et al.
Published: (2026)
FootFormer: Estimating Stability from Visual Input
by: Kraiger, Keaton, et al.
Published: (2025)
by: Kraiger, Keaton, et al.
Published: (2025)
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
by: Yan, Yibin, et al.
Published: (2024)
by: Yan, Yibin, et al.
Published: (2024)
High-Fidelity Mask-free Neural Surface Reconstruction for Virtual Reality
by: Bai, Haotian, et al.
Published: (2024)
by: Bai, Haotian, et al.
Published: (2024)
Generalizable Non-Line-of-Sight Imaging with Learnable Physical Priors
by: Sun, Shida, et al.
Published: (2024)
by: Sun, Shida, et al.
Published: (2024)
Scalable Visual State Space Model with Fractal Scanning
by: Tang, Lv, et al.
Published: (2024)
by: Tang, Lv, et al.
Published: (2024)
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
StableIdentity: Inserting Anybody into Anywhere at First Sight
by: Wang, Qinghe, et al.
Published: (2024)
by: Wang, Qinghe, et al.
Published: (2024)
Implicit Counterfactual Learning for Audio-Visual Segmentation
by: Zha, Mingfeng, et al.
Published: (2025)
by: Zha, Mingfeng, et al.
Published: (2025)
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning
by: Chen, Tieyuan, et al.
Published: (2025)
by: Chen, Tieyuan, et al.
Published: (2025)
Few-shot Unknown Class Discovery of Hyperspectral Images with Prototype Learning and Clustering
by: Liu, Chun, et al.
Published: (2025)
by: Liu, Chun, et al.
Published: (2025)
Similar Items
-
CELLO: Causal Evaluation of Large Vision-Language Models
by: Chen, Meiqi, et al.
Published: (2024) -
Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective
by: Chen, Meiqi, et al.
Published: (2024) -
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
by: Peng, Bo, et al.
Published: (2026) -
Solving Spatial Supersensing Without Spatial Supersensing
by: Udandarao, Vishaal, et al.
Published: (2025) -
Cambrian-S: Towards Spatial Supersensing in Video
by: Yang, Shusheng, et al.
Published: (2025)