Robust Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Mengshi, Lv, Changsheng, Ma, Huadong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving
by: Lv, Changsheng, et al.
Published: (2024)
by: Lv, Changsheng, et al.
Published: (2024)
Question-Aware Evidence Ledgers for Video Relational Reasoning
by: Ou, Yilin, et al.
Published: (2026)
by: Ou, Yilin, et al.
Published: (2026)
Multi-Stage Contrastive Regression for Action Quality Assessment
by: An, Qi, et al.
Published: (2024)
by: An, Qi, et al.
Published: (2024)
Learning Group Interactions and Semantic Intentions for Multi-Object Trajectory Prediction
by: Qi, Mengshi, et al.
Published: (2024)
by: Qi, Mengshi, et al.
Published: (2024)
Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph Generation
by: Lv, Changsheng, et al.
Published: (2025)
by: Lv, Changsheng, et al.
Published: (2025)
Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Improving Batch Normalization with TTA for Robust Object Detection in Self-Driving
by: Liao, Dacheng, et al.
Published: (2024)
by: Liao, Dacheng, et al.
Published: (2024)
Decomposed Vector-Quantized Variational Autoencoder for Human Grasp Generation
by: Zhao, Zhe, et al.
Published: (2024)
by: Zhao, Zhe, et al.
Published: (2024)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
Towards Robust Unsupervised Attention Prediction in Autonomous Driving
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
by: Lin, Yuxin, et al.
Published: (2025)
by: Lin, Yuxin, et al.
Published: (2025)
A New Teacher-Reviewer-Student Framework for Semi-supervised 2D Human Pose Estimation
by: Yun, Wulian, et al.
Published: (2025)
by: Yun, Wulian, et al.
Published: (2025)
Towards Efficient Object Re-Identification with A Novel Cloud-Edge Collaborative Framework
by: Wang, Chuanming, et al.
Published: (2024)
by: Wang, Chuanming, et al.
Published: (2024)
Semi-Supervised Teacher-Reference-Student Architecture for Action Quality Assessment
by: Yun, Wulian, et al.
Published: (2024)
by: Yun, Wulian, et al.
Published: (2024)
Uncovering the human motion pattern: Pattern Memory-based Diffusion Model for Trajectory Prediction
by: Yang, Yuxin, et al.
Published: (2024)
by: Yang, Yuxin, et al.
Published: (2024)
Mutual Distillation Learning For Person Re-Identification
by: Fu, Huiyuan, et al.
Published: (2024)
by: Fu, Huiyuan, et al.
Published: (2024)
Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
VIoTGPT: Learning to Schedule Vision Tools in LLMs towards Intelligent Video Internet of Things
by: Zhong, Yaoyao, et al.
Published: (2023)
by: Zhong, Yaoyao, et al.
Published: (2023)
DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models
by: Deng, Wei, et al.
Published: (2026)
by: Deng, Wei, et al.
Published: (2026)
CausalDisenSeg: A Causality-Guided Disentanglement Framework with Counterfactual Reasoning for Robust Brain Tumor Segmentation Under Missing Modalities
by: Liu, Bo, et al.
Published: (2026)
by: Liu, Bo, et al.
Published: (2026)
Causal Debiasing for Visual Commonsense Reasoning
by: Zou, Jiayi, et al.
Published: (2025)
by: Zou, Jiayi, et al.
Published: (2025)
Interpretable Perception and Reasoning for Audiovisual Geolocation
by: Su, Yiyang, et al.
Published: (2026)
by: Su, Yiyang, et al.
Published: (2026)
Disentanglement-Based Equivariant Learning for Compositional VQA
by: Du, Zhou, et al.
Published: (2026)
by: Du, Zhou, et al.
Published: (2026)
Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation
by: Li, Xiaochuan, et al.
Published: (2025)
by: Li, Xiaochuan, et al.
Published: (2025)
EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning
by: Ma, Mingjie, et al.
Published: (2024)
by: Ma, Mingjie, et al.
Published: (2024)
Improving Partially Observed Trajectories Forecasting by Target-driven Self-Distillation
by: Shu, Peng, et al.
Published: (2025)
by: Shu, Peng, et al.
Published: (2025)
Multi-modal Document Presentation Attack Detection With Forensics Trace Disentanglement
by: Chen, Changsheng, et al.
Published: (2024)
by: Chen, Changsheng, et al.
Published: (2024)
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
by: Yang, Siqi, et al.
Published: (2025)
by: Yang, Siqi, et al.
Published: (2025)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
by: Chen, Jiali, et al.
Published: (2024)
by: Chen, Jiali, et al.
Published: (2024)
Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models
by: Xia, Yuexuan, et al.
Published: (2025)
by: Xia, Yuexuan, et al.
Published: (2025)
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
by: Kundu, Sanjoy, et al.
Published: (2024)
by: Kundu, Sanjoy, et al.
Published: (2024)
DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
by: Zou, Jiayi, et al.
Published: (2025)
by: Zou, Jiayi, et al.
Published: (2025)
Augmented Commonsense Knowledge for Remote Object Grounding
by: Mohammadi, Bahram, et al.
Published: (2024)
by: Mohammadi, Bahram, et al.
Published: (2024)
Learning Exposure Correction in Dynamic Scenes
by: Liu, Jin, et al.
Published: (2024)
by: Liu, Jin, et al.
Published: (2024)
A Study of Commonsense Reasoning over Visual Object Properties
by: Kolari, Abhishek, et al.
Published: (2025)
by: Kolari, Abhishek, et al.
Published: (2025)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Similar Items
-
T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving
by: Lv, Changsheng, et al.
Published: (2024) -
Question-Aware Evidence Ledgers for Video Relational Reasoning
by: Ou, Yilin, et al.
Published: (2026) -
Multi-Stage Contrastive Regression for Action Quality Assessment
by: An, Qi, et al.
Published: (2024) -
Learning Group Interactions and Semantic Intentions for Multi-Object Trajectory Prediction
by: Qi, Mengshi, et al.
Published: (2024) -
Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph Generation
by: Lv, Changsheng, et al.
Published: (2025)