MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Tieyuan, Liu, Huabin, He, Tianyao, Chen, Yihang, Gan, Chaofan, Ma, Xiao, Zhong, Cheng, Zhang, Yang, Wang, Yingxue, Lin, Hui, Lin, Weiyao |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning
par: Chen, Tieyuan, et autres
Publié: (2025)
par: Chen, Tieyuan, et autres
Publié: (2025)
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
par: He, Zhihao, et autres
Publié: (2025)
par: He, Zhihao, et autres
Publié: (2025)
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
par: Chen, Tieyuan, et autres
Publié: (2025)
par: Chen, Tieyuan, et autres
Publié: (2025)
CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental Learning
par: Chen, Tieyuan, et autres
Publié: (2025)
par: Chen, Tieyuan, et autres
Publié: (2025)
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
par: He, Zhihao, et autres
Publié: (2026)
par: He, Zhihao, et autres
Publié: (2026)
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
par: Gan, Chaofan, et autres
Publié: (2025)
par: Gan, Chaofan, et autres
Publié: (2025)
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
par: Gan, Chaofan, et autres
Publié: (2025)
par: Gan, Chaofan, et autres
Publié: (2025)
From Priors to Perception: Grounding Video-LLMs in Physical Reality
par: Zhao, Zicheng, et autres
Publié: (2026)
par: Zhao, Zicheng, et autres
Publié: (2026)
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
par: Qin, Ziran, et autres
Publié: (2025)
par: Qin, Ziran, et autres
Publié: (2025)
DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction
par: Gan, Chaofan, et autres
Publié: (2024)
par: Gan, Chaofan, et autres
Publié: (2024)
DND: Boosting Large Language Models with Dynamic Nested Depth
par: Chen, Tieyuan, et autres
Publié: (2025)
par: Chen, Tieyuan, et autres
Publié: (2025)
MCA: 2D-3D Retrieval with Noisy Labels via Multi-level Adaptive Correction and Alignment
par: Zou, Gui, et autres
Publié: (2025)
par: Zou, Gui, et autres
Publié: (2025)
Unlocking the Value of Text: Event-Driven Reasoning and Multi-Level Alignment for Time Series Forecasting
par: Wang, Siyuan, et autres
Publié: (2026)
par: Wang, Siyuan, et autres
Publié: (2026)
CogStream: Context-guided Streaming Video Question Answering
par: Zhao, Zicheng, et autres
Publié: (2025)
par: Zhao, Zicheng, et autres
Publié: (2025)
EventRR: Event Referential Reasoning for Referring Video Object Segmentation
par: Xu, Huihui, et autres
Publié: (2025)
par: Xu, Huihui, et autres
Publié: (2025)
HAC++: Towards 100X Compression of 3D Gaussian Splatting
par: Chen, Yihang, et autres
Publié: (2025)
par: Chen, Yihang, et autres
Publié: (2025)
HAC: Hash-grid Assisted Context for 3D Gaussian Splatting Compression
par: Chen, Yihang, et autres
Publié: (2024)
par: Chen, Yihang, et autres
Publié: (2024)
DHEA-MECD: An Embodied Intelligence-Powered DRL Algorithm for AUV Tracking in Underwater Environments with High-Dimensional Features
par: Tian, Kai, et autres
Publié: (2026)
par: Tian, Kai, et autres
Publié: (2026)
DIGIC: Domain Generalizable Imitation Learning by Causal Discovery
par: Chen, Yang, et autres
Publié: (2024)
par: Chen, Yang, et autres
Publié: (2024)
ProgD: Progressive Multi-scale Decoding with Dynamic Graphs for Joint Multi-agent Motion Forecasting
par: Gao, Xing, et autres
Publié: (2025)
par: Gao, Xing, et autres
Publié: (2025)
Fast Feedforward 3D Gaussian Splatting Compression
par: Chen, Yihang, et autres
Publié: (2024)
par: Chen, Yihang, et autres
Publié: (2024)
PCGS: Progressive Compression of 3D Gaussian Splatting
par: Chen, Yihang, et autres
Publié: (2025)
par: Chen, Yihang, et autres
Publié: (2025)
ACCESS : A Benchmark for Abstract Causal Event Discovery and Reasoning
par: Vo, Vy, et autres
Publié: (2025)
par: Vo, Vy, et autres
Publié: (2025)
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
par: Zhang, Zhiwei, et autres
Publié: (2025)
par: Zhang, Zhiwei, et autres
Publié: (2025)
Unlocking the Potential of Model Merging for Low-Resource Languages
par: Tao, Mingxu, et autres
Publié: (2024)
par: Tao, Mingxu, et autres
Publié: (2024)
Multi-Task Anti-Causal Learning for Reconstructing Urban Events from Residents' Reports
par: Zhou, Liangkai, et autres
Publié: (2026)
par: Zhou, Liangkai, et autres
Publié: (2026)
YingVideo-MV: Music-Driven Multi-Stage Video Generation
par: Chen, Jiahui, et autres
Publié: (2025)
par: Chen, Jiahui, et autres
Publié: (2025)
Finding the Trigger: Causal Abductive Reasoning on Video Events
par: Le, Thao Minh, et autres
Publié: (2025)
par: Le, Thao Minh, et autres
Publié: (2025)
Compositional Physical Reasoning of Objects and Events from Videos
par: Chen, Zhenfang, et autres
Publié: (2024)
par: Chen, Zhenfang, et autres
Publié: (2024)
Towards CausalGPT: A Multi-Agent Approach for Faithful Knowledge Reasoning via Promoting Causal Consistency in LLMs
par: Tang, Ziyi, et autres
Publié: (2023)
par: Tang, Ziyi, et autres
Publié: (2023)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
par: Liu, Huabin, et autres
Publié: (2025)
par: Liu, Huabin, et autres
Publié: (2025)
Robust Causal Discovery under Imperfect Structural Constraints
par: Wang, Zidong, et autres
Publié: (2025)
par: Wang, Zidong, et autres
Publié: (2025)
Codes for reproduce figures in deep clustering rain-induced noise paper
par: Shen, Junzhu, et autres
Publié: (2025)
par: Shen, Junzhu, et autres
Publié: (2025)
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
par: Wang, Enguang, et autres
Publié: (2024)
par: Wang, Enguang, et autres
Publié: (2024)
Generalized Category Discovery in Event-Centric Contexts: Latent Pattern Mining with LLMs
par: Luo, Yi, et autres
Publié: (2025)
par: Luo, Yi, et autres
Publié: (2025)
Cross-modal Causal Relation Alignment for Video Question Grounding
par: Chen, Weixing, et autres
Publié: (2025)
par: Chen, Weixing, et autres
Publié: (2025)
Local Causal Discovery for Structural Evidence of Direct Discrimination
par: Maasch, Jacqueline, et autres
Publié: (2024)
par: Maasch, Jacqueline, et autres
Publié: (2024)
CRAT: A Multi-Agent Framework for Causality-Enhanced Reflective and Retrieval-Augmented Translation with Large Language Models
par: Chen, Meiqi, et autres
Publié: (2024)
par: Chen, Meiqi, et autres
Publié: (2024)
GenAR: Next-Scale Autoregressive Generation for Spatial Gene Expression Prediction
par: Ouyang, Jiarui, et autres
Publié: (2025)
par: Ouyang, Jiarui, et autres
Publié: (2025)
Exploring Multi-Modal Data with Tool-Augmented LLM Agents for Precise Causal Discovery
par: Shen, ChengAo, et autres
Publié: (2024)
par: Shen, ChengAo, et autres
Publié: (2024)
Documents similaires
-
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning
par: Chen, Tieyuan, et autres
Publié: (2025) -
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
par: He, Zhihao, et autres
Publié: (2025) -
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
par: Chen, Tieyuan, et autres
Publié: (2025) -
CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental Learning
par: Chen, Tieyuan, et autres
Publié: (2025) -
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
par: He, Zhihao, et autres
Publié: (2026)