Counterfactual World Models via Digital Twin-conditioned Video Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Yiqing, Maksutova, Aiza, Li, Chenjia, Unberath, Mathias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text-Driven Reasoning Video Editing via Reinforcement Learning on Digital Twin Representations
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Online Reasoning Video Segmentation with Just-in-Time Digital Twins
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Fast Reasoning Segmentation for Images and Videos
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
RVTBench: A Benchmark for Visual Reasoning Tasks
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
by: Maksutova, Aiza, et al.
Published: (2026)
by: Maksutova, Aiza, et al.
Published: (2026)
Reasoning Segmentation for Images and Videos: A Survey
by: Shen, Yiqing, et al.
Published: (2025)
by: Shen, Yiqing, et al.
Published: (2025)
Memorizing SAM: 3D Medical Segment Anything Model with Memorizing Transformer
by: Shao, Xinyuan, et al.
Published: (2024)
by: Shao, Xinyuan, et al.
Published: (2024)
TwinOR: Photorealistic Digital Twins of Dynamic Operating Rooms for Embodied AI Research
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
Towards Controllable Video Synthesis of Routine and Rare OR Events
by: Schneider, Dominik, et al.
Published: (2026)
by: Schneider, Dominik, et al.
Published: (2026)
Towards Robust Algorithms for Surgical Phase Recognition via Digital Twin Representation
by: Ding, Hao, et al.
Published: (2024)
by: Ding, Hao, et al.
Published: (2024)
Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations
by: Li, Yizhen, et al.
Published: (2025)
by: Li, Yizhen, et al.
Published: (2025)
DiffuseReg: Denoising Diffusion Model for Obtaining Deformation Fields in Unsupervised Deformable Image Registration
by: Zhuo, Yongtai, et al.
Published: (2024)
by: Zhuo, Yongtai, et al.
Published: (2024)
A Causal Framework for Aligning Image Quality Metrics and Deep Neural Network Robustness
by: Drenkow, Nathan, et al.
Published: (2025)
by: Drenkow, Nathan, et al.
Published: (2025)
FastSAM3D: An Efficient Segment Anything Model for 3D Volumetric Medical Images
by: Shen, Yiqing, et al.
Published: (2024)
by: Shen, Yiqing, et al.
Published: (2024)
MoSFormer: Augmenting Temporal Context with Memory of Surgery for Surgical Phase Recognition
by: Ding, Hao, et al.
Published: (2025)
by: Ding, Hao, et al.
Published: (2025)
An Intrinsically Explainable Approach to Detecting Vertebral Compression Fractures in CT Scans via Neurosymbolic Modeling
by: Inigo, Blanca, et al.
Published: (2024)
by: Inigo, Blanca, et al.
Published: (2024)
Point Prompting: Counterfactual Tracking with Video Diffusion Models
by: Shrivastava, Ayush, et al.
Published: (2025)
by: Shrivastava, Ayush, et al.
Published: (2025)
Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion
by: Chen, Ting-Hsuan, et al.
Published: (2026)
by: Chen, Ting-Hsuan, et al.
Published: (2026)
Intelligent Control of Robotic X-ray Devices using a Language-promptable Digital Twin
by: Killeen, Benjamin D., et al.
Published: (2024)
by: Killeen, Benjamin D., et al.
Published: (2024)
Privacy-Preserving Operating Room Workflow Analysis using Digital Twins
by: Perez, Alejandra, et al.
Published: (2025)
by: Perez, Alejandra, et al.
Published: (2025)
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
by: He, Haoran, et al.
Published: (2024)
by: He, Haoran, et al.
Published: (2024)
Understanding Physical Dynamics with Counterfactual World Modeling
by: Venkatesh, Rahul, et al.
Published: (2023)
by: Venkatesh, Rahul, et al.
Published: (2023)
Causality-Driven Audits of Model Robustness
by: Drenkow, Nathan, et al.
Published: (2024)
by: Drenkow, Nathan, et al.
Published: (2024)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
From Generalization to Precision: Exploring SAM for Tool Segmentation in Surgical Environments
by: Oguine, Kanyifeechukwu J., et al.
Published: (2024)
by: Oguine, Kanyifeechukwu J., et al.
Published: (2024)
LD-ViCE: Latent Diffusion Model for Video Counterfactual Explanations
by: Varshney, Payal, et al.
Published: (2025)
by: Varshney, Payal, et al.
Published: (2025)
PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
by: Yin, Yuyang, et al.
Published: (2025)
by: Yin, Yuyang, et al.
Published: (2025)
Understanding the Implicit User Intention via Reasoning with Large Language Model for Image Editing
by: Wang, Yijia, et al.
Published: (2025)
by: Wang, Yijia, et al.
Published: (2025)
CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models
by: Begiristain, León, et al.
Published: (2026)
by: Begiristain, León, et al.
Published: (2026)
Causally Steered Diffusion for Automated Video Counterfactual Generation
by: Spyrou, Nikos, et al.
Published: (2025)
by: Spyrou, Nikos, et al.
Published: (2025)
Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis
by: Gao, Jianzhe, et al.
Published: (2026)
by: Gao, Jianzhe, et al.
Published: (2026)
World-consistent Video Diffusion with Explicit 3D Modeling
by: Zhang, Qihang, et al.
Published: (2024)
by: Zhang, Qihang, et al.
Published: (2024)
VideoShield: Regulating Diffusion-based Video Generation Models via Watermarking
by: Hu, Runyi, et al.
Published: (2025)
by: Hu, Runyi, et al.
Published: (2025)
Blind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion Model
by: Wang, Shuyun, et al.
Published: (2026)
by: Wang, Shuyun, et al.
Published: (2026)
Hyperspectral Image Recovery Constrained by Multi-Granularity Non-Local Self-Similarity Priors
by: Peng, Zhuoran, et al.
Published: (2025)
by: Peng, Zhuoran, et al.
Published: (2025)
TDM: Temporally-Consistent Diffusion Model for All-in-One Real-World Video Restoration
by: Li, Yizhou, et al.
Published: (2025)
by: Li, Yizhou, et al.
Published: (2025)
Similar Items
-
Text-Driven Reasoning Video Editing via Reinforcement Learning on Digital Twin Representations
by: Shen, Yiqing, et al.
Published: (2025) -
Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models
by: Shen, Yiqing, et al.
Published: (2025) -
Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
by: Shen, Yiqing, et al.
Published: (2025) -
Online Reasoning Video Segmentation with Just-in-Time Digital Twins
by: Shen, Yiqing, et al.
Published: (2025) -
Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction
by: Shen, Yiqing, et al.
Published: (2025)