Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xingrui, Ma, Wufei, Wang, Angtian, Chen, Shuo, Kortylewski, Adam, Yuille, Alan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NOVUM: Neural Object Volumes for Robust Object Classification
von: Jesslen, Artur, et al.
Veröffentlicht: (2023)
von: Jesslen, Artur, et al.
Veröffentlicht: (2023)
VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis
von: Wang, Angtian, et al.
Veröffentlicht: (2022)
von: Wang, Angtian, et al.
Veröffentlicht: (2022)
PASR: Pose-Aware 3D Shape Retrieval from Occluded Single Views
von: Shi, Jiaxin, et al.
Veröffentlicht: (2026)
von: Shi, Jiaxin, et al.
Veröffentlicht: (2026)
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
Learning a Category-level Object Pose Estimator without Pose Annotations
von: Tian, Fengrui, et al.
Veröffentlicht: (2024)
von: Tian, Fengrui, et al.
Veröffentlicht: (2024)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
Generating Images with 3D Annotations Using Diffusion Models
von: Ma, Wufei, et al.
Veröffentlicht: (2023)
von: Ma, Wufei, et al.
Veröffentlicht: (2023)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
DINeMo: Learning Neural Mesh Models with no 3D Annotations
von: Guo, Weijie, et al.
Veröffentlicht: (2025)
von: Guo, Weijie, et al.
Veröffentlicht: (2025)
A Bayesian Approach to OOD Robustness in Image Classification
von: Kaushik, Prakhar, et al.
Veröffentlicht: (2024)
von: Kaushik, Prakhar, et al.
Veröffentlicht: (2024)
HECTOR: Hybrid Editable Compositional Object References for Video Generation
von: Zhang, Guofeng, et al.
Veröffentlicht: (2026)
von: Zhang, Guofeng, et al.
Veröffentlicht: (2026)
DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
von: Zhong, Shanshan, et al.
Veröffentlicht: (2025)
von: Zhong, Shanshan, et al.
Veröffentlicht: (2025)
TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
von: Sheung, Eddie Pokming, et al.
Veröffentlicht: (2025)
von: Sheung, Eddie Pokming, et al.
Veröffentlicht: (2025)
Prompt-Based Exemplar Super-Compression and Regeneration for Class-Incremental Learning
von: Duan, Ruxiao, et al.
Veröffentlicht: (2023)
von: Duan, Ruxiao, et al.
Veröffentlicht: (2023)
Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose Estimation
von: Kaushik, Prakhar, et al.
Veröffentlicht: (2024)
von: Kaushik, Prakhar, et al.
Veröffentlicht: (2024)
iNeMo: Incremental Neural Mesh Models for Robust Class-Incremental Learning
von: Fischer, Tom, et al.
Veröffentlicht: (2024)
von: Fischer, Tom, et al.
Veröffentlicht: (2024)
3D Question Answering for City Scene Understanding
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
Structure-Aware Sparse-View X-ray 3D Reconstruction
von: Cai, Yuanhao, et al.
Veröffentlicht: (2023)
von: Cai, Yuanhao, et al.
Veröffentlicht: (2023)
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
Gaussian Scenes: Pose-Free Sparse-View Scene Reconstruction using Depth-Enhanced Diffusion Priors
von: Paul, Soumava, et al.
Veröffentlicht: (2024)
von: Paul, Soumava, et al.
Veröffentlicht: (2024)
LychSim: A Controllable and Interactive Simulation Framework for Vision Research
von: Ma, Wufei, et al.
Veröffentlicht: (2026)
von: Ma, Wufei, et al.
Veröffentlicht: (2026)
Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence
von: Jesslen, Artur, et al.
Veröffentlicht: (2026)
von: Jesslen, Artur, et al.
Veröffentlicht: (2026)
Semantic Flow: Learning Semantic Field of Dynamic Scenes from Monocular Videos
von: Tian, Fengrui, et al.
Veröffentlicht: (2024)
von: Tian, Fengrui, et al.
Veröffentlicht: (2024)
From Pixel to Cancer: Cellular Automata in Computed Tomography
von: Lai, Yuxiang, et al.
Veröffentlicht: (2024)
von: Lai, Yuxiang, et al.
Veröffentlicht: (2024)
CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models
von: Begiristain, León, et al.
Veröffentlicht: (2026)
von: Begiristain, León, et al.
Veröffentlicht: (2026)
Category-Level 3D Correspondence in Camera Space via Morphable Object Priors
von: Sommer, Leonhard, et al.
Veröffentlicht: (2026)
von: Sommer, Leonhard, et al.
Veröffentlicht: (2026)
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
von: Paul, Soumava, et al.
Veröffentlicht: (2026)
von: Paul, Soumava, et al.
Veröffentlicht: (2026)
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
von: Zhang, Guofeng, et al.
Veröffentlicht: (2025)
von: Zhang, Guofeng, et al.
Veröffentlicht: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
von: Li, Zechuan, et al.
Veröffentlicht: (2025)
von: Li, Zechuan, et al.
Veröffentlicht: (2025)
The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
von: Wu, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhuoyuan, et al.
Veröffentlicht: (2025)
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
von: Luo, Jingzhou, et al.
Veröffentlicht: (2025)
von: Luo, Jingzhou, et al.
Veröffentlicht: (2025)
Unsupervised Learning of Category-Level 3D Pose from Object-Centric Videos
von: Sommer, Leonhard, et al.
Veröffentlicht: (2024)
von: Sommer, Leonhard, et al.
Veröffentlicht: (2024)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
NOVUM: Neural Object Volumes for Robust Object Classification
von: Jesslen, Artur, et al.
Veröffentlicht: (2023) -
VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis
von: Wang, Angtian, et al.
Veröffentlicht: (2022) -
PASR: Pose-Aware 3D Shape Retrieval from Occluded Single Views
von: Shi, Jiaxin, et al.
Veröffentlicht: (2026) -
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
von: Ma, Wufei, et al.
Veröffentlicht: (2024) -
Learning a Category-level Object Pose Estimator without Pose Annotations
von: Tian, Fengrui, et al.
Veröffentlicht: (2024)