Video Spatial Reasoning with Object-Centric 3D Rollout
Fuente:
arXiv
Salvato in:
| Autori principali: | Tang, Haoran, Cao, Meng, Liu, Ruyang, Liang, Xiaoxi, Li, Linglong, Li, Ge, Liang, Xiaodan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval
di: Tang, Haoran, et al.
Pubblicazione: (2024)
di: Tang, Haoran, et al.
Pubblicazione: (2024)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
di: Cao, Meng, et al.
Pubblicazione: (2024)
di: Cao, Meng, et al.
Pubblicazione: (2024)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
di: Cao, Meng, et al.
Pubblicazione: (2024)
di: Cao, Meng, et al.
Pubblicazione: (2024)
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
di: Cao, Meng, et al.
Pubblicazione: (2026)
di: Cao, Meng, et al.
Pubblicazione: (2026)
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
di: Cao, Meng, et al.
Pubblicazione: (2025)
di: Cao, Meng, et al.
Pubblicazione: (2025)
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
di: Liu, Ruyang, et al.
Pubblicazione: (2025)
di: Liu, Ruyang, et al.
Pubblicazione: (2025)
ST-LLM: Large Language Models Are Effective Temporal Learners
di: Liu, Ruyang, et al.
Pubblicazione: (2024)
di: Liu, Ruyang, et al.
Pubblicazione: (2024)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
di: Liu, Ruyang, et al.
Pubblicazione: (2023)
di: Liu, Ruyang, et al.
Pubblicazione: (2023)
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
di: Cao, Meng, et al.
Pubblicazione: (2025)
di: Cao, Meng, et al.
Pubblicazione: (2025)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
di: Sun, Shangkun, et al.
Pubblicazione: (2024)
di: Sun, Shangkun, et al.
Pubblicazione: (2024)
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
di: Li, Haoyuan, et al.
Pubblicazione: (2026)
di: Li, Haoyuan, et al.
Pubblicazione: (2026)
Reasoning-Enhanced Object-Centric Learning for Videos
di: Li, Jian, et al.
Pubblicazione: (2024)
di: Li, Jian, et al.
Pubblicazione: (2024)
Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
di: Mirjalili, Vahid, et al.
Pubblicazione: (2025)
di: Mirjalili, Vahid, et al.
Pubblicazione: (2025)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
di: Ma, Wenxin, et al.
Pubblicazione: (2026)
di: Ma, Wenxin, et al.
Pubblicazione: (2026)
Object-Centric Framework for Video Moment Retrieval
di: Li, Zongyao, et al.
Pubblicazione: (2025)
di: Li, Zongyao, et al.
Pubblicazione: (2025)
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
di: Cao, Meng, et al.
Pubblicazione: (2025)
di: Cao, Meng, et al.
Pubblicazione: (2025)
Simultaneous Detection and Interaction Reasoning for Object-Centric Action Recognition
di: Li, Xunsong, et al.
Pubblicazione: (2024)
di: Li, Xunsong, et al.
Pubblicazione: (2024)
Motion-aware Memory Network for Fast Video Salient Object Detection
di: Zhao, Xing, et al.
Pubblicazione: (2022)
di: Zhao, Xing, et al.
Pubblicazione: (2022)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
di: Hu, Pengfei, et al.
Pubblicazione: (2025)
di: Hu, Pengfei, et al.
Pubblicazione: (2025)
Cycle Consistency in Video Object-Centric Learning
di: Zhao, Rongzhen, et al.
Pubblicazione: (2026)
di: Zhao, Rongzhen, et al.
Pubblicazione: (2026)
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
di: Cao, Meng, et al.
Pubblicazione: (2025)
di: Cao, Meng, et al.
Pubblicazione: (2025)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
di: Liu, Yuanxin, et al.
Pubblicazione: (2025)
di: Liu, Yuanxin, et al.
Pubblicazione: (2025)
Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
di: Huang, Shaofei, et al.
Pubblicazione: (2024)
di: Huang, Shaofei, et al.
Pubblicazione: (2024)
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
di: Ge, Mengmeng, et al.
Pubblicazione: (2026)
di: Ge, Mengmeng, et al.
Pubblicazione: (2026)
3D Feature Distillation with Object-Centric Priors
di: Tziafas, Georgios, et al.
Pubblicazione: (2024)
di: Tziafas, Georgios, et al.
Pubblicazione: (2024)
Unsupervised Learning of Category-Level 3D Pose from Object-Centric Videos
di: Sommer, Leonhard, et al.
Pubblicazione: (2024)
di: Sommer, Leonhard, et al.
Pubblicazione: (2024)
LucidDreaming: Controllable Object-Centric 3D Generation
di: Wang, Zhaoning, et al.
Pubblicazione: (2023)
di: Wang, Zhaoning, et al.
Pubblicazione: (2023)
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
di: Xu, Rongtao, et al.
Pubblicazione: (2025)
di: Xu, Rongtao, et al.
Pubblicazione: (2025)
3D Visibility-aware Generalizable Neural Radiance Fields for Interacting Hands
di: Huang, Xuan, et al.
Pubblicazione: (2024)
di: Huang, Xuan, et al.
Pubblicazione: (2024)
Online Reasoning Video Object Segmentation
di: Liu, Jinyuan, et al.
Pubblicazione: (2026)
di: Liu, Jinyuan, et al.
Pubblicazione: (2026)
SFGFusion: Surface Fitting Guided 3D Object Detection with 4D Radar and Camera Fusion
di: Li, Xiaozhi, et al.
Pubblicazione: (2025)
di: Li, Xiaozhi, et al.
Pubblicazione: (2025)
Towards 3D Object-Centric Feature Learning for Semantic Scene Completion
di: Wang, Weihua, et al.
Pubblicazione: (2025)
di: Wang, Weihua, et al.
Pubblicazione: (2025)
R2G: Reasoning to Ground in 3D Scenes
di: Li, Yixuan, et al.
Pubblicazione: (2024)
di: Li, Yixuan, et al.
Pubblicazione: (2024)
NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
di: Liu, Xiangyan, et al.
Pubblicazione: (2025)
di: Liu, Xiangyan, et al.
Pubblicazione: (2025)
GS-CLIP: Gaussian Splatting for Contrastive Language-Image-3D Pretraining from Real-World Data
di: Li, Haoyuan, et al.
Pubblicazione: (2024)
di: Li, Haoyuan, et al.
Pubblicazione: (2024)
WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
di: Guan, Runwei, et al.
Pubblicazione: (2026)
di: Guan, Runwei, et al.
Pubblicazione: (2026)
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
di: Qi, Zekun, et al.
Pubblicazione: (2024)
di: Qi, Zekun, et al.
Pubblicazione: (2024)
Internalizing Temporal Consistency in Video Object-Centric Learning without Explicit Regularization
di: Zhao, Rongzhen, et al.
Pubblicazione: (2026)
di: Zhao, Rongzhen, et al.
Pubblicazione: (2026)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
di: Ouyang, Kun, et al.
Pubblicazione: (2025)
di: Ouyang, Kun, et al.
Pubblicazione: (2025)
HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
di: Li, Yayuan, et al.
Pubblicazione: (2024)
di: Li, Yayuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval
di: Tang, Haoran, et al.
Pubblicazione: (2024) -
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
di: Cao, Meng, et al.
Pubblicazione: (2024) -
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
di: Cao, Meng, et al.
Pubblicazione: (2024) -
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
di: Cao, Meng, et al.
Pubblicazione: (2026) -
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
di: Cao, Meng, et al.
Pubblicazione: (2025)