Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Peddi, Rohith, Saurabh, Shanmugam, Shravan, Pallapothula, Likhitha, Xiang, Yu, Singla, Parag, Gogate, Vibhav |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation
by: Peddi, Rohith, et al.
Published: (2024)
by: Peddi, Rohith, et al.
Published: (2024)
Towards Scene Graph Anticipation
by: Peddi, Rohith, et al.
Published: (2024)
by: Peddi, Rohith, et al.
Published: (2024)
Grasping Trajectory Optimization with Point Clouds
by: Xiang, Yu, et al.
Published: (2024)
by: Xiang, Yu, et al.
Published: (2024)
CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities
by: Peddi, Rohith, et al.
Published: (2023)
by: Peddi, Rohith, et al.
Published: (2023)
Deep Dependency Networks and Advanced Inference Schemes for Multi-Label Classification
by: Arya, Shivvrat, et al.
Published: (2024)
by: Arya, Shivvrat, et al.
Published: (2024)
Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
by: Gao, Ziqi, et al.
Published: (2026)
by: Gao, Ziqi, et al.
Published: (2026)
Towards Long-Form Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2026)
by: Gu, Xin, et al.
Published: (2026)
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
by: Sohn, Tin Stribor, et al.
Published: (2025)
by: Sohn, Tin Stribor, et al.
Published: (2025)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
One World, Dual Timeline: Decoupled Spatio-Temporal Gaussian Scene Graph for 4D Cooperative Driving Reconstruction
by: Chen, Yulong, et al.
Published: (2026)
by: Chen, Yulong, et al.
Published: (2026)
Video-Language Alignment via Spatio-Temporal Graph Transformer
by: Zhang, Shi-Xue, et al.
Published: (2024)
by: Zhang, Shi-Xue, et al.
Published: (2024)
Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
by: Huang, Junchao, et al.
Published: (2025)
by: Huang, Junchao, et al.
Published: (2025)
DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
by: Chu, Wen-Hsuan, et al.
Published: (2024)
by: Chu, Wen-Hsuan, et al.
Published: (2024)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
Depth-Aware Rover: A Study of Edge AI and Monocular Vision for Real-World Implementation
by: Relia, Lomash, et al.
Published: (2026)
by: Relia, Lomash, et al.
Published: (2026)
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
by: Grover, Shresth, et al.
Published: (2025)
by: Grover, Shresth, et al.
Published: (2025)
SpatioTemporal Difference Network for Video Depth Super-Resolution
by: Wang, Zhengxue, et al.
Published: (2025)
by: Wang, Zhengxue, et al.
Published: (2025)
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
by: Xing, Ke, et al.
Published: (2025)
by: Xing, Ke, et al.
Published: (2025)
Toward Scene Graph and Layout Guided Complex 3D Scene Generation
by: Huang, Yu-Hsiang, et al.
Published: (2024)
by: Huang, Yu-Hsiang, et al.
Published: (2024)
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
by: Yao, Jiali, et al.
Published: (2025)
by: Yao, Jiali, et al.
Published: (2025)
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
by: Guo, Xiangyu, et al.
Published: (2025)
by: Guo, Xiangyu, et al.
Published: (2025)
G-MSGINet: A Grouped Multi-Scale Graph-Involution Network for Contactless Fingerprint Recognition
by: Peddi, Santhoshkumar, et al.
Published: (2025)
by: Peddi, Santhoshkumar, et al.
Published: (2025)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
by: Ma, Jingtian, et al.
Published: (2025)
by: Ma, Jingtian, et al.
Published: (2025)
WorldTree: Towards 4D Dynamic Worlds from Monocular Video using Tree-Chains
by: Wang, Qisen, et al.
Published: (2026)
by: Wang, Qisen, et al.
Published: (2026)
OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition
by: Chen, Tongjia, et al.
Published: (2023)
by: Chen, Tongjia, et al.
Published: (2023)
DriveFix: Spatio-Temporally Coherent Driving Scene Restoration
by: Si, Heyu, et al.
Published: (2026)
by: Si, Heyu, et al.
Published: (2026)
STGFormer: Spatio-Temporal GraphFormer for 3D Human Pose Estimation in Video
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Context-Guided Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
SAMJAM: Zero-Shot Video Scene Graph Generation for Egocentric Kitchen Videos
by: Li, Joshua, et al.
Published: (2025)
by: Li, Joshua, et al.
Published: (2025)
Broadening View Synthesis of Dynamic Scenes from Constrained Monocular Videos
by: Jiang, Le, et al.
Published: (2025)
by: Jiang, Le, et al.
Published: (2025)
Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
by: Li, Qirui, et al.
Published: (2025)
by: Li, Qirui, et al.
Published: (2025)
Monocular Occupancy Prediction for Scalable Indoor Scenes
by: Yu, Hongxiao, et al.
Published: (2024)
by: Yu, Hongxiao, et al.
Published: (2024)
StableDPT: Temporal Stable Monocular Video Depth Estimation
by: Sobko, Ivan, et al.
Published: (2026)
by: Sobko, Ivan, et al.
Published: (2026)
SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models
by: Makarov, Vladislav, et al.
Published: (2026)
by: Makarov, Vladislav, et al.
Published: (2026)
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
by: Aparcedo, Alejandro, et al.
Published: (2026)
by: Aparcedo, Alejandro, et al.
Published: (2026)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
Similar Items
-
Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation
by: Peddi, Rohith, et al.
Published: (2024) -
Towards Scene Graph Anticipation
by: Peddi, Rohith, et al.
Published: (2024) -
Grasping Trajectory Optimization with Point Clouds
by: Xiang, Yu, et al.
Published: (2024) -
CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities
by: Peddi, Rohith, et al.
Published: (2023) -
Deep Dependency Networks and Advanced Inference Schemes for Multi-Label Classification
by: Arya, Shivvrat, et al.
Published: (2024)