Unified 3D Scene Understanding Through Physical World Modeling
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Wanhee, Kotar, Klemen, Venkatesh, Rahul Mysore, Watrous, Jared, Chen, Honglin, Aw, Khai Loong, Yamins, Daniel L. K. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
3D Scene Understanding Through Local Random Access Sequence Modeling
por: Lee, Wanhee, et al.
Publicado: (2025)
por: Lee, Wanhee, et al.
Publicado: (2025)
Zero-shot World Models Are Developmentally Efficient Learners
por: Aw, Khai Loong, et al.
Publicado: (2026)
por: Aw, Khai Loong, et al.
Publicado: (2026)
Physical Object Understanding with a Physically Controllable World Model
por: Venkatesh, Rahul, et al.
Publicado: (2026)
por: Venkatesh, Rahul, et al.
Publicado: (2026)
World Modeling with Probabilistic Structure Integration
por: Kotar, Klemen, et al.
Publicado: (2025)
por: Kotar, Klemen, et al.
Publicado: (2025)
Taming generative video models for zero-shot optical flow extraction
por: Kim, Seungwoo, et al.
Publicado: (2025)
por: Kim, Seungwoo, et al.
Publicado: (2025)
Understanding Physical Dynamics with Counterfactual World Modeling
por: Venkatesh, Rahul, et al.
Publicado: (2023)
por: Venkatesh, Rahul, et al.
Publicado: (2023)
Discovering and using Spelke segments
por: Venkatesh, Rahul, et al.
Publicado: (2025)
por: Venkatesh, Rahul, et al.
Publicado: (2025)
Representing Speech Through Autoregressive Prediction of Cochlear Tokens
por: Tuckute, Greta, et al.
Publicado: (2025)
por: Tuckute, Greta, et al.
Publicado: (2025)
Model Connectomes: A Generational Approach to Data-Efficient Language Models
por: Kotar, Klemen, et al.
Publicado: (2025)
por: Kotar, Klemen, et al.
Publicado: (2025)
Understanding Quantum Information and Computation
por: Watrous, John
Publicado: (2025)
por: Watrous, John
Publicado: (2025)
Instruction-tuning Aligns LLMs to the Human Brain
por: Aw, Khai Loong, et al.
Publicado: (2023)
por: Aw, Khai Loong, et al.
Publicado: (2023)
Preference optimization of protein language models as a multi-objective binder design paradigm
por: Mistani, Pouria, et al.
Publicado: (2024)
por: Mistani, Pouria, et al.
Publicado: (2024)
Characterizing the visual representation of objects from the child's view
por: Yang, Jane, et al.
Publicado: (2026)
por: Yang, Jane, et al.
Publicado: (2026)
Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals
por: Stojanov, Stefan, et al.
Publicado: (2025)
por: Stojanov, Stefan, et al.
Publicado: (2025)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
por: Zhou, Xin, et al.
Publicado: (2026)
por: Zhou, Xin, et al.
Publicado: (2026)
A Unified Framework for 3D Scene Understanding
por: Xu, Wei, et al.
Publicado: (2024)
por: Xu, Wei, et al.
Publicado: (2024)
Unified Semantic Transformer for 3D Scene Understanding
por: Koch, Sebastian, et al.
Publicado: (2025)
por: Koch, Sebastian, et al.
Publicado: (2025)
Toward Quantum Utility in Finance: A Robust Data-Driven Algorithm for Asset Clustering
por: Sharma, Shivam, et al.
Publicado: (2025)
por: Sharma, Shivam, et al.
Publicado: (2025)
Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions
por: Cross, Logan, et al.
Publicado: (2025)
por: Cross, Logan, et al.
Publicado: (2025)
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
por: Zhou, Xin, et al.
Publicado: (2025)
por: Zhou, Xin, et al.
Publicado: (2025)
Know Your Library; A Guide for Education Students. Know Your Library Series, No. 5.
por: Watrous, Lyle C., Comp.
Publicado: (1975)
por: Watrous, Lyle C., Comp.
Publicado: (1975)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
por: Deng, Tianchen, et al.
Publicado: (2025)
por: Deng, Tianchen, et al.
Publicado: (2025)
Improved Detection Performance of Cognitive Radio Networks in AWGN and Rayleigh Fading Environments
por: Ying Loong Lee
Publicado: (2013)
por: Ying Loong Lee
Publicado: (2013)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
por: Huang, Ting, et al.
Publicado: (2025)
por: Huang, Ting, et al.
Publicado: (2025)
Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding
por: Mao, Ye, et al.
Publicado: (2026)
por: Mao, Ye, et al.
Publicado: (2026)
Understanding Human Limits in Pattern Recognition: A Computational Model of Sequential Reasoning in Rock, Paper, Scissors
por: Cross, Logan, et al.
Publicado: (2025)
por: Cross, Logan, et al.
Publicado: (2025)
OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
por: Gao, Sensen, et al.
Publicado: (2026)
por: Gao, Sensen, et al.
Publicado: (2026)
Emergence of Fluctuation Relations in UNO
por: Sidajaya, Peter, et al.
Publicado: (2024)
por: Sidajaya, Peter, et al.
Publicado: (2024)
R3DS: Reality-linked 3D Scenes for Panoramic Scene Understanding
por: Wu, Qirui, et al.
Publicado: (2024)
por: Wu, Qirui, et al.
Publicado: (2024)
Prediction-Based Markov Violation Scores for Detecting Non-Markovian Observations in Reinforcement Learning
por: Mysore, Naveen
Publicado: (2026)
por: Mysore, Naveen
Publicado: (2026)
Quantifying First-Order Markov Violations in Noisy Reinforcement Learning: A Causal Discovery Approach
por: Mysore, Naveen
Publicado: (2025)
por: Mysore, Naveen
Publicado: (2025)
Temporal Functional Circuits: From Spline Plots to Faithful Explanations in KAN Forecasting
por: Mysore, Naveen
Publicado: (2026)
por: Mysore, Naveen
Publicado: (2026)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
por: Mysore, Naveen
Publicado: (2026)
por: Mysore, Naveen
Publicado: (2026)
OpenSU3D: Open World 3D Scene Understanding using Foundation Models
por: Mohiuddin, Rafay, et al.
Publicado: (2024)
por: Mohiuddin, Rafay, et al.
Publicado: (2024)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
por: Tan, Alvin Wei Ming, et al.
Publicado: (2025)
por: Tan, Alvin Wei Ming, et al.
Publicado: (2025)
TUN3D: Towards Real-World Scene Understanding from Unposed Images
por: Konushin, Anton, et al.
Publicado: (2025)
por: Konushin, Anton, et al.
Publicado: (2025)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
por: Zhuo, Dong, et al.
Publicado: (2026)
por: Zhuo, Dong, et al.
Publicado: (2026)
Unified Editing of Panorama, 3D Scenes, and Videos Through Disentangled Self-Attention Injection
por: Kwon, Gihyun, et al.
Publicado: (2024)
por: Kwon, Gihyun, et al.
Publicado: (2024)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
por: Halacheva, Anna-Maria, et al.
Publicado: (2024)
por: Halacheva, Anna-Maria, et al.
Publicado: (2024)
Quantum Annealing-Based Algorithm for Efficient Coalition Formation Among LEO Satellites
por: Venkatesh, Supreeth Mysore, et al.
Publicado: (2024)
por: Venkatesh, Supreeth Mysore, et al.
Publicado: (2024)
Ejemplares similares
-
3D Scene Understanding Through Local Random Access Sequence Modeling
por: Lee, Wanhee, et al.
Publicado: (2025) -
Zero-shot World Models Are Developmentally Efficient Learners
por: Aw, Khai Loong, et al.
Publicado: (2026) -
Physical Object Understanding with a Physically Controllable World Model
por: Venkatesh, Rahul, et al.
Publicado: (2026) -
World Modeling with Probabilistic Structure Integration
por: Kotar, Klemen, et al.
Publicado: (2025) -
Taming generative video models for zero-shot optical flow extraction
por: Kim, Seungwoo, et al.
Publicado: (2025)