VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation
Fuente:
arXiv
Saved in:
| Main Authors: | O'Mahony, Felix, Cipolla, Roberto, Tewari, Ayush |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Color Equivariant Representations
by: Yang, Yulong, et al.
Published: (2024)
by: Yang, Yulong, et al.
Published: (2024)
On the Detection of Anomalous or Out-Of-Distribution Data in Vision Models Using Statistical Techniques
by: O'Mahony, Laura, et al.
Published: (2024)
by: O'Mahony, Laura, et al.
Published: (2024)
Towards Utilising a Range of Neural Activations for Comprehending Representational Associations
by: O'Mahony, Laura, et al.
Published: (2024)
by: O'Mahony, Laura, et al.
Published: (2024)
FOCUS -- Multi-View Foot Reconstruction From Synthetically Trained Dense Correspondences
by: Boyne, Oliver, et al.
Published: (2025)
by: Boyne, Oliver, et al.
Published: (2025)
FIND: An Unsupervised Implicit 3D Model of Articulated Human Feet
by: Boyne, Oliver, et al.
Published: (2022)
by: Boyne, Oliver, et al.
Published: (2022)
FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent
by: Smith, Cameron, et al.
Published: (2024)
by: Smith, Cameron, et al.
Published: (2024)
Efficient Camera-Controlled Video Generation of Static Scenes via Sparse Diffusion and 3D Rendering
by: Chen, Jieying, et al.
Published: (2026)
by: Chen, Jieying, et al.
Published: (2026)
PRAM: Place Recognition Anywhere Model for Efficient Visual Localization
by: Xue, Fei, et al.
Published: (2024)
by: Xue, Fei, et al.
Published: (2024)
NPLMV-PS: Neural Point-Light Multi-View Photometric Stereo
by: Logothetis, Fotios, et al.
Published: (2024)
by: Logothetis, Fotios, et al.
Published: (2024)
FOUND: Foot Optimization with Uncertain Normals for Surface Deformation Using Synthetic Data
by: Boyne, Oliver, et al.
Published: (2023)
by: Boyne, Oliver, et al.
Published: (2023)
LUCES-MV: A Multi-View Dataset for Near-Field Point Light Source Photometric Stereo
by: Logothetis, Fotios, et al.
Published: (2024)
by: Logothetis, Fotios, et al.
Published: (2024)
Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends
by: Liu, Jiuming, et al.
Published: (2026)
by: Liu, Jiuming, et al.
Published: (2026)
ReCoRe: Regularized Contrastive Representation Learning of World Model
by: Poudel, Rudra P. K., et al.
Published: (2023)
by: Poudel, Rudra P. K., et al.
Published: (2023)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
by: Zhang, Wanyue, et al.
Published: (2026)
by: Zhang, Wanyue, et al.
Published: (2026)
WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning
by: Englmeier, Stefan, et al.
Published: (2026)
by: Englmeier, Stefan, et al.
Published: (2026)
Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Faster 3D Gaussian Splatting Convergence via Structure-Aware Densification
by: Lyu, Linjie, et al.
Published: (2026)
by: Lyu, Linjie, et al.
Published: (2026)
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
by: Gu, Renjie, et al.
Published: (2026)
by: Gu, Renjie, et al.
Published: (2026)
VRS-NeRF: Visual Relocalization with Sparse Neural Radiance Field
by: Xue, Fei, et al.
Published: (2024)
by: Xue, Fei, et al.
Published: (2024)
GAURA: Generalizable Approach for Unified Restoration and Rendering of Arbitrary Views
by: Gupta, Vinayak, et al.
Published: (2024)
by: Gupta, Vinayak, et al.
Published: (2024)
DivAS: Interactive 3D Segmentation of NeRFs via Depth-Weighted Voxel Aggregation
by: Pande, Ayush
Published: (2026)
by: Pande, Ayush
Published: (2026)
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
Abstraction in Style
by: Lu, Min, et al.
Published: (2026)
by: Lu, Min, et al.
Published: (2026)
VLM-3D:End-to-End Vision-Language Models for Open-World 3D Perception
by: Chang, Fuhao, et al.
Published: (2025)
by: Chang, Fuhao, et al.
Published: (2025)
Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold
by: Pan, Xingang, et al.
Published: (2023)
by: Pan, Xingang, et al.
Published: (2023)
Understanding Multi-View Transformers
by: Stary, Michal, et al.
Published: (2025)
by: Stary, Michal, et al.
Published: (2025)
CubeletWorld: A New Abstraction for Scalable 3D Modeling
by: Samad, Azlaan Mustafa, et al.
Published: (2025)
by: Samad, Azlaan Mustafa, et al.
Published: (2025)
Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
by: Berman, Nimrod, et al.
Published: (2025)
by: Berman, Nimrod, et al.
Published: (2025)
VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models
by: Kang, Shuhao, et al.
Published: (2026)
by: Kang, Shuhao, et al.
Published: (2026)
LongVLM: Efficient Long Video Understanding via Large Language Models
by: Weng, Yuetian, et al.
Published: (2024)
by: Weng, Yuetian, et al.
Published: (2024)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Grounding World Simulation Models in a Real-World Metropolis
by: Seo, Junyoung, et al.
Published: (2026)
by: Seo, Junyoung, et al.
Published: (2026)
VisionArena: 230K Real World User-VLM Conversations with Preference Labels
by: Chou, Christopher, et al.
Published: (2024)
by: Chou, Christopher, et al.
Published: (2024)
Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding
by: Omasa, Takamitsu, et al.
Published: (2025)
by: Omasa, Takamitsu, et al.
Published: (2025)
WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models
by: Upadhyay, Rishi, et al.
Published: (2026)
by: Upadhyay, Rishi, et al.
Published: (2026)
VLM Models and Automated Grading of Atopic Dermatitis
by: Lalonde, Marc, et al.
Published: (2025)
by: Lalonde, Marc, et al.
Published: (2025)
Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
by: Tan, Kaizhen, et al.
Published: (2026)
by: Tan, Kaizhen, et al.
Published: (2026)
Real-time Structure Flow
by: Adarve, Juan David, et al.
Published: (2024)
by: Adarve, Juan David, et al.
Published: (2024)
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
Similar Items
-
Learning Color Equivariant Representations
by: Yang, Yulong, et al.
Published: (2024) -
On the Detection of Anomalous or Out-Of-Distribution Data in Vision Models Using Statistical Techniques
by: O'Mahony, Laura, et al.
Published: (2024) -
Towards Utilising a Range of Neural Activations for Comprehending Representational Associations
by: O'Mahony, Laura, et al.
Published: (2024) -
FOCUS -- Multi-View Foot Reconstruction From Synthetically Trained Dense Correspondences
by: Boyne, Oliver, et al.
Published: (2025) -
FIND: An Unsupervised Implicit 3D Model of Articulated Human Feet
by: Boyne, Oliver, et al.
Published: (2022)