Aether: Geometric-Aware Unified World Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Aether Team, Zhu, Haoyi, Wang, Yifan, Zhou, Jianjun, Chang, Wenzheng, Zhou, Yang, Li, Zizun, Chen, Junyi, Shen, Chunhua, Pang, Jiangmiao, He, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$π^3$: Permutation-Equivariant Visual Geometry Learning
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
DeepVerse: 4D Autoregressive Video Generation as a World Model
by: Chen, Junyi, et al.
Published: (2025)
by: Chen, Junyi, et al.
Published: (2025)
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
by: Li, Zizun, et al.
Published: (2025)
by: Li, Zizun, et al.
Published: (2025)
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
by: Peng, Jiaqi, et al.
Published: (2025)
by: Peng, Jiaqi, et al.
Published: (2025)
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
by: Zhu, Haoyi, et al.
Published: (2024)
by: Zhu, Haoyi, et al.
Published: (2024)
Geo-Align: Video Generation Alignment via Metric Geometry Reward
by: Li, Zizun, et al.
Published: (2026)
by: Li, Zizun, et al.
Published: (2026)
LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry
by: Peng, Jiaqi, et al.
Published: (2025)
by: Peng, Jiaqi, et al.
Published: (2025)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
by: Wang, Yating, et al.
Published: (2025)
by: Wang, Yating, et al.
Published: (2025)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds
by: Zhou, Yunsong, et al.
Published: (2026)
by: Zhou, Yunsong, et al.
Published: (2026)
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
by: Liu, Mingyu, et al.
Published: (2025)
by: Liu, Mingyu, et al.
Published: (2025)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
by: GigaWorld Team, et al.
Published: (2025)
by: GigaWorld Team, et al.
Published: (2025)
Embodiment-Aware Generalist Specialist Distillation for Unified Humanoid Whole-Body Control
by: Peng, Quanquan, et al.
Published: (2026)
by: Peng, Quanquan, et al.
Published: (2026)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
GeoWorld: Geometric World Models
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
(MGS)$^2$-Net: Unifying Micro-Geometric Scale and Macro-Geometric Structure for Cross-View Geo-Localization
by: Li, Minglei, et al.
Published: (2026)
by: Li, Minglei, et al.
Published: (2026)
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
by: Yang, Jiange, et al.
Published: (2025)
by: Yang, Jiange, et al.
Published: (2025)
GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction
by: Chen, Xiao, et al.
Published: (2024)
by: Chen, Xiao, et al.
Published: (2024)
Tac2Real: Reliable and GPU Visuotactile Simulation for Online Reinforcement Learning and Zero-Shot Real-World Deployment
by: Yan, Ningyu, et al.
Published: (2026)
by: Yan, Ningyu, et al.
Published: (2026)
STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System
by: Luo, Zhen, et al.
Published: (2026)
by: Luo, Zhen, et al.
Published: (2026)
DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
by: Zhou, Yang, et al.
Published: (2026)
by: Zhou, Yang, et al.
Published: (2026)
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
by: Liao, Yue, et al.
Published: (2025)
by: Liao, Yue, et al.
Published: (2025)
Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
by: Zhu, Haoyi, et al.
Published: (2024)
by: Zhu, Haoyi, et al.
Published: (2024)
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
by: Chen, Jiahe, et al.
Published: (2026)
by: Chen, Jiahe, et al.
Published: (2026)
Matching Distance and Geometric Distribution Aided Learning Multiview Point Cloud Registration
by: Li, Shiqi, et al.
Published: (2025)
by: Li, Shiqi, et al.
Published: (2025)
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
by: Dong, Yifei, et al.
Published: (2025)
by: Dong, Yifei, et al.
Published: (2025)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes
by: Chen, Xiao, et al.
Published: (2025)
by: Chen, Xiao, et al.
Published: (2025)
LCPR: A Multi-Scale Attention-Based LiDAR-Camera Fusion Network for Place Recognition
by: Zhou, Zijie, et al.
Published: (2023)
by: Zhou, Zijie, et al.
Published: (2023)
PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
by: Zhu, Zhihao, et al.
Published: (2025)
by: Zhu, Zhihao, et al.
Published: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
by: Wang, Kaijun, et al.
Published: (2025)
by: Wang, Kaijun, et al.
Published: (2025)
WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
by: Shang, Yu, et al.
Published: (2026)
by: Shang, Yu, et al.
Published: (2026)
A Unified and General Humanoid Whole-Body Controller for Versatile Locomotion
by: Xue, Yufei, et al.
Published: (2025)
by: Xue, Yufei, et al.
Published: (2025)
UniCon: A Unified System for Efficient Robot Learning Transfers
by: Lin, Yunfeng, et al.
Published: (2026)
by: Lin, Yunfeng, et al.
Published: (2026)
Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming
by: Tong, Baoshun, et al.
Published: (2026)
by: Tong, Baoshun, et al.
Published: (2026)
Demystifying Action Space Design for Robotic Manipulation Policies
by: Feng, Yuchun, et al.
Published: (2026)
by: Feng, Yuchun, et al.
Published: (2026)
Similar Items
-
$π^3$: Permutation-Equivariant Visual Geometry Learning
by: Wang, Yifan, et al.
Published: (2025) -
DeepVerse: 4D Autoregressive Video Generation as a World Model
by: Chen, Junyi, et al.
Published: (2025) -
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
by: Li, Zizun, et al.
Published: (2025) -
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
by: Zhou, Yang, et al.
Published: (2025) -
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
by: Peng, Jiaqi, et al.
Published: (2025)