Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Weichen, Tang, Peizhi, Zeng, Xin, Man, Fanhang, Yu, Shiquan, Dai, Zichao, Zhao, Baining, Chen, Hongjin, Shang, Yu, Wu, Wei, Gao, Chen, Chen, Xinlei, Wang, Xin, Li, Yong, Zhu, Wenwu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
AirScape: An Aerial Generative World Model with Motion Controllability
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
VAEER: Visual Attention-Inspired Emotion Elicitation Reasoning
by: Man, Fanhang, et al.
Published: (2025)
by: Man, Fanhang, et al.
Published: (2025)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
by: Zhao, Baining, et al.
Published: (2026)
by: Zhao, Baining, et al.
Published: (2026)
Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
Context-Aware Sentiment Forecasting via LLM-based Multi-Perspective Role-Playing Agents
by: Man, Fanhang, et al.
Published: (2025)
by: Man, Fanhang, et al.
Published: (2025)
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment
by: Gao, Chen, et al.
Published: (2024)
by: Gao, Chen, et al.
Published: (2024)
WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models
by: Chen, Hongjin, et al.
Published: (2026)
by: Chen, Hongjin, et al.
Published: (2026)
iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
by: Fang, Jianjie, et al.
Published: (2026)
by: Fang, Jianjie, et al.
Published: (2026)
Understanding and Evaluating Hallucinations in 3D Visual Language Models
by: Peng, Ruiying, et al.
Published: (2025)
by: Peng, Ruiying, et al.
Published: (2025)
LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE
by: Shang, Yu, et al.
Published: (2025)
by: Shang, Yu, et al.
Published: (2025)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace
by: Zhao, Baining, et al.
Published: (2026)
by: Zhao, Baining, et al.
Published: (2026)
U2UData+: A Scalable Swarm UAVs Autonomous Flight Dataset for Embodied Long-horizon Tasks
by: Feng, Tongtong, et al.
Published: (2025)
by: Feng, Tongtong, et al.
Published: (2025)
RoboScape: Physics-informed Embodied World Model
by: Shang, Yu, et al.
Published: (2025)
by: Shang, Yu, et al.
Published: (2025)
Digital Twin-Empowered Task Assignment in Aerial MEC Network: A Resource Coalition Cooperation Approach with Generative Model
by: Tang, Xin, et al.
Published: (2024)
by: Tang, Xin, et al.
Published: (2024)
The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space
by: Yang, Jinrong, et al.
Published: (2025)
by: Yang, Jinrong, et al.
Published: (2025)
Embodied AI: From LLMs to World Models
by: Feng, Tongtong, et al.
Published: (2025)
by: Feng, Tongtong, et al.
Published: (2025)
QUIDS: Quality-informed Incentive-driven Multi-agent Dispatching System for Mobile Crowdsensing
by: Zhou, Nan, et al.
Published: (2025)
by: Zhou, Nan, et al.
Published: (2025)
Multi-sentence Video Grounding for Long Video Generation
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
Towards Precise Intent-Aligned VLA Aerial Navigation via Expert-Guided GRPO
by: Chen, Tianyang, et al.
Published: (2026)
by: Chen, Tianyang, et al.
Published: (2026)
Semantic Audio-Visual Navigation in Continuous Environments
by: Zeng, Yichen, et al.
Published: (2026)
by: Zeng, Yichen, et al.
Published: (2026)
RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL
by: Tang, Yinzhou, et al.
Published: (2025)
by: Tang, Yinzhou, et al.
Published: (2025)
Efficient SLAM Algorithm Based on Quaternions for Real‐Time Navigation of Unmanned Aerial Vehicles
by: Xin Wang, et al.
Published: (2026)
by: Xin Wang, et al.
Published: (2026)
MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation
by: Yu, Yangcheng, et al.
Published: (2025)
by: Yu, Yangcheng, et al.
Published: (2025)
Language-Conditioned World Modeling for Visual Navigation
by: Dong, Yifei, et al.
Published: (2026)
by: Dong, Yifei, et al.
Published: (2026)
CCMamba: Topologically-Informed Selective State-Space Networks on Combinatorial Complexes for Higher-Order Graph Learning
by: Chen, Jiawen, et al.
Published: (2026)
by: Chen, Jiawen, et al.
Published: (2026)
WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
by: Chen, Delong, et al.
Published: (2025)
by: Chen, Delong, et al.
Published: (2025)
Visual-Geometry GP-based Navigable Space for Autonomous Navigation
by: Ali, Mahmoud, et al.
Published: (2024)
by: Ali, Mahmoud, et al.
Published: (2024)
CCA: Collaborative Competitive Agents for Image Editing
by: Hang, Tiankai, et al.
Published: (2024)
by: Hang, Tiankai, et al.
Published: (2024)
A Simple Aerial Detection Baseline of Multimodal Language Models
by: Li, Qingyun, et al.
Published: (2025)
by: Li, Qingyun, et al.
Published: (2025)
Distributed fast F‐T control for UAV formation in the presence of unknown input disturbances
by: Hongjin Liao, et al.
Published: (2024)
by: Hongjin Liao, et al.
Published: (2024)
$\imath$Hopf algebras associated with self-dual Hopf algebras
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
by: Shang, Yu, et al.
Published: (2026)
by: Shang, Yu, et al.
Published: (2026)
LVIC: Multi-modality segmentation by Lifting Visual Info as Cue
by: Dong, Zichao, et al.
Published: (2024)
by: Dong, Zichao, et al.
Published: (2024)
HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation
by: Zhang, Conglang, et al.
Published: (2026)
by: Zhang, Conglang, et al.
Published: (2026)
Long‐lasting UV‐blocking Mechanism of Lignin: Origin and Stabilization of Semiquinone Radicals
by: Yu Fu, et al.
Published: (2024)
by: Yu Fu, et al.
Published: (2024)
Similar Items
-
CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory
by: Zhang, Weichen, et al.
Published: (2025) -
AirScape: An Aerial Generative World Model with Motion Controllability
by: Zhao, Baining, et al.
Published: (2025) -
VAEER: Visual Attention-Inspired Emotion Elicitation Reasoning
by: Man, Fanhang, et al.
Published: (2025) -
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
by: Zhao, Baining, et al.
Published: (2026) -
Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
by: Zhao, Baining, et al.
Published: (2025)