AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Tianling, Gan, Shengzhe, Gu, Leslie, Li, Yuelei, Zhan, Fangneng, Pfister, Hanspeter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
Advances in Feed‐Forward 3D Reconstruction and View Synthesis: A Survey
by: Jiahui Zhang, et al.
Published: (2026)
by: Jiahui Zhang, et al.
Published: (2026)
RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
by: Yang, Sizhe, et al.
Published: (2026)
by: Yang, Sizhe, et al.
Published: (2026)
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
by: Liu, Zhenyang, et al.
Published: (2026)
by: Liu, Zhenyang, et al.
Published: (2026)
Any4D: Unified Feed-Forward Metric 4D Reconstruction
by: Karhade, Jay, et al.
Published: (2025)
by: Karhade, Jay, et al.
Published: (2025)
EC3R-SLAM: Efficient and Consistent Monocular Dense SLAM with Feed-Forward 3D Reconstruction
by: Hu, Lingxiang, et al.
Published: (2025)
by: Hu, Lingxiang, et al.
Published: (2025)
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
by: Keetha, Nikhil, et al.
Published: (2025)
by: Keetha, Nikhil, et al.
Published: (2025)
LEXI-SG: Monocular 3D Scene Graph Mapping with Room-Guided Feed-Forward Reconstruction
by: Kassab, Christina, et al.
Published: (2026)
by: Kassab, Christina, et al.
Published: (2026)
GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation
by: Zhang, Zijian, et al.
Published: (2026)
by: Zhang, Zijian, et al.
Published: (2026)
Active3D: Active High-Fidelity 3D Reconstruction via Hierarchical Uncertainty Quantification
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
HGS-Planner: Hierarchical Planning Framework for Active Scene Reconstruction Using 3D Gaussian Splatting
by: Xu, Zijun, et al.
Published: (2024)
by: Xu, Zijun, et al.
Published: (2024)
Unifying 2D and 3D Vision-Language Understanding
by: Jain, Ayush, et al.
Published: (2025)
by: Jain, Ayush, et al.
Published: (2025)
3D Vision-tactile Reconstruction from Infrared and Visible Images for Robotic Fine-grained Tactile Perception
by: Lin, Yuankai, et al.
Published: (2025)
by: Lin, Yuankai, et al.
Published: (2025)
ReconDrive: Fast Feed-Forward 4D Gaussian Splatting for Autonomous Driving Scene Reconstruction
by: Yu, Haibao, et al.
Published: (2026)
by: Yu, Haibao, et al.
Published: (2026)
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
by: Gu, Leslie, et al.
Published: (2025)
by: Gu, Leslie, et al.
Published: (2025)
Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation
by: Li, Yuelei, et al.
Published: (2025)
by: Li, Yuelei, et al.
Published: (2025)
Learning Occlusion-aware Decision-making from Agent Interaction via Active Perception
by: Jia, Jie, et al.
Published: (2024)
by: Jia, Jie, et al.
Published: (2024)
UniScale: Unified Scale-Aware 3D Reconstruction for Multi-View Understanding via Prior Injection for Robotic Perception
by: Mahdavian, Mohammad, et al.
Published: (2026)
by: Mahdavian, Mohammad, et al.
Published: (2026)
Stable Language Guidance for Vision-Language-Action Models
by: Zhan, Zhihao, et al.
Published: (2026)
by: Zhan, Zhihao, et al.
Published: (2026)
3D Active Metric-Semantic SLAM
by: Tao, Yuezhan, et al.
Published: (2023)
by: Tao, Yuezhan, et al.
Published: (2023)
Which Reconstruction Model Should a Robot Use? Routing Image-to-3D Models for Cost-Aware Robotic Manipulation
by: Anand, Akash, et al.
Published: (2026)
by: Anand, Akash, et al.
Published: (2026)
Whisker-based Active Tactile Perception for Contour Reconstruction
by: Dang, Yixuan, et al.
Published: (2025)
by: Dang, Yixuan, et al.
Published: (2025)
ScaRF-SLAM: Scale-Consistent Reconstruction with Feed-Forward Models and Classical Visual SLAM
by: Zhang, Yuhao, et al.
Published: (2026)
by: Zhang, Yuhao, et al.
Published: (2026)
Multi-Agent Monocular Dense SLAM With 3D Reconstruction Priors
by: Zhou, Yuchen, et al.
Published: (2025)
by: Zhou, Yuchen, et al.
Published: (2025)
LangFlash: Feed-forward 3D Language Gaussian Splatting from Sparse Unposed Images
by: Liu, Yilong, et al.
Published: (2026)
by: Liu, Yilong, et al.
Published: (2026)
UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception
by: Wang, Ziming
Published: (2026)
by: Wang, Ziming
Published: (2026)
Real-time 3D Semantic Scene Perception for Egocentric Robots with Binocular Vision
by: Nguyen, K., et al.
Published: (2024)
by: Nguyen, K., et al.
Published: (2024)
3DGSNav: Enhancing Vision-Language Model Reasoning for Object Navigation via Active 3D Gaussian Splatting
by: Zheng, Wancai, et al.
Published: (2026)
by: Zheng, Wancai, et al.
Published: (2026)
3D-VLA: A 3D Vision-Language-Action Generative World Model
by: Zhen, Haoyu, et al.
Published: (2024)
by: Zhen, Haoyu, et al.
Published: (2024)
3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation
by: Gkanatsios, Nikolaos, et al.
Published: (2025)
by: Gkanatsios, Nikolaos, et al.
Published: (2025)
EgoAVFlow: Robot Policy Learning with Active Vision from Human Egocentric Videos via 3D Flow
by: Cho, Daesol, et al.
Published: (2026)
by: Cho, Daesol, et al.
Published: (2026)
Vision in Action: Learning Active Perception from Human Demonstrations
by: Xiong, Haoyu, et al.
Published: (2025)
by: Xiong, Haoyu, et al.
Published: (2025)
Real-Time 3D Vision-Language Embedding Mapping
by: Rauch, Christian, et al.
Published: (2025)
by: Rauch, Christian, et al.
Published: (2025)
P$^{3}$Nav: End-to-End Perception, Prediction and Planning for Vision-and-Language Navigation
by: Li, Tianfu, et al.
Published: (2026)
by: Li, Tianfu, et al.
Published: (2026)
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
by: Yang, Jianing, et al.
Published: (2025)
by: Yang, Jianing, et al.
Published: (2025)
Physical Priors Augmented Event-Based 3D Reconstruction
by: Wang, Jiaxu, et al.
Published: (2024)
by: Wang, Jiaxu, et al.
Published: (2024)
Rejecting Outliers in 2D-3D Point Correspondences from 2D Forward-Looking Sonar Observations
by: Su, Jiayi, et al.
Published: (2025)
by: Su, Jiayi, et al.
Published: (2025)
Conflict-Aware Active Perception and Control in 3D Gaussian Splatting Fields via Control Barrier Functions
by: Khass, Amirhossein Mollaei, et al.
Published: (2026)
by: Khass, Amirhossein Mollaei, et al.
Published: (2026)
Similar Items
-
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
by: Liu, Yifan, et al.
Published: (2025) -
Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
by: Zhang, Jiahui, et al.
Published: (2025) -
Advances in Feed‐Forward 3D Reconstruction and View Synthesis: A Survey
by: Jiahui Zhang, et al.
Published: (2026) -
RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph
by: Liu, Yifan, et al.
Published: (2025) -
Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
by: Yang, Sizhe, et al.
Published: (2026)