Embodied Navigation with Auxiliary Task of Action Description Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Kondoh, Haru, Kanezaki, Asako |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation
by: Sun, Leyuan, et al.
Published: (2024)
by: Sun, Leyuan, et al.
Published: (2024)
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
by: Yajima, Masaru, et al.
Published: (2025)
by: Yajima, Masaru, et al.
Published: (2025)
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
by: Zhang, Jiazhao, et al.
Published: (2024)
by: Zhang, Jiazhao, et al.
Published: (2024)
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
by: Liu, Dayong, et al.
Published: (2025)
by: Liu, Dayong, et al.
Published: (2025)
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
by: Ding, Hongyu, et al.
Published: (2026)
by: Ding, Hongyu, et al.
Published: (2026)
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
by: Li, Yixuan, et al.
Published: (2025)
by: Li, Yixuan, et al.
Published: (2025)
FlowLoss: Dynamic Flow-Conditioned Loss Strategy for Video Diffusion Models
by: Wu, Kuanting, et al.
Published: (2025)
by: Wu, Kuanting, et al.
Published: (2025)
OP-Align: Object-level and Part-level Alignment for Self-supervised Category-level Articulated Object Pose Estimation
by: Che, Yuchen, et al.
Published: (2024)
by: Che, Yuchen, et al.
Published: (2024)
RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action
by: Sun, Xiaoquan, et al.
Published: (2025)
by: Sun, Xiaoquan, et al.
Published: (2025)
Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
by: Sohn, Tin Stribor, et al.
Published: (2025)
by: Sohn, Tin Stribor, et al.
Published: (2025)
Personalized Embodied Navigation for Portable Object Finding
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
by: Zou, Ding, et al.
Published: (2025)
by: Zou, Ding, et al.
Published: (2025)
Nav-R1: Reasoning and Navigation in Embodied Scenes
by: Liu, Qingxiang, et al.
Published: (2025)
by: Liu, Qingxiang, et al.
Published: (2025)
Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation
by: Chen, Meng, et al.
Published: (2024)
by: Chen, Meng, et al.
Published: (2024)
EnerVerse-AC: Envisioning Embodied Environments with Action Condition
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
Embodied Navigation at the Art Gallery
by: Bigazzi, Roberto, et al.
Published: (2022)
by: Bigazzi, Roberto, et al.
Published: (2022)
COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation
by: Che, Yuchen, et al.
Published: (2026)
by: Che, Yuchen, et al.
Published: (2026)
Zero-shot Degree of Ill-posedness Estimation for Active Small Object Change Detection
by: Takeda, Koji, et al.
Published: (2024)
by: Takeda, Koji, et al.
Published: (2024)
From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding
by: Halacheva, Anna-Maria, et al.
Published: (2025)
by: Halacheva, Anna-Maria, et al.
Published: (2025)
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
by: Lang, Xiaolei, et al.
Published: (2026)
by: Lang, Xiaolei, et al.
Published: (2026)
RoboAgent: Chaining Basic Capabilities for Embodied Task Planning
by: Xu, Peiran, et al.
Published: (2026)
by: Xu, Peiran, et al.
Published: (2026)
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
by: Liu, Xinhao, et al.
Published: (2024)
by: Liu, Xinhao, et al.
Published: (2024)
VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory
by: Wang, Shaoan, et al.
Published: (2026)
by: Wang, Shaoan, et al.
Published: (2026)
Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions
by: Galliena, Tommaso, et al.
Published: (2025)
by: Galliena, Tommaso, et al.
Published: (2025)
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
by: Hung, Chia-Yu, et al.
Published: (2025)
by: Hung, Chia-Yu, et al.
Published: (2025)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
by: Ling, Yiran, et al.
Published: (2026)
by: Ling, Yiran, et al.
Published: (2026)
VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
by: Lin, Sihao, et al.
Published: (2025)
by: Lin, Sihao, et al.
Published: (2025)
ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation
by: Ao, Yuzhuo, et al.
Published: (2026)
by: Ao, Yuzhuo, et al.
Published: (2026)
World Action Models: The Next Frontier in Embodied AI
by: Wang, Siyin, et al.
Published: (2026)
by: Wang, Siyin, et al.
Published: (2026)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
by: Zhang, Qiyao, et al.
Published: (2026)
by: Zhang, Qiyao, et al.
Published: (2026)
SABER: A Scalable Action-Based Embodied Dataset for Real-World VLA Adaptation
by: Menga, Narsimha, et al.
Published: (2026)
by: Menga, Narsimha, et al.
Published: (2026)
Expand Your SCOPE: Semantic Cognition over Potential-Based Exploration for Embodied Visual Navigation
by: Wang, Ningnan, et al.
Published: (2025)
by: Wang, Ningnan, et al.
Published: (2025)
Bridging the Indoor-Outdoor Gap: Vision-Centric Instruction-Guided Embodied Navigation for the Last Meters
by: Zhao, Yuxiang, et al.
Published: (2026)
by: Zhao, Yuxiang, et al.
Published: (2026)
OctoNav: Towards Generalist Embodied Navigation
by: Gao, Chen, et al.
Published: (2025)
by: Gao, Chen, et al.
Published: (2025)
SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning
by: Yang, Yuyuan, et al.
Published: (2026)
by: Yang, Yuyuan, et al.
Published: (2026)
Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning
by: Ganai, Milan, et al.
Published: (2026)
by: Ganai, Milan, et al.
Published: (2026)
DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
by: Fang, Zhen, et al.
Published: (2025)
by: Fang, Zhen, et al.
Published: (2025)
EmbodiedSplat: Personalized Real-to-Sim-to-Real Navigation with Gaussian Splats from a Mobile Device
by: Chhablani, Gunjan, et al.
Published: (2025)
by: Chhablani, Gunjan, et al.
Published: (2025)
MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
by: Huang, Xun, et al.
Published: (2025)
by: Huang, Xun, et al.
Published: (2025)
Similar Items
-
Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation
by: Sun, Leyuan, et al.
Published: (2024) -
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
by: Yajima, Masaru, et al.
Published: (2025) -
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
by: Zhang, Jiazhao, et al.
Published: (2024) -
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
by: Liu, Dayong, et al.
Published: (2025) -
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
by: Ding, Hongyu, et al.
Published: (2026)