C2F-Space: Coarse-to-Fine Space Grounding for Spatial Instructions using Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Oh, Nayoung, Kim, Dohyun, Bang, Junhyeong, Paul, Rohan, Park, Daehyung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LINGO-Space: Language-Conditioned Incremental Grounding for Space
by: Kim, Dohyun, et al.
Published: (2024)
by: Kim, Dohyun, et al.
Published: (2024)
DiSPo: Diffusion-SSM based Policy Learning for Coarse-to-Fine Action Discretization
by: Oh, Nayoung, et al.
Published: (2024)
by: Oh, Nayoung, et al.
Published: (2024)
A Visuo-Tactile Data Collection System with Haptic Feedback for Coarse-to-Fine Imitation Learning
by: Kim, Yeseung, et al.
Published: (2026)
by: Kim, Yeseung, et al.
Published: (2026)
A Survey on Integration of Large Language Models with Intelligent Robots
by: Kim, Yeseung, et al.
Published: (2024)
by: Kim, Yeseung, et al.
Published: (2024)
SGGNet$^2$: Speech-Scene Graph Grounding Network for Speech-guided Navigation
by: Kim, Dohyun, et al.
Published: (2023)
by: Kim, Dohyun, et al.
Published: (2023)
A Reachability Tree-Based Algorithm for Robot Task and Motion Planning
by: Kim, Kanghyun, et al.
Published: (2023)
by: Kim, Kanghyun, et al.
Published: (2023)
Graph-based 3D Collision-distance Estimation Network with Probabilistic Graph Rewiring
by: Song, Minjae, et al.
Published: (2023)
by: Song, Minjae, et al.
Published: (2023)
ILCL: Inverse Logic-Constraint Learning from Temporally Constrained Demonstrations
by: Cho, Minwoo, et al.
Published: (2025)
by: Cho, Minwoo, et al.
Published: (2025)
G$^{2}$TR: Generalized Grounded Temporal Reasoning for Robot Instruction Following by Combining Large Pre-trained Models
by: Arora, Riya, et al.
Published: (2024)
by: Arora, Riya, et al.
Published: (2024)
Reinforcement Learning-based Fault-Tolerant Control for Quadrotor with Online Transformer Adaptation
by: Kim, Dohyun, et al.
Published: (2025)
by: Kim, Dohyun, et al.
Published: (2025)
Implicit Neural-Representation Learning for Elastic Deformable-Object Manipulations
by: Song, Minseok, et al.
Published: (2025)
by: Song, Minseok, et al.
Published: (2025)
Learning-based Initialization of Trajectory Optimization for Path-following Problems of Redundant Manipulators
by: Yoon, Minsung, et al.
Published: (2026)
by: Yoon, Minsung, et al.
Published: (2026)
RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
by: Koo, Jiyeon, et al.
Published: (2025)
by: Koo, Jiyeon, et al.
Published: (2025)
Event-Grounded Sparse Autoencoders for Vision-Language-Action Policies
by: Jin, Xinchen, et al.
Published: (2026)
by: Jin, Xinchen, et al.
Published: (2026)
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
by: Tu, Ruisen, et al.
Published: (2026)
by: Tu, Ruisen, et al.
Published: (2026)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
by: Hu, Jianshu, et al.
Published: (2025)
by: Hu, Jianshu, et al.
Published: (2025)
C2F-TP: A Coarse-to-Fine Denoising Framework for Uncertainty-Aware Trajectory Prediction
by: Wang, Zichen, et al.
Published: (2024)
by: Wang, Zichen, et al.
Published: (2024)
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
by: Hu, Xintong, et al.
Published: (2026)
by: Hu, Xintong, et al.
Published: (2026)
Coarse-to-Fine Compositional Diffusion for Long-Horizon Planning
by: Park, Byoungwoo, et al.
Published: (2026)
by: Park, Byoungwoo, et al.
Published: (2026)
TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
by: Zhang, Zongzheng, et al.
Published: (2025)
by: Zhang, Zongzheng, et al.
Published: (2025)
SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model
by: Zhu, Shaoting, et al.
Published: (2024)
by: Zhu, Shaoting, et al.
Published: (2024)
Dynamic Control Barrier Function Regulation with Vision-Language Models for Safe, Adaptive, and Realtime Visual Navigation
by: Chen, Jeffrey, et al.
Published: (2026)
by: Chen, Jeffrey, et al.
Published: (2026)
SuReNav: Superpixel Graph-based Constraint Relaxation for Navigation in Over-constrained Environments
by: Koh, Keonyoung, et al.
Published: (2026)
by: Koh, Keonyoung, et al.
Published: (2026)
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
by: Shi, Zhan, et al.
Published: (2025)
by: Shi, Zhan, et al.
Published: (2025)
Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces
by: Kåsene, Vebjørn Haug, et al.
Published: (2025)
by: Kåsene, Vebjørn Haug, et al.
Published: (2025)
Space-LLaVA: a Vision-Language Model Adapted to Extraterrestrial Applications
by: Foutter, Matthew, et al.
Published: (2024)
by: Foutter, Matthew, et al.
Published: (2024)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
by: Zhang, Zhengshen, et al.
Published: (2025)
by: Zhang, Zhengshen, et al.
Published: (2025)
Premover: Fast Vision-Language-Action Control by Acting Before Instructions Are Complete
by: Park, Joonha, et al.
Published: (2026)
by: Park, Joonha, et al.
Published: (2026)
Haptic-Based Bilateral Teleoperation of Aerial Manipulator for Extracting Wedged Object with Compensation of Human Reaction Time
by: Byun, Jeonghyun, et al.
Published: (2024)
by: Byun, Jeonghyun, et al.
Published: (2024)
CaFe-TeleVision: A Coarse-to-Fine Teleoperation System with Immersive Situated Visualization for Enhanced Ergonomics
by: Tang, Zixin, et al.
Published: (2025)
by: Tang, Zixin, et al.
Published: (2025)
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
by: Huang, Helong, et al.
Published: (2025)
by: Huang, Helong, et al.
Published: (2025)
NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Adaptive Capacity Allocation for Vision Language Action Fine-tuning
by: Kim, Donghoon, et al.
Published: (2026)
by: Kim, Donghoon, et al.
Published: (2026)
AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
by: Takagi, Yusuke, et al.
Published: (2026)
by: Takagi, Yusuke, et al.
Published: (2026)
Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment
by: Wulff, Theodor, et al.
Published: (2026)
by: Wulff, Theodor, et al.
Published: (2026)
Instantaneous Planning, Control and Safety for Navigation in Unknown Underwater Spaces
by: Karthik, Veejay, et al.
Published: (2026)
by: Karthik, Veejay, et al.
Published: (2026)
TACS-Graphs: Traversability-Aware Consistent Scene Graphs for Ground Robot Localization and Mapping
by: Kim, Jeewon, et al.
Published: (2025)
by: Kim, Jeewon, et al.
Published: (2025)
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
by: Glossop, Catherine, et al.
Published: (2025)
by: Glossop, Catherine, et al.
Published: (2025)
Similar Items
-
LINGO-Space: Language-Conditioned Incremental Grounding for Space
by: Kim, Dohyun, et al.
Published: (2024) -
DiSPo: Diffusion-SSM based Policy Learning for Coarse-to-Fine Action Discretization
by: Oh, Nayoung, et al.
Published: (2024) -
A Visuo-Tactile Data Collection System with Haptic Feedback for Coarse-to-Fine Imitation Learning
by: Kim, Yeseung, et al.
Published: (2026) -
A Survey on Integration of Large Language Models with Intelligent Robots
by: Kim, Yeseung, et al.
Published: (2024) -
SGGNet$^2$: Speech-Scene Graph Grounding Network for Speech-guided Navigation
by: Kim, Dohyun, et al.
Published: (2023)