AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Zhaofeng, Zhou, Sifan, Zhang, Qinbo, Xu, Rongtao, Su, Qi, Mendez-Mendz, Jorge, Liang, Ci-Jyun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sample-Efficient Robot Skill Learning for Construction Tasks: Benchmarking Hierarchical Reinforcement Learning and Vision-Language-Action VLA Model
by: Hu, Zhaofeng, et al.
Published: (2025)
by: Hu, Zhaofeng, et al.
Published: (2025)
MVCTrack: Boosting 3D Point Cloud Tracking via Multimodal-Guided Virtual Cues
by: Hu, Zhaofeng, et al.
Published: (2024)
by: Hu, Zhaofeng, et al.
Published: (2024)
Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations
by: Spieler, Jonathan, et al.
Published: (2026)
by: Spieler, Jonathan, et al.
Published: (2026)
Slot-Level Robotic Placement via Visual Imitation from Single Human Video
by: Shan, Dandan, et al.
Published: (2025)
by: Shan, Dandan, et al.
Published: (2025)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
by: Chen, Kehan, et al.
Published: (2024)
by: Chen, Kehan, et al.
Published: (2024)
AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models
by: Heo, Hyeongjun, et al.
Published: (2026)
by: Heo, Hyeongjun, et al.
Published: (2026)
AniTrack: A Power-Efficient, Time-Slotted and Robust UWB Localization System for Animal Tracking in a Controlled Setting
by: Luder, Victor, et al.
Published: (2025)
by: Luder, Victor, et al.
Published: (2025)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Few-Shot Vision-Language Action-Incremental Policy Learning
by: Song, Mingchen, et al.
Published: (2025)
by: Song, Mingchen, et al.
Published: (2025)
GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
by: Abouzeid, Ali, et al.
Published: (2025)
by: Abouzeid, Ali, et al.
Published: (2025)
STORM: Slot-based Task-aware Object-centric Representation for robotic Manipulation
by: Chapin, Alexandre, et al.
Published: (2026)
by: Chapin, Alexandre, et al.
Published: (2026)
VIPS-Odom: Visual-Inertial Odometry Tightly-coupled with Parking Slots for Autonomous Parking
by: Jiang, Xuefeng, et al.
Published: (2024)
by: Jiang, Xuefeng, et al.
Published: (2024)
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
by: Guo, Minghao, et al.
Published: (2025)
by: Guo, Minghao, et al.
Published: (2025)
Slot-hopping Enabled Loiter Guidance and Automation for Fixed-wing UAV Corridors
by: J, Pradeep, et al.
Published: (2026)
by: J, Pradeep, et al.
Published: (2026)
Manipulation Planning for Construction Activities with Repetitive Tasks
by: Liu, Wangyi, et al.
Published: (2026)
by: Liu, Wangyi, et al.
Published: (2026)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
by: Hanyu, Taisei, et al.
Published: (2025)
by: Hanyu, Taisei, et al.
Published: (2025)
Temporally Consistent Object-Centric Learning by Contrasting Slots
by: Manasyan, Anna, et al.
Published: (2024)
by: Manasyan, Anna, et al.
Published: (2024)
RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models
by: Liufu, Weijia, et al.
Published: (2026)
by: Liufu, Weijia, et al.
Published: (2026)
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
by: Yang, Rushuai, et al.
Published: (2026)
by: Yang, Rushuai, et al.
Published: (2026)
LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
by: Ding, Hongyu, et al.
Published: (2025)
by: Ding, Hongyu, et al.
Published: (2025)
CL-CoTNav: Closed-Loop Hierarchical Chain-of-Thought for Zero-Shot Object-Goal Navigation with Vision-Language Models
by: Cai, Yuxin, et al.
Published: (2025)
by: Cai, Yuxin, et al.
Published: (2025)
Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera
by: Guo, Yuliang, et al.
Published: (2025)
by: Guo, Yuliang, et al.
Published: (2025)
MoWe: Motion Observation for Wind Estimation of Sailing Robots
by: Qinbo Sun, et al.
Published: (2025)
by: Qinbo Sun, et al.
Published: (2025)
High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions
by: Cao, Yongpeng, et al.
Published: (2026)
by: Cao, Yongpeng, et al.
Published: (2026)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
by: Villar-Corrales, Angel, et al.
Published: (2025)
by: Villar-Corrales, Angel, et al.
Published: (2025)
SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from Pixels
by: Mosbach, Malte, et al.
Published: (2024)
by: Mosbach, Malte, et al.
Published: (2024)
GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance
by: Yuan, Shuaihang, et al.
Published: (2024)
by: Yuan, Shuaihang, et al.
Published: (2024)
Hyper-GoalNet: Goal-Conditioned Manipulation Policy Learning with HyperNetworks
by: Zhou, Pei, et al.
Published: (2025)
by: Zhou, Pei, et al.
Published: (2025)
SemNav: A Model-Based Planner for Zero-Shot Object Goal Navigation Using Vision-Foundation Models
by: Debnath, Arnab, et al.
Published: (2025)
by: Debnath, Arnab, et al.
Published: (2025)
Multi-Floor Zero-Shot Object Navigation Policy
by: Zhang, Lingfeng, et al.
Published: (2024)
by: Zhang, Lingfeng, et al.
Published: (2024)
Explicit Stair Geometry Conditioning for Robust Humanoid Locomotion
by: Zhang, Jianguo, et al.
Published: (2026)
by: Zhang, Jianguo, et al.
Published: (2026)
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
by: Zuo, Kuangji, et al.
Published: (2026)
by: Zuo, Kuangji, et al.
Published: (2026)
GraSP-STL: A Graph-Based Framework for Zero-Shot Signal Temporal Logic Planning via Offline Goal-Conditioned Reinforcement Learning
by: Hou, Ancheng, et al.
Published: (2026)
by: Hou, Ancheng, et al.
Published: (2026)
Exploring the Reliability of Foundation Model-Based Frontier Selection in Zero-Shot Object Goal Navigation
by: Yuan, Shuaihang, et al.
Published: (2024)
by: Yuan, Shuaihang, et al.
Published: (2024)
InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
by: Yan, Yu, et al.
Published: (2024)
by: Yan, Yu, et al.
Published: (2024)
FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation
by: Tan, Mingao, et al.
Published: (2026)
by: Tan, Mingao, et al.
Published: (2026)
Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration
by: Gong, Zeying, et al.
Published: (2025)
by: Gong, Zeying, et al.
Published: (2025)
GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy
by: Li, Peiyan, et al.
Published: (2024)
by: Li, Peiyan, et al.
Published: (2024)
DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
by: Fekri, Pedram, et al.
Published: (2025)
by: Fekri, Pedram, et al.
Published: (2025)
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model
by: Guo, Yanjiang, et al.
Published: (2026)
by: Guo, Yanjiang, et al.
Published: (2026)
Similar Items
-
Sample-Efficient Robot Skill Learning for Construction Tasks: Benchmarking Hierarchical Reinforcement Learning and Vision-Language-Action VLA Model
by: Hu, Zhaofeng, et al.
Published: (2025) -
MVCTrack: Boosting 3D Point Cloud Tracking via Multimodal-Guided Virtual Cues
by: Hu, Zhaofeng, et al.
Published: (2024) -
Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations
by: Spieler, Jonathan, et al.
Published: (2026) -
Slot-Level Robotic Placement via Visual Imitation from Single Human Video
by: Shan, Dandan, et al.
Published: (2025) -
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
by: Chen, Kehan, et al.
Published: (2024)