AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Zhaofeng, Zhou, Sifan, Zhang, Qinbo, Xu, Rongtao, Su, Qi, Mendez-Mendz, Jorge, Liang, Ci-Jyun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sample-Efficient Robot Skill Learning for Construction Tasks: Benchmarking Hierarchical Reinforcement Learning and Vision-Language-Action VLA Model
por: Hu, Zhaofeng, et al.
Publicado: (2025)
por: Hu, Zhaofeng, et al.
Publicado: (2025)
MVCTrack: Boosting 3D Point Cloud Tracking via Multimodal-Guided Virtual Cues
por: Hu, Zhaofeng, et al.
Publicado: (2024)
por: Hu, Zhaofeng, et al.
Publicado: (2024)
Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations
por: Spieler, Jonathan, et al.
Publicado: (2026)
por: Spieler, Jonathan, et al.
Publicado: (2026)
Slot-Level Robotic Placement via Visual Imitation from Single Human Video
por: Shan, Dandan, et al.
Publicado: (2025)
por: Shan, Dandan, et al.
Publicado: (2025)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
por: Chen, Kehan, et al.
Publicado: (2024)
por: Chen, Kehan, et al.
Publicado: (2024)
AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models
por: Heo, Hyeongjun, et al.
Publicado: (2026)
por: Heo, Hyeongjun, et al.
Publicado: (2026)
AniTrack: A Power-Efficient, Time-Slotted and Robust UWB Localization System for Animal Tracking in a Controlled Setting
por: Luder, Victor, et al.
Publicado: (2025)
por: Luder, Victor, et al.
Publicado: (2025)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
por: Liu, Yu, et al.
Publicado: (2024)
por: Liu, Yu, et al.
Publicado: (2024)
Few-Shot Vision-Language Action-Incremental Policy Learning
por: Song, Mingchen, et al.
Publicado: (2025)
por: Song, Mingchen, et al.
Publicado: (2025)
GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
por: Abouzeid, Ali, et al.
Publicado: (2025)
por: Abouzeid, Ali, et al.
Publicado: (2025)
STORM: Slot-based Task-aware Object-centric Representation for robotic Manipulation
por: Chapin, Alexandre, et al.
Publicado: (2026)
por: Chapin, Alexandre, et al.
Publicado: (2026)
VIPS-Odom: Visual-Inertial Odometry Tightly-coupled with Parking Slots for Autonomous Parking
por: Jiang, Xuefeng, et al.
Publicado: (2024)
por: Jiang, Xuefeng, et al.
Publicado: (2024)
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
por: Guo, Minghao, et al.
Publicado: (2025)
por: Guo, Minghao, et al.
Publicado: (2025)
Slot-hopping Enabled Loiter Guidance and Automation for Fixed-wing UAV Corridors
por: J, Pradeep, et al.
Publicado: (2026)
por: J, Pradeep, et al.
Publicado: (2026)
Manipulation Planning for Construction Activities with Repetitive Tasks
por: Liu, Wangyi, et al.
Publicado: (2026)
por: Liu, Wangyi, et al.
Publicado: (2026)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
por: Hanyu, Taisei, et al.
Publicado: (2025)
por: Hanyu, Taisei, et al.
Publicado: (2025)
Temporally Consistent Object-Centric Learning by Contrasting Slots
por: Manasyan, Anna, et al.
Publicado: (2024)
por: Manasyan, Anna, et al.
Publicado: (2024)
RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models
por: Liufu, Weijia, et al.
Publicado: (2026)
por: Liufu, Weijia, et al.
Publicado: (2026)
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
por: Yang, Rushuai, et al.
Publicado: (2026)
por: Yang, Rushuai, et al.
Publicado: (2026)
LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
por: Ding, Hongyu, et al.
Publicado: (2025)
por: Ding, Hongyu, et al.
Publicado: (2025)
CL-CoTNav: Closed-Loop Hierarchical Chain-of-Thought for Zero-Shot Object-Goal Navigation with Vision-Language Models
por: Cai, Yuxin, et al.
Publicado: (2025)
por: Cai, Yuxin, et al.
Publicado: (2025)
Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera
por: Guo, Yuliang, et al.
Publicado: (2025)
por: Guo, Yuliang, et al.
Publicado: (2025)
MoWe: Motion Observation for Wind Estimation of Sailing Robots
por: Qinbo Sun, et al.
Publicado: (2025)
por: Qinbo Sun, et al.
Publicado: (2025)
High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions
por: Cao, Yongpeng, et al.
Publicado: (2026)
por: Cao, Yongpeng, et al.
Publicado: (2026)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
por: Villar-Corrales, Angel, et al.
Publicado: (2025)
por: Villar-Corrales, Angel, et al.
Publicado: (2025)
SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from Pixels
por: Mosbach, Malte, et al.
Publicado: (2024)
por: Mosbach, Malte, et al.
Publicado: (2024)
GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance
por: Yuan, Shuaihang, et al.
Publicado: (2024)
por: Yuan, Shuaihang, et al.
Publicado: (2024)
Hyper-GoalNet: Goal-Conditioned Manipulation Policy Learning with HyperNetworks
por: Zhou, Pei, et al.
Publicado: (2025)
por: Zhou, Pei, et al.
Publicado: (2025)
SemNav: A Model-Based Planner for Zero-Shot Object Goal Navigation Using Vision-Foundation Models
por: Debnath, Arnab, et al.
Publicado: (2025)
por: Debnath, Arnab, et al.
Publicado: (2025)
Multi-Floor Zero-Shot Object Navigation Policy
por: Zhang, Lingfeng, et al.
Publicado: (2024)
por: Zhang, Lingfeng, et al.
Publicado: (2024)
Explicit Stair Geometry Conditioning for Robust Humanoid Locomotion
por: Zhang, Jianguo, et al.
Publicado: (2026)
por: Zhang, Jianguo, et al.
Publicado: (2026)
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
por: Zuo, Kuangji, et al.
Publicado: (2026)
por: Zuo, Kuangji, et al.
Publicado: (2026)
GraSP-STL: A Graph-Based Framework for Zero-Shot Signal Temporal Logic Planning via Offline Goal-Conditioned Reinforcement Learning
por: Hou, Ancheng, et al.
Publicado: (2026)
por: Hou, Ancheng, et al.
Publicado: (2026)
Exploring the Reliability of Foundation Model-Based Frontier Selection in Zero-Shot Object Goal Navigation
por: Yuan, Shuaihang, et al.
Publicado: (2024)
por: Yuan, Shuaihang, et al.
Publicado: (2024)
InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
por: Yan, Yu, et al.
Publicado: (2024)
por: Yan, Yu, et al.
Publicado: (2024)
FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation
por: Tan, Mingao, et al.
Publicado: (2026)
por: Tan, Mingao, et al.
Publicado: (2026)
Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration
por: Gong, Zeying, et al.
Publicado: (2025)
por: Gong, Zeying, et al.
Publicado: (2025)
GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy
por: Li, Peiyan, et al.
Publicado: (2024)
por: Li, Peiyan, et al.
Publicado: (2024)
DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
por: Fekri, Pedram, et al.
Publicado: (2025)
por: Fekri, Pedram, et al.
Publicado: (2025)
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model
por: Guo, Yanjiang, et al.
Publicado: (2026)
por: Guo, Yanjiang, et al.
Publicado: (2026)
Ejemplares similares
-
Sample-Efficient Robot Skill Learning for Construction Tasks: Benchmarking Hierarchical Reinforcement Learning and Vision-Language-Action VLA Model
por: Hu, Zhaofeng, et al.
Publicado: (2025) -
MVCTrack: Boosting 3D Point Cloud Tracking via Multimodal-Guided Virtual Cues
por: Hu, Zhaofeng, et al.
Publicado: (2024) -
Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations
por: Spieler, Jonathan, et al.
Publicado: (2026) -
Slot-Level Robotic Placement via Visual Imitation from Single Human Video
por: Shan, Dandan, et al.
Publicado: (2025) -
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
por: Chen, Kehan, et al.
Publicado: (2024)