PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Nasiriany, Soroush, Xia, Fei, Yu, Wenhao, Xiao, Ted, Liang, Jacky, Dasgupta, Ishita, Xie, Annie, Driess, Danny, Wahid, Ayzaan, Xu, Zhuo, Vuong, Quan, Zhang, Tingnan, Lee, Tsang-Wei Edward, Lee, Kuang-Huei, Xu, Peng, Kirmani, Sean, Zhu, Yuke, Zeng, Andy, Hausman, Karol, Heess, Nicolas, Finn, Chelsea, Levine, Sergey, Ichter, Brian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
by: Nasiriany, Soroush, et al.
Published: (2026)
by: Nasiriany, Soroush, et al.
Published: (2026)
Vision Language Models are In-Context Value Learners
by: Ma, Yecheng Jason, et al.
Published: (2024)
by: Ma, Yecheng Jason, et al.
Published: (2024)
ALOHA Unleashed: A Simple Recipe for Robot Dexterity
by: Zhao, Tony Z., et al.
Published: (2024)
by: Zhao, Tony Z., et al.
Published: (2024)
PRIME: Scaffolding Manipulation Tasks with Behavior Primitives for Data-Efficient Imitation Learning
by: Gao, Tian, et al.
Published: (2024)
by: Gao, Tian, et al.
Published: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs
by: Chiang, Hao-Tien Lewis, et al.
Published: (2024)
by: Chiang, Hao-Tien Lewis, et al.
Published: (2024)
RT-Sketch: Goal-Conditioned Imitation Learning from Hand-Drawn Sketches
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers
by: Jain, Vidhi, et al.
Published: (2024)
by: Jain, Vidhi, et al.
Published: (2024)
Learning to Learn Faster from Human Feedback with Language Model Predictive Control
by: Liang, Jacky, et al.
Published: (2024)
by: Liang, Jacky, et al.
Published: (2024)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
by: Ahn, Michael, et al.
Published: (2024)
by: Ahn, Michael, et al.
Published: (2024)
Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
by: Li, Chengshu, et al.
Published: (2023)
by: Li, Chengshu, et al.
Published: (2023)
PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement
by: Zhang, Tuo, et al.
Published: (2026)
by: Zhang, Tuo, et al.
Published: (2026)
Self-Improving Embodied Foundation Models
by: Ghasemipour, Seyed Kamyar Seyed, et al.
Published: (2025)
by: Ghasemipour, Seyed Kamyar Seyed, et al.
Published: (2025)
robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
by: Zhu, Yuke, et al.
Published: (2020)
by: Zhu, Yuke, et al.
Published: (2020)
PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation
by: Zhang, Kaidong, et al.
Published: (2024)
by: Zhang, Kaidong, et al.
Published: (2024)
SPARSE-PIVOT: Dynamic correlation clustering for node insertions
by: Dalirrooyfard, Mina, et al.
Published: (2025)
by: Dalirrooyfard, Mina, et al.
Published: (2025)
A Simple Evaluation Model for Feature Subset Selection Algorithms
by: Huei Diana Lee
Published: (2006)
by: Huei Diana Lee
Published: (2006)
Producción en la radio moderna / Carl Hausman, Philip Benoit, Lewis B. O’Donnell ; traducción de Eloy Pineda Rojas
by: Hausman, Carl, 1953-
Published: (1953)
by: Hausman, Carl, 1953-
Published: (1953)
Global logistics indicators, supply chain metrics, and bilateral trade patterns / Warren H. Hausman, Hau L. Lee, Uma Subramanian
by: Hausman, Warren H
Published: (2005)
by: Hausman, Warren H
Published: (2005)
CORRECCIÓN DE LA EVALUACIÓN ERRÓNEA DE LA OCDE ACERCA DE LA COMPETENCIA EN EL SECTOR DE LAS TELECOMUNICACIONES EN MÉXICO
by: Jerry A. Hausman
Published: (2013)
by: Jerry A. Hausman
Published: (2013)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
CHEMICAL CONTROL OVER THE SYNTHESIS OF DEFECT-FREE AND HETEROMETAL-DOPED METAL OXIDE NANOPARTICLES BY THE SINGLE-SOURCE PRECURSOR APPROACH
by: M. Driess
Published: (2005)
by: M. Driess
Published: (2005)
Neural network-based CUSUM for online change-point detection
by: Gong, Tingnan, et al.
Published: (2022)
by: Gong, Tingnan, et al.
Published: (2022)
Fostering riparian cooperation in international river basins : the World Bank at its best in development diplomacy / Syed Kirmani, Guy Le Moigne
by: Kirmani, Syed
by: Kirmani, Syed
What Matters in Learning from Large-Scale Datasets for Robot Manipulation
by: Saxena, Vaibhav, et al.
Published: (2025)
by: Saxena, Vaibhav, et al.
Published: (2025)
Actionable Warning Is Not Enough: Recommending Valid Actionable Warnings with Weak Supervision
by: Xue, Zhipeng, et al.
Published: (2025)
by: Xue, Zhipeng, et al.
Published: (2025)
Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
by: Driess, Danny, et al.
Published: (2025)
by: Driess, Danny, et al.
Published: (2025)
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
$π_0$: A Vision-Language-Action Flow Model for General Robot Control
by: Black, Kevin, et al.
Published: (2024)
by: Black, Kevin, et al.
Published: (2024)
Interpretability Can Be Actionable
by: Orgad, Hadas, et al.
Published: (2026)
by: Orgad, Hadas, et al.
Published: (2026)
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
by: Torne, Marcel, et al.
Published: (2026)
by: Torne, Marcel, et al.
Published: (2026)
Do We Need Large VLMs for Spotting Soccer Actions?
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
PIVOT-Net: Heterogeneous Point-Voxel-Tree-based Framework for Point Cloud Compression
by: Pang, Jiahao, et al.
Published: (2024)
by: Pang, Jiahao, et al.
Published: (2024)
KTCF: Actionable Recourse in Knowledge Tracing via Counterfactual Explanations for Education
by: Kim, Woojin, et al.
Published: (2026)
by: Kim, Woojin, et al.
Published: (2026)
Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation
by: Maddukuri, Abhiram, et al.
Published: (2025)
by: Maddukuri, Abhiram, et al.
Published: (2025)
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
Similar Items
-
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024) -
RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
by: Nasiriany, Soroush, et al.
Published: (2026) -
Vision Language Models are In-Context Value Learners
by: Ma, Yecheng Jason, et al.
Published: (2024) -
ALOHA Unleashed: A Simple Recipe for Robot Dexterity
by: Zhao, Tony Z., et al.
Published: (2024) -
PRIME: Scaffolding Manipulation Tasks with Behavior Primitives for Data-Efficient Imitation Learning
by: Gao, Tian, et al.
Published: (2024)