In-Context Learning Enables Robot Action Prediction in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Yin, Yida, Wang, Zekai, Sharma, Yuvan, Niu, Dantong, Darrell, Trevor, Herzig, Roei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
por: Niu, Dantong, et al.
Publicado: (2024)
por: Niu, Dantong, et al.
Publicado: (2024)
Pre-training Auto-regressive Robotic Models with 4D Representations
por: Niu, Dantong, et al.
Publicado: (2025)
por: Niu, Dantong, et al.
Publicado: (2025)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
por: Mitra, Chancharik, et al.
Publicado: (2025)
por: Mitra, Chancharik, et al.
Publicado: (2025)
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
por: Hsieh, Wen-Han, et al.
Publicado: (2025)
por: Hsieh, Wen-Han, et al.
Publicado: (2025)
Learning to Grasp Anything by Playing with Random Toys
por: Niu, Dantong, et al.
Publicado: (2025)
por: Niu, Dantong, et al.
Publicado: (2025)
From Generated Human Videos to Physically Plausible Robot Trajectories
por: Ni, James, et al.
Publicado: (2025)
por: Ni, James, et al.
Publicado: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
por: Huang, Brandon, et al.
Publicado: (2024)
por: Huang, Brandon, et al.
Publicado: (2024)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
por: Mitra, Chancharik, et al.
Publicado: (2023)
por: Mitra, Chancharik, et al.
Publicado: (2023)
Recursive Visual Programming
por: Ge, Jiaxin, et al.
Publicado: (2023)
por: Ge, Jiaxin, et al.
Publicado: (2023)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
por: Shang, Chuyi, et al.
Publicado: (2024)
por: Shang, Chuyi, et al.
Publicado: (2024)
AlphaSpace: Enabling Robotic Actions through Semantic Tokenization and Symbolic Reasoning
por: Dao, Alan, et al.
Publicado: (2025)
por: Dao, Alan, et al.
Publicado: (2025)
Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
por: Mitra, Chancharik, et al.
Publicado: (2023)
por: Mitra, Chancharik, et al.
Publicado: (2023)
Retrieval-Augmented Hierarchical in-Context Reinforcement Learning and Hindsight Modular Reflections for Task Planning with LLMs
por: Sun, Chuanneng, et al.
Publicado: (2024)
por: Sun, Chuanneng, et al.
Publicado: (2024)
InCoRo: In-Context Learning for Robotics Control with Feedback Loops
por: Zhu, Jiaqiang Ye, et al.
Publicado: (2024)
por: Zhu, Jiaqiang Ye, et al.
Publicado: (2024)
LLMs for Robotic Object Disambiguation
por: Jiang, Connie, et al.
Publicado: (2024)
por: Jiang, Connie, et al.
Publicado: (2024)
Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics
por: Di Palo, Norman, et al.
Publicado: (2024)
por: Di Palo, Norman, et al.
Publicado: (2024)
ViPRA: Video Prediction for Robot Actions
por: Routray, Sandeep, et al.
Publicado: (2025)
por: Routray, Sandeep, et al.
Publicado: (2025)
Ain't Misbehavin' -- Using LLMs to Generate Expressive Robot Behavior in Conversations with the Tabletop Robot Haru
por: Wang, Zining, et al.
Publicado: (2024)
por: Wang, Zining, et al.
Publicado: (2024)
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
por: Yang, Yandan, et al.
Publicado: (2026)
por: Yang, Yandan, et al.
Publicado: (2026)
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
por: Wang, Qiuyue, et al.
Publicado: (2026)
por: Wang, Qiuyue, et al.
Publicado: (2026)
Stable Language Guidance for Vision-Language-Action Models
por: Zhan, Zhihao, et al.
Publicado: (2026)
por: Zhan, Zhihao, et al.
Publicado: (2026)
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
por: Wang, Siyin, et al.
Publicado: (2025)
por: Wang, Siyin, et al.
Publicado: (2025)
VoicePilot: Harnessing LLMs as Speech Interfaces for Physically Assistive Robots
por: Padmanabha, Akhil, et al.
Publicado: (2024)
por: Padmanabha, Akhil, et al.
Publicado: (2024)
LLM-Driven Robots Risk Enacting Discrimination, Violence, and Unlawful Actions
por: Hundt, Andrew, et al.
Publicado: (2024)
por: Hundt, Andrew, et al.
Publicado: (2024)
ASMR: Augmenting Life Scenario using Large Generative Models for Robotic Action Reflection
por: Tsai, Shang-Chi, et al.
Publicado: (2025)
por: Tsai, Shang-Chi, et al.
Publicado: (2025)
Hazards in Daily Life? Enabling Robots to Proactively Detect and Resolve Anomalies
por: Song, Zirui, et al.
Publicado: (2024)
por: Song, Zirui, et al.
Publicado: (2024)
Scaling Laws in Scientific Discovery with AI and Robot Scientists
por: Zhang, Pengsong, et al.
Publicado: (2025)
por: Zhang, Pengsong, et al.
Publicado: (2025)
Joint Action Language Modelling for Transparent Policy Execution
por: Wulff, Theodor, et al.
Publicado: (2025)
por: Wulff, Theodor, et al.
Publicado: (2025)
EdgeVLA: Efficient Vision-Language-Action Models
por: Budzianowski, Paweł, et al.
Publicado: (2025)
por: Budzianowski, Paweł, et al.
Publicado: (2025)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
por: Mitra, Chancharik, et al.
Publicado: (2024)
por: Mitra, Chancharik, et al.
Publicado: (2024)
Can LLMs Translate Human Instructions into a Reinforcement Learning Agent's Internal Emergent Symbolic Representation?
por: Ma, Ziqi, et al.
Publicado: (2025)
por: Ma, Ziqi, et al.
Publicado: (2025)
Red-Teaming Vision-Language-Action Models via Quality Diversity Prompt Generation for Robust Robot Policies
por: Srikanth, Siddharth, et al.
Publicado: (2026)
por: Srikanth, Siddharth, et al.
Publicado: (2026)
TULIP: Towards Unified Language-Image Pretraining
por: Tang, Zineng, et al.
Publicado: (2025)
por: Tang, Zineng, et al.
Publicado: (2025)
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
por: Borazjanizadeh, Nasim, et al.
Publicado: (2024)
por: Borazjanizadeh, Nasim, et al.
Publicado: (2024)
ALRM: Agentic LLM for Robotic Manipulation
por: Santos, Vitor Gaboardi dos, et al.
Publicado: (2026)
por: Santos, Vitor Gaboardi dos, et al.
Publicado: (2026)
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
por: Zhao, Zhida, et al.
Publicado: (2025)
por: Zhao, Zhida, et al.
Publicado: (2025)
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
por: Zheng, Huan, et al.
Publicado: (2025)
por: Zheng, Huan, et al.
Publicado: (2025)
Robot Detection System 1: Front-Following
por: Lin, Jinwei
Publicado: (2024)
por: Lin, Jinwei
Publicado: (2024)
Foundation Models for Autonomous Robots in Unstructured Environments
por: Naderi, Hossein, et al.
Publicado: (2024)
por: Naderi, Hossein, et al.
Publicado: (2024)
PROGrasp: Pragmatic Human-Robot Communication for Object Grasping
por: Kang, Gi-Cheon, et al.
Publicado: (2023)
por: Kang, Gi-Cheon, et al.
Publicado: (2023)
Ejemplares similares
-
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
por: Niu, Dantong, et al.
Publicado: (2024) -
Pre-training Auto-regressive Robotic Models with 4D Representations
por: Niu, Dantong, et al.
Publicado: (2025) -
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
por: Mitra, Chancharik, et al.
Publicado: (2025) -
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
por: Hsieh, Wen-Han, et al.
Publicado: (2025) -
Learning to Grasp Anything by Playing with Random Toys
por: Niu, Dantong, et al.
Publicado: (2025)