Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
Fuente:
arXiv
Guardado en:
| Autores principales: | Mitra, Chancharik, Anwar, Abrar, Corona, Rodolfo, Klein, Dan, Darrell, Trevor, Thomason, Jesse |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
por: Mitra, Chancharik, et al.
Publicado: (2025)
por: Mitra, Chancharik, et al.
Publicado: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
por: Mitra, Chancharik, et al.
Publicado: (2023)
por: Mitra, Chancharik, et al.
Publicado: (2023)
In-Context Learning Enables Robot Action Prediction in LLMs
por: Yin, Yida, et al.
Publicado: (2024)
por: Yin, Yida, et al.
Publicado: (2024)
PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language Models
por: Zhu, Wang Bill, et al.
Publicado: (2025)
por: Zhu, Wang Bill, et al.
Publicado: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
por: Huang, Brandon, et al.
Publicado: (2024)
por: Huang, Brandon, et al.
Publicado: (2024)
Contrast Sets for Evaluating Language-Guided Robot Policies
por: Anwar, Abrar, et al.
Publicado: (2024)
por: Anwar, Abrar, et al.
Publicado: (2024)
Language Models can Infer Action Semantics for Symbolic Planners from Environment Feedback
por: Zhu, Wang, et al.
Publicado: (2024)
por: Zhu, Wang, et al.
Publicado: (2024)
Leveraging Large Language Models in Human-Robot Interaction: A Critical Analysis of Potential and Pitfalls
por: Atuhurra, Jesse
Publicado: (2024)
por: Atuhurra, Jesse
Publicado: (2024)
Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People
por: Merchant, Zain, et al.
Publicado: (2024)
por: Merchant, Zain, et al.
Publicado: (2024)
TwoStep: Multi-agent Task Planning using Classical Planners and Large Language Models
por: Bai, David, et al.
Publicado: (2024)
por: Bai, David, et al.
Publicado: (2024)
M3PT: A Transformer for Multimodal, Multi-Party Social Signal Prediction with Person-aware Blockwise Attention
por: Tang, Yiming, et al.
Publicado: (2025)
por: Tang, Yiming, et al.
Publicado: (2025)
ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation
por: Anwar, Abrar, et al.
Publicado: (2024)
por: Anwar, Abrar, et al.
Publicado: (2024)
RobotFleet: An Open-Source Framework for Centralized Multi-Robot Task Planning
por: Gupta, Rohan, et al.
Publicado: (2025)
por: Gupta, Rohan, et al.
Publicado: (2025)
Enough Coin Flips Can Make LLMs Act Bayesian
por: Gupta, Ritwik, et al.
Publicado: (2025)
por: Gupta, Ritwik, et al.
Publicado: (2025)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
por: Mitra, Chancharik, et al.
Publicado: (2024)
por: Mitra, Chancharik, et al.
Publicado: (2024)
Pose Priors from Language Models
por: Subramanian, Sanjay, et al.
Publicado: (2024)
por: Subramanian, Sanjay, et al.
Publicado: (2024)
SmallPlan: Leverage Small Language Models for Sequential Path Planning with Simulation-Powered, LLM-Guided Distillation
por: Pham, Quang P. M., et al.
Publicado: (2025)
por: Pham, Quang P. M., et al.
Publicado: (2025)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
por: Liu, Fanfan, et al.
Publicado: (2024)
por: Liu, Fanfan, et al.
Publicado: (2024)
WinoViz: Probing Visual Properties of Objects Under Different States
por: Jin, Woojeong, et al.
Publicado: (2024)
por: Jin, Woojeong, et al.
Publicado: (2024)
OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision-Language Foundation Models
por: Kuang, Yuxuan, et al.
Publicado: (2024)
por: Kuang, Yuxuan, et al.
Publicado: (2024)
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
por: Lu, Jinghui, et al.
Publicado: (2026)
por: Lu, Jinghui, et al.
Publicado: (2026)
Empathic Grounding: Explorations using Multimodal Interaction and Large Language Models with Conversational Agents
por: Arjmand, Mehdi, et al.
Publicado: (2024)
por: Arjmand, Mehdi, et al.
Publicado: (2024)
Analyzing The Language of Visual Tokens
por: Chan, David M., et al.
Publicado: (2024)
por: Chan, David M., et al.
Publicado: (2024)
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
por: Srinivasan, Tejas, et al.
Publicado: (2025)
por: Srinivasan, Tejas, et al.
Publicado: (2025)
Can LLMs plan paths with extra hints from solvers?
por: Wu, Erik, et al.
Publicado: (2024)
por: Wu, Erik, et al.
Publicado: (2024)
NERsocial: Efficient Named Entity Recognition Dataset Construction for Human-Robot Interaction Utilizing RapidNER
por: Atuhurra, Jesse, et al.
Publicado: (2024)
por: Atuhurra, Jesse, et al.
Publicado: (2024)
A Survey of Robotic Language Grounding: Tradeoffs between Symbols and Embeddings
por: Cohen, Vanya, et al.
Publicado: (2024)
por: Cohen, Vanya, et al.
Publicado: (2024)
Grounding Large Language Models In Embodied Environment With Imperfect World Models
por: Liu, Haolan, et al.
Publicado: (2024)
por: Liu, Haolan, et al.
Publicado: (2024)
Unsupervised, Bottom-up Category Discovery for Symbol Grounding with a Curious Robot
por: Henry, Catherine, et al.
Publicado: (2024)
por: Henry, Catherine, et al.
Publicado: (2024)
PROGrasp: Pragmatic Human-Robot Communication for Object Grasping
por: Kang, Gi-Cheon, et al.
Publicado: (2023)
por: Kang, Gi-Cheon, et al.
Publicado: (2023)
Leveraging Adaptive Group Negotiation for Heterogeneous Multi-Robot Collaboration with Large Language Models
por: Song, Siqi, et al.
Publicado: (2025)
por: Song, Siqi, et al.
Publicado: (2025)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
por: Darabi, Nastaran, et al.
Publicado: (2026)
por: Darabi, Nastaran, et al.
Publicado: (2026)
ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
por: Zhang, Jiahui, et al.
Publicado: (2025)
por: Zhang, Jiahui, et al.
Publicado: (2025)
LINGO-Space: Language-Conditioned Incremental Grounding for Space
por: Kim, Dohyun, et al.
Publicado: (2024)
por: Kim, Dohyun, et al.
Publicado: (2024)
Retrieval-Augmented Hierarchical in-Context Reinforcement Learning and Hindsight Modular Reflections for Task Planning with LLMs
por: Sun, Chuanneng, et al.
Publicado: (2024)
por: Sun, Chuanneng, et al.
Publicado: (2024)
Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models
por: Wang, Tianyu, et al.
Publicado: (2024)
por: Wang, Tianyu, et al.
Publicado: (2024)
Grasp Multiple Objects with One Hand
por: Li, Yuyang, et al.
Publicado: (2023)
por: Li, Yuyang, et al.
Publicado: (2023)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
por: Li, Zerui, et al.
Publicado: (2025)
por: Li, Zerui, et al.
Publicado: (2025)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
por: Wang, Zhaowei, et al.
Publicado: (2024)
por: Wang, Zhaowei, et al.
Publicado: (2024)
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
por: Waheed, Abdul, et al.
Publicado: (2025)
por: Waheed, Abdul, et al.
Publicado: (2025)
Ejemplares similares
-
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
por: Mitra, Chancharik, et al.
Publicado: (2025) -
Compositional Chain-of-Thought Prompting for Large Multimodal Models
por: Mitra, Chancharik, et al.
Publicado: (2023) -
In-Context Learning Enables Robot Action Prediction in LLMs
por: Yin, Yida, et al.
Publicado: (2024) -
PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language Models
por: Zhu, Wang Bill, et al.
Publicado: (2025) -
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
por: Huang, Brandon, et al.
Publicado: (2024)