Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Gopinathan, Muraleekrishna, Masek, Martin, Abu-Khalaf, Jumana, Suter, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
di: Li, Danyang, et al.
Pubblicazione: (2025)
di: Li, Danyang, et al.
Pubblicazione: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
di: Lim, Shoon Kit, et al.
Pubblicazione: (2025)
di: Lim, Shoon Kit, et al.
Pubblicazione: (2025)
Autonomous Navigation and Collision Avoidance for Mobile Robots: Classification and Review
di: de Carvalho, Marcus Vinicius Leal, et al.
Pubblicazione: (2024)
di: de Carvalho, Marcus Vinicius Leal, et al.
Pubblicazione: (2024)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
di: Agia, Christopher, et al.
Pubblicazione: (2024)
di: Agia, Christopher, et al.
Pubblicazione: (2024)
EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
di: Chen, Junting, et al.
Pubblicazione: (2024)
di: Chen, Junting, et al.
Pubblicazione: (2024)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
di: Gasparino, Mateus Valverde, et al.
Pubblicazione: (2024)
di: Gasparino, Mateus Valverde, et al.
Pubblicazione: (2024)
Industrial Robot Motion Planning with GPUs: Integration of cuRobo for Extended DOF Systems
di: Abuelsamen, Luai, et al.
Pubblicazione: (2025)
di: Abuelsamen, Luai, et al.
Pubblicazione: (2025)
A Surveillance Based Interactive Robot
di: Kavimandan, Kshitij, et al.
Pubblicazione: (2025)
di: Kavimandan, Kshitij, et al.
Pubblicazione: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
di: Farzulla, Murad
Pubblicazione: (2026)
di: Farzulla, Murad
Pubblicazione: (2026)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
di: Tourani, Ali, et al.
Pubblicazione: (2023)
di: Tourani, Ali, et al.
Pubblicazione: (2023)
Deep Probabilistic Traversability with Test-time Adaptation for Uncertainty-aware Planetary Rover Navigation
di: Endo, Masafumi, et al.
Pubblicazione: (2024)
di: Endo, Masafumi, et al.
Pubblicazione: (2024)
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
di: Riva, Paolo, et al.
Pubblicazione: (2026)
di: Riva, Paolo, et al.
Pubblicazione: (2026)
Deployment-Time Reliability of Learned Robot Policies
di: Agia, Christopher
Pubblicazione: (2026)
di: Agia, Christopher
Pubblicazione: (2026)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
di: Romero, Angel, et al.
Pubblicazione: (2025)
di: Romero, Angel, et al.
Pubblicazione: (2025)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
di: Babu, Abhijith, et al.
Pubblicazione: (2026)
di: Babu, Abhijith, et al.
Pubblicazione: (2026)
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
di: Bonial, Claire, et al.
Pubblicazione: (2024)
di: Bonial, Claire, et al.
Pubblicazione: (2024)
SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
di: Lukin, Stephanie M., et al.
Pubblicazione: (2024)
di: Lukin, Stephanie M., et al.
Pubblicazione: (2024)
Motion Perceiver: Real-Time Occupancy Forecasting for Embedded Systems
di: Ferenczi, Bryce, et al.
Pubblicazione: (2023)
di: Ferenczi, Bryce, et al.
Pubblicazione: (2023)
PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision
di: Patania, Sabrina, et al.
Pubblicazione: (2025)
di: Patania, Sabrina, et al.
Pubblicazione: (2025)
RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation
di: Chen, Junting, et al.
Pubblicazione: (2024)
di: Chen, Junting, et al.
Pubblicazione: (2024)
Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos
di: Hu, X., et al.
Pubblicazione: (2025)
di: Hu, X., et al.
Pubblicazione: (2025)
VLA Foundry: A Unified Framework for Training Vision-Language-Action Models
di: Mercat, Jean, et al.
Pubblicazione: (2026)
di: Mercat, Jean, et al.
Pubblicazione: (2026)
Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs
di: Annese, Luca, et al.
Pubblicazione: (2025)
di: Annese, Luca, et al.
Pubblicazione: (2025)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
di: Zhang, Sinin, et al.
Pubblicazione: (2026)
di: Zhang, Sinin, et al.
Pubblicazione: (2026)
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
di: Deichler, Anna, et al.
Pubblicazione: (2025)
di: Deichler, Anna, et al.
Pubblicazione: (2025)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
di: Hu, Pan
Pubblicazione: (2025)
di: Hu, Pan
Pubblicazione: (2025)
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
di: Römer, Ralf, et al.
Pubblicazione: (2026)
di: Römer, Ralf, et al.
Pubblicazione: (2026)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
di: Syed, Shahram Najam, et al.
Pubblicazione: (2025)
di: Syed, Shahram Najam, et al.
Pubblicazione: (2025)
A Survey of Spatial Memory Representations for Efficient Robot Navigation
di: Pangaliman, Ma. Madecheen S., et al.
Pubblicazione: (2026)
di: Pangaliman, Ma. Madecheen S., et al.
Pubblicazione: (2026)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
A Segmented Robot Grasping Perception Neural Network for Edge AI
di: Bröcheler, Casper, et al.
Pubblicazione: (2025)
di: Bröcheler, Casper, et al.
Pubblicazione: (2025)
Ego-Motion Aware Target Prediction Module for Robust Multi-Object Tracking
di: Mahdian, Navid, et al.
Pubblicazione: (2024)
di: Mahdian, Navid, et al.
Pubblicazione: (2024)
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
di: Chen, Yiteng, et al.
Pubblicazione: (2025)
di: Chen, Yiteng, et al.
Pubblicazione: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025)
di: Dai, Song, et al.
Pubblicazione: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
di: Rahmatullaev, Temurbek, et al.
Pubblicazione: (2025)
di: Rahmatullaev, Temurbek, et al.
Pubblicazione: (2025)
Memory-Efficient Differentially Private Training with Gradient Random Projection
di: Mulrooney, Alex, et al.
Pubblicazione: (2025)
di: Mulrooney, Alex, et al.
Pubblicazione: (2025)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
di: Menon, Anjali R., et al.
Pubblicazione: (2025)
di: Menon, Anjali R., et al.
Pubblicazione: (2025)
Documenti analoghi
-
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024) -
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
di: Li, Danyang, et al.
Pubblicazione: (2025) -
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
di: Lim, Shoon Kit, et al.
Pubblicazione: (2025) -
Autonomous Navigation and Collision Avoidance for Mobile Robots: Classification and Review
di: de Carvalho, Marcus Vinicius Leal, et al.
Pubblicazione: (2024) -
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
di: Agia, Christopher, et al.
Pubblicazione: (2024)