Goal-Based Vision-Language Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Patapati, Santosh, Srinivasan, Trisanth |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-Language Cross-Attention for Real-Time Autonomous Driving
by: Patapati, Santosh, et al.
Published: (2025)
by: Patapati, Santosh, et al.
Published: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
by: Romero, Angel, et al.
Published: (2025)
by: Romero, Angel, et al.
Published: (2025)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
by: Gasparino, Mateus Valverde, et al.
Published: (2024)
by: Gasparino, Mateus Valverde, et al.
Published: (2024)
Pointing-Guided Target Estimation via Transformer-Based Attention
by: Müller, Luca, et al.
Published: (2025)
by: Müller, Luca, et al.
Published: (2025)
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
by: Römer, Ralf, et al.
Published: (2026)
by: Römer, Ralf, et al.
Published: (2026)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
by: Syed, Shahram Najam, et al.
Published: (2025)
by: Syed, Shahram Najam, et al.
Published: (2025)
Structured Interfaces for Automated Reasoning with 3D Scene Graphs
by: Ray, Aaron, et al.
Published: (2025)
by: Ray, Aaron, et al.
Published: (2025)
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
by: Ziakas, Christos, et al.
Published: (2025)
by: Ziakas, Christos, et al.
Published: (2025)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
by: Liu, Zhi
Published: (2026)
by: Liu, Zhi
Published: (2026)
Optimizing Neurorobot Policy under Limited Demonstration Data through Preference Regret
by: Nguyen, Viet Dung, et al.
Published: (2026)
by: Nguyen, Viet Dung, et al.
Published: (2026)
Bayesian Data Augmentation and Training for Perception DNN in Autonomous Aerial Vehicles
by: Rasul, Ashik E, et al.
Published: (2024)
by: Rasul, Ashik E, et al.
Published: (2024)
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
by: Arnaud, Sergio, et al.
Published: (2025)
by: Arnaud, Sergio, et al.
Published: (2025)
Curb Your Attention: Causal Attention Gating for Robust Trajectory Prediction in Autonomous Driving
by: Ahmadi, Ehsan, et al.
Published: (2024)
by: Ahmadi, Ehsan, et al.
Published: (2024)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
by: Koh, Hyunseo, et al.
Published: (2026)
by: Koh, Hyunseo, et al.
Published: (2026)
Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos
by: Hu, X., et al.
Published: (2025)
by: Hu, X., et al.
Published: (2025)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
by: Mansour, Jad, et al.
Published: (2025)
by: Mansour, Jad, et al.
Published: (2025)
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction
by: Mansour, Jad, et al.
Published: (2024)
by: Mansour, Jad, et al.
Published: (2024)
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
by: Aida, Adriana, et al.
Published: (2026)
by: Aida, Adriana, et al.
Published: (2026)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
by: Babu, Abhijith, et al.
Published: (2026)
by: Babu, Abhijith, et al.
Published: (2026)
Zero-Shot Object Goal Visual Navigation With Class-Independent Relationship Network
by: Li, Xinting, et al.
Published: (2023)
by: Li, Xinting, et al.
Published: (2023)
HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
by: Wang, Zhi, et al.
Published: (2026)
by: Wang, Zhi, et al.
Published: (2026)
M3CAD: Towards Generic Cooperative Autonomous Driving Benchmark
by: Zhu, Morui, et al.
Published: (2025)
by: Zhu, Morui, et al.
Published: (2025)
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
by: Kang, Li, et al.
Published: (2026)
by: Kang, Li, et al.
Published: (2026)
The Impact of 2D Segmentation Backbones on Point Cloud Predictions Using 4D Radar
by: Muckelroy III, William, et al.
Published: (2025)
by: Muckelroy III, William, et al.
Published: (2025)
From Photons to Physics: Autonomous Indoor Drones and the Future of Objective Property Assessment
by: Teikari, Petteri, et al.
Published: (2025)
by: Teikari, Petteri, et al.
Published: (2025)
CC-SGG: Corner Case Scenario Generation using Learned Scene Graphs
by: Drayson, George, et al.
Published: (2023)
by: Drayson, George, et al.
Published: (2023)
VLA Foundry: A Unified Framework for Training Vision-Language-Action Models
by: Mercat, Jean, et al.
Published: (2026)
by: Mercat, Jean, et al.
Published: (2026)
Tulip Agent -- Enabling LLM-Based Agents to Solve Tasks Using Large Tool Libraries
by: Ocker, Felix, et al.
Published: (2024)
by: Ocker, Felix, et al.
Published: (2024)
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
by: Römer, Ralf, et al.
Published: (2025)
by: Römer, Ralf, et al.
Published: (2025)
Contextual Graph Representations for Task-Driven 3D Perception and Planning
by: Agia, Christopher
Published: (2026)
by: Agia, Christopher
Published: (2026)
RoboPack: Learning Tactile-Informed Dynamics Models for Dense Packing
by: Ai, Bo, et al.
Published: (2024)
by: Ai, Bo, et al.
Published: (2024)
Can Robots "Taste" Grapes? Estimating SSC with Simple RGB Sensors
by: Ciarfuglia, Thomas Alessandro, et al.
Published: (2024)
by: Ciarfuglia, Thomas Alessandro, et al.
Published: (2024)
Baby Sophia: A Developmental Approach to Self-Exploration through Self-Touch and Hand Regard
by: Zarifis, Stelios, et al.
Published: (2025)
by: Zarifis, Stelios, et al.
Published: (2025)
Transformers for Image-Goal Navigation
by: Pelluri, Nikhilanj
Published: (2024)
by: Pelluri, Nikhilanj
Published: (2024)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
by: Mehta, Vinit, et al.
Published: (2025)
by: Mehta, Vinit, et al.
Published: (2025)
Spot-Compose: A Framework for Open-Vocabulary Object Retrieval and Drawer Manipulation in Point Clouds
by: Lemke, Oliver, et al.
Published: (2024)
by: Lemke, Oliver, et al.
Published: (2024)
VIN-NBV: A View Introspection Network for Next-Best-View Selection
by: Frahm, Noah, et al.
Published: (2025)
by: Frahm, Noah, et al.
Published: (2025)
A Survey of Spatial Memory Representations for Efficient Robot Navigation
by: Pangaliman, Ma. Madecheen S., et al.
Published: (2026)
by: Pangaliman, Ma. Madecheen S., et al.
Published: (2026)
RAE-NWM: Navigation World Model in Dense Visual Representation Space
by: Zhang, Mingkun, et al.
Published: (2026)
by: Zhang, Mingkun, et al.
Published: (2026)
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
by: Deichler, Anna
Published: (2026)
by: Deichler, Anna
Published: (2026)
Similar Items
-
Vision-Language Cross-Attention for Real-Time Autonomous Driving
by: Patapati, Santosh, et al.
Published: (2025) -
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
by: Romero, Angel, et al.
Published: (2025) -
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
by: Gasparino, Mateus Valverde, et al.
Published: (2024) -
Pointing-Guided Target Estimation via Transformer-Based Attention
by: Müller, Luca, et al.
Published: (2025) -
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
by: Römer, Ralf, et al.
Published: (2026)