Goal-Based Vision-Language Driving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Patapati, Santosh, Srinivasan, Trisanth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Vision-Language Cross-Attention for Real-Time Autonomous Driving
von: Patapati, Santosh, et al.
Veröffentlicht: (2025)
von: Patapati, Santosh, et al.
Veröffentlicht: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
von: Romero, Angel, et al.
Veröffentlicht: (2025)
von: Romero, Angel, et al.
Veröffentlicht: (2025)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
von: Gasparino, Mateus Valverde, et al.
Veröffentlicht: (2024)
von: Gasparino, Mateus Valverde, et al.
Veröffentlicht: (2024)
Pointing-Guided Target Estimation via Transformer-Based Attention
von: Müller, Luca, et al.
Veröffentlicht: (2025)
von: Müller, Luca, et al.
Veröffentlicht: (2025)
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
von: Römer, Ralf, et al.
Veröffentlicht: (2026)
von: Römer, Ralf, et al.
Veröffentlicht: (2026)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
Structured Interfaces for Automated Reasoning with 3D Scene Graphs
von: Ray, Aaron, et al.
Veröffentlicht: (2025)
von: Ray, Aaron, et al.
Veröffentlicht: (2025)
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
von: Ziakas, Christos, et al.
Veröffentlicht: (2025)
von: Ziakas, Christos, et al.
Veröffentlicht: (2025)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
von: Liu, Zhi
Veröffentlicht: (2026)
von: Liu, Zhi
Veröffentlicht: (2026)
Optimizing Neurorobot Policy under Limited Demonstration Data through Preference Regret
von: Nguyen, Viet Dung, et al.
Veröffentlicht: (2026)
von: Nguyen, Viet Dung, et al.
Veröffentlicht: (2026)
Bayesian Data Augmentation and Training for Perception DNN in Autonomous Aerial Vehicles
von: Rasul, Ashik E, et al.
Veröffentlicht: (2024)
von: Rasul, Ashik E, et al.
Veröffentlicht: (2024)
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
von: Arnaud, Sergio, et al.
Veröffentlicht: (2025)
von: Arnaud, Sergio, et al.
Veröffentlicht: (2025)
Curb Your Attention: Causal Attention Gating for Robust Trajectory Prediction in Autonomous Driving
von: Ahmadi, Ehsan, et al.
Veröffentlicht: (2024)
von: Ahmadi, Ehsan, et al.
Veröffentlicht: (2024)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
von: Koh, Hyunseo, et al.
Veröffentlicht: (2026)
von: Koh, Hyunseo, et al.
Veröffentlicht: (2026)
Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos
von: Hu, X., et al.
Veröffentlicht: (2025)
von: Hu, X., et al.
Veröffentlicht: (2025)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
von: Mansour, Jad, et al.
Veröffentlicht: (2025)
von: Mansour, Jad, et al.
Veröffentlicht: (2025)
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction
von: Mansour, Jad, et al.
Veröffentlicht: (2024)
von: Mansour, Jad, et al.
Veröffentlicht: (2024)
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
von: Aida, Adriana, et al.
Veröffentlicht: (2026)
von: Aida, Adriana, et al.
Veröffentlicht: (2026)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
von: Babu, Abhijith, et al.
Veröffentlicht: (2026)
von: Babu, Abhijith, et al.
Veröffentlicht: (2026)
Zero-Shot Object Goal Visual Navigation With Class-Independent Relationship Network
von: Li, Xinting, et al.
Veröffentlicht: (2023)
von: Li, Xinting, et al.
Veröffentlicht: (2023)
HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
von: Wang, Zhi, et al.
Veröffentlicht: (2026)
von: Wang, Zhi, et al.
Veröffentlicht: (2026)
M3CAD: Towards Generic Cooperative Autonomous Driving Benchmark
von: Zhu, Morui, et al.
Veröffentlicht: (2025)
von: Zhu, Morui, et al.
Veröffentlicht: (2025)
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
von: Kang, Li, et al.
Veröffentlicht: (2026)
von: Kang, Li, et al.
Veröffentlicht: (2026)
The Impact of 2D Segmentation Backbones on Point Cloud Predictions Using 4D Radar
von: Muckelroy III, William, et al.
Veröffentlicht: (2025)
von: Muckelroy III, William, et al.
Veröffentlicht: (2025)
From Photons to Physics: Autonomous Indoor Drones and the Future of Objective Property Assessment
von: Teikari, Petteri, et al.
Veröffentlicht: (2025)
von: Teikari, Petteri, et al.
Veröffentlicht: (2025)
CC-SGG: Corner Case Scenario Generation using Learned Scene Graphs
von: Drayson, George, et al.
Veröffentlicht: (2023)
von: Drayson, George, et al.
Veröffentlicht: (2023)
VLA Foundry: A Unified Framework for Training Vision-Language-Action Models
von: Mercat, Jean, et al.
Veröffentlicht: (2026)
von: Mercat, Jean, et al.
Veröffentlicht: (2026)
Tulip Agent -- Enabling LLM-Based Agents to Solve Tasks Using Large Tool Libraries
von: Ocker, Felix, et al.
Veröffentlicht: (2024)
von: Ocker, Felix, et al.
Veröffentlicht: (2024)
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
von: Römer, Ralf, et al.
Veröffentlicht: (2025)
von: Römer, Ralf, et al.
Veröffentlicht: (2025)
Contextual Graph Representations for Task-Driven 3D Perception and Planning
von: Agia, Christopher
Veröffentlicht: (2026)
von: Agia, Christopher
Veröffentlicht: (2026)
RoboPack: Learning Tactile-Informed Dynamics Models for Dense Packing
von: Ai, Bo, et al.
Veröffentlicht: (2024)
von: Ai, Bo, et al.
Veröffentlicht: (2024)
Can Robots "Taste" Grapes? Estimating SSC with Simple RGB Sensors
von: Ciarfuglia, Thomas Alessandro, et al.
Veröffentlicht: (2024)
von: Ciarfuglia, Thomas Alessandro, et al.
Veröffentlicht: (2024)
Baby Sophia: A Developmental Approach to Self-Exploration through Self-Touch and Hand Regard
von: Zarifis, Stelios, et al.
Veröffentlicht: (2025)
von: Zarifis, Stelios, et al.
Veröffentlicht: (2025)
Transformers for Image-Goal Navigation
von: Pelluri, Nikhilanj
Veröffentlicht: (2024)
von: Pelluri, Nikhilanj
Veröffentlicht: (2024)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
von: Mehta, Vinit, et al.
Veröffentlicht: (2025)
von: Mehta, Vinit, et al.
Veröffentlicht: (2025)
Spot-Compose: A Framework for Open-Vocabulary Object Retrieval and Drawer Manipulation in Point Clouds
von: Lemke, Oliver, et al.
Veröffentlicht: (2024)
von: Lemke, Oliver, et al.
Veröffentlicht: (2024)
VIN-NBV: A View Introspection Network for Next-Best-View Selection
von: Frahm, Noah, et al.
Veröffentlicht: (2025)
von: Frahm, Noah, et al.
Veröffentlicht: (2025)
A Survey of Spatial Memory Representations for Efficient Robot Navigation
von: Pangaliman, Ma. Madecheen S., et al.
Veröffentlicht: (2026)
von: Pangaliman, Ma. Madecheen S., et al.
Veröffentlicht: (2026)
RAE-NWM: Navigation World Model in Dense Visual Representation Space
von: Zhang, Mingkun, et al.
Veröffentlicht: (2026)
von: Zhang, Mingkun, et al.
Veröffentlicht: (2026)
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
von: Deichler, Anna
Veröffentlicht: (2026)
von: Deichler, Anna
Veröffentlicht: (2026)
Ähnliche Einträge
-
Vision-Language Cross-Attention for Real-Time Autonomous Driving
von: Patapati, Santosh, et al.
Veröffentlicht: (2025) -
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
von: Romero, Angel, et al.
Veröffentlicht: (2025) -
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
von: Gasparino, Mateus Valverde, et al.
Veröffentlicht: (2024) -
Pointing-Guided Target Estimation via Transformer-Based Attention
von: Müller, Luca, et al.
Veröffentlicht: (2025) -
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
von: Römer, Ralf, et al.
Veröffentlicht: (2026)