Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Ünal, Berkehan, Dierend, Hauke, Fazlija, Dren, Plachetka, Christopher |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
por: Mehta, Vinit, et al.
Publicado: (2025)
por: Mehta, Vinit, et al.
Publicado: (2025)
FusionSense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction
por: Fang, Irving, et al.
Publicado: (2024)
por: Fang, Irving, et al.
Publicado: (2024)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
por: Hou, Zhiyi, et al.
Publicado: (2025)
por: Hou, Zhiyi, et al.
Publicado: (2025)
RCooper: A Real-world Large-scale Dataset for Roadside Cooperative Perception
por: Hao, Ruiyang, et al.
Publicado: (2024)
por: Hao, Ruiyang, et al.
Publicado: (2024)
A Computer Vision Pipeline for Iterative Bullet Hole Tracking in Rifle Zeroing
por: Belcher, Robert M., et al.
Publicado: (2026)
por: Belcher, Robert M., et al.
Publicado: (2026)
Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception
por: Meng, Siyuan, et al.
Publicado: (2026)
por: Meng, Siyuan, et al.
Publicado: (2026)
Model Agnostic Defense against Adversarial Patch Attacks on Object Detection in Unmanned Aerial Vehicles
por: Pathak, Saurabh, et al.
Publicado: (2024)
por: Pathak, Saurabh, et al.
Publicado: (2024)
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs
por: Cao, Yihan, et al.
Publicado: (2024)
por: Cao, Yihan, et al.
Publicado: (2024)
Learning the Pedestrian-Vehicle Interaction for Pedestrian Trajectory Prediction
por: Zhang, Chi, et al.
Publicado: (2022)
por: Zhang, Chi, et al.
Publicado: (2022)
Sparse 3D Perception for Rose Harvesting Robots: A Two-Stage Approach Bridging Simulation and Real-World Applications
por: Samavati, Taha, et al.
Publicado: (2025)
por: Samavati, Taha, et al.
Publicado: (2025)
Perception-to-Pursuit: Track-Centric Temporal Reasoning for Open-World Drone Detection and Autonomous Chasing
por: Oruganti, Venkatakrishna Reddy
Publicado: (2026)
por: Oruganti, Venkatakrishna Reddy
Publicado: (2026)
A Novel Camera-to-Robot Calibration Method for Vision-Based Floor Measurements
por: Rudolph, Jan Andre, et al.
Publicado: (2026)
por: Rudolph, Jan Andre, et al.
Publicado: (2026)
Grounding Synthetic Data Generation With Vision and Language Models
por: Çağlar, Ümit Mert, et al.
Publicado: (2026)
por: Çağlar, Ümit Mert, et al.
Publicado: (2026)
Vision-Language Cross-Attention for Real-Time Autonomous Driving
por: Patapati, Santosh, et al.
Publicado: (2025)
por: Patapati, Santosh, et al.
Publicado: (2025)
FALCON: Few-Shot Adversarial Learning for Cross-Domain Medical Image Segmentation
por: Fayjie, Abdur R., et al.
Publicado: (2026)
por: Fayjie, Abdur R., et al.
Publicado: (2026)
Can VLMs Unlock Semantic Anomaly Detection? A Framework for Structured Reasoning
por: Brusnicki, Roberto, et al.
Publicado: (2025)
por: Brusnicki, Roberto, et al.
Publicado: (2025)
Distant Object Localisation from Noisy Image Segmentation Sequences
por: Pesonen, Julius, et al.
Publicado: (2025)
por: Pesonen, Julius, et al.
Publicado: (2025)
Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors
por: Sun, Jiachen, et al.
Publicado: (2024)
por: Sun, Jiachen, et al.
Publicado: (2024)
On the Domain Robustness of Contrastive Vision-Language Models
por: Koddenbrock, Mario, et al.
Publicado: (2025)
por: Koddenbrock, Mario, et al.
Publicado: (2025)
GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
por: Hacheme, Gilles Quentin, et al.
Publicado: (2025)
por: Hacheme, Gilles Quentin, et al.
Publicado: (2025)
VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models
por: Bastien, JF, et al.
Publicado: (2026)
por: Bastien, JF, et al.
Publicado: (2026)
Efficient Multi-Band Temporal Video Filter for Reducing Human-Robot Interaction
por: O'Gorman, Lawrence
Publicado: (2024)
por: O'Gorman, Lawrence
Publicado: (2024)
Conceptual Evaluation of Deep Visual Stereo Odometry for the MARWIN Radiation Monitoring Robot in Accelerator Tunnels
por: Dehne, André, et al.
Publicado: (2025)
por: Dehne, André, et al.
Publicado: (2025)
Pathological Primitive Segmentation Based on Visual Foundation Model with Zero-Shot Mask Generation
por: Arnob, Abu Bakor Hayat, et al.
Publicado: (2024)
por: Arnob, Abu Bakor Hayat, et al.
Publicado: (2024)
Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation
por: Reddy, Ruturaj, et al.
Publicado: (2026)
por: Reddy, Ruturaj, et al.
Publicado: (2026)
Thermal and RGB Images Work Better Together in Wind Turbine Damage Detection
por: Svystun, Serhii, et al.
Publicado: (2024)
por: Svystun, Serhii, et al.
Publicado: (2024)
F3DGS: Federated 3D Gaussian Splatting for Decentralized Multi-Agent World Modeling
por: Zhu, Morui, et al.
Publicado: (2026)
por: Zhu, Morui, et al.
Publicado: (2026)
Gaze Estimation for Human-Robot Interaction: Analysis Using the NICO Platform
por: Palider, Matej, et al.
Publicado: (2025)
por: Palider, Matej, et al.
Publicado: (2025)
A Probabilistic Formulation of LiDAR Mapping with Neural Radiance Fields
por: McDermott, Matthew, et al.
Publicado: (2024)
por: McDermott, Matthew, et al.
Publicado: (2024)
CudaSIFT-SLAM: multiple-map visual SLAM for full procedure mapping in real human endoscopy
por: Elvira, Richard, et al.
Publicado: (2024)
por: Elvira, Richard, et al.
Publicado: (2024)
MARVO: Marine-Adaptive Radiance-aware Visual Odometry
por: Sundar, Sacchin, et al.
Publicado: (2025)
por: Sundar, Sacchin, et al.
Publicado: (2025)
Research Challenges and Progress in the End-to-End V2X Cooperative Autonomous Driving Competition
por: Hao, Ruiyang, et al.
Publicado: (2025)
por: Hao, Ruiyang, et al.
Publicado: (2025)
Bio-inspired visual relative localization for large swarms of UAVs
por: Křížek, Martin, et al.
Publicado: (2024)
por: Křížek, Martin, et al.
Publicado: (2024)
StyleDrive: Towards Driving-Style Aware Benchmarking of End-To-End Autonomous Driving
por: Hao, Ruiyang, et al.
Publicado: (2025)
por: Hao, Ruiyang, et al.
Publicado: (2025)
ICET Online Accuracy Characterization for Geometry-Based Laser Scan Matching
por: McDermott, Matthew, et al.
Publicado: (2023)
por: McDermott, Matthew, et al.
Publicado: (2023)
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
por: Xiao, Jiasong, et al.
Publicado: (2026)
por: Xiao, Jiasong, et al.
Publicado: (2026)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
por: Zhang, Jiwen, et al.
Publicado: (2026)
por: Zhang, Jiwen, et al.
Publicado: (2026)
The Impact of Synthetic Data on Object Detection Model Performance: A Comparative Analysis with Real-World Data
por: Bay, Muammer, et al.
Publicado: (2025)
por: Bay, Muammer, et al.
Publicado: (2025)
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
por: Yajima, Masaru, et al.
Publicado: (2025)
por: Yajima, Masaru, et al.
Publicado: (2025)
CODEI: Resource-Efficient Task-Driven Co-Design of Perception and Decision Making for Mobile Robots Applied to Autonomous Vehicles
por: Milojevic, Dejan, et al.
Publicado: (2025)
por: Milojevic, Dejan, et al.
Publicado: (2025)
Ejemplares similares
-
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
por: Mehta, Vinit, et al.
Publicado: (2025) -
FusionSense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction
por: Fang, Irving, et al.
Publicado: (2024) -
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
por: Hou, Zhiyi, et al.
Publicado: (2025) -
RCooper: A Real-world Large-scale Dataset for Roadside Cooperative Perception
por: Hao, Ruiyang, et al.
Publicado: (2024) -
A Computer Vision Pipeline for Iterative Bullet Hole Tracking in Rifle Zeroing
por: Belcher, Robert M., et al.
Publicado: (2026)