Semantic-Drive: Democratizing Long-Tail Data Curation via Open-Vocabulary Grounding and Neuro-Symbolic VLM Consensus
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Guillen-Perez, Antonio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
Neuro-Symbolic Concepts
von: Mao, Jiayuan, et al.
Veröffentlicht: (2025)
von: Mao, Jiayuan, et al.
Veröffentlicht: (2025)
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
von: Werby, Abdelrhman, et al.
Veröffentlicht: (2024)
von: Werby, Abdelrhman, et al.
Veröffentlicht: (2024)
Context-based Motion Retrieval using Open Vocabulary Methods for Autonomous Driving
von: Englmeier, Stefan, et al.
Veröffentlicht: (2025)
von: Englmeier, Stefan, et al.
Veröffentlicht: (2025)
A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
von: Jahangard, Simindokht, et al.
Veröffentlicht: (2025)
von: Jahangard, Simindokht, et al.
Veröffentlicht: (2025)
OSMa-Bench: Evaluating Open Semantic Mapping Under Varying Lighting Conditions
von: Popov, Maxim, et al.
Veröffentlicht: (2025)
von: Popov, Maxim, et al.
Veröffentlicht: (2025)
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
von: Atuhurra, Jesse, et al.
Veröffentlicht: (2025)
von: Atuhurra, Jesse, et al.
Veröffentlicht: (2025)
LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset
von: Wagner, Royden, et al.
Veröffentlicht: (2026)
von: Wagner, Royden, et al.
Veröffentlicht: (2026)
Neuro-Symbolic Manipulation Understanding with Enriched Semantic Event Chains
von: Ziaeetabar, Fatemeh
Veröffentlicht: (2026)
von: Ziaeetabar, Fatemeh
Veröffentlicht: (2026)
Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning with Dense Labeling
von: Yashima, Daichi, et al.
Veröffentlicht: (2024)
von: Yashima, Daichi, et al.
Veröffentlicht: (2024)
VERDI: VLM-Embedded Reasoning for Autonomous Driving
von: Feng, Bowen, et al.
Veröffentlicht: (2025)
von: Feng, Bowen, et al.
Veröffentlicht: (2025)
InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects
von: Mohamed, Ibrahim Sheikh, et al.
Veröffentlicht: (2025)
von: Mohamed, Ibrahim Sheikh, et al.
Veröffentlicht: (2025)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
Open-Vocabulary Online Semantic Mapping for SLAM
von: Martins, Tomas Berriel, et al.
Veröffentlicht: (2024)
von: Martins, Tomas Berriel, et al.
Veröffentlicht: (2024)
DM2RM: Dual-Mode Multimodal Ranking for Target Objects and Receptacles Based on Open-Vocabulary Instructions
von: Korekata, Ryosuke, et al.
Veröffentlicht: (2024)
von: Korekata, Ryosuke, et al.
Veröffentlicht: (2024)
Grounding Driving VLA via Inverse Kinematics
von: Park, Junsung, et al.
Veröffentlicht: (2026)
von: Park, Junsung, et al.
Veröffentlicht: (2026)
OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
von: Liu, Yuecheng, et al.
Veröffentlicht: (2025)
von: Liu, Yuecheng, et al.
Veröffentlicht: (2025)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
von: Li, Rong, et al.
Veröffentlicht: (2024)
von: Li, Rong, et al.
Veröffentlicht: (2024)
LangCoop: Collaborative Driving with Language
von: Gao, Xiangbo, et al.
Veröffentlicht: (2025)
von: Gao, Xiangbo, et al.
Veröffentlicht: (2025)
A Language Agent for Autonomous Driving
von: Mao, Jiageng, et al.
Veröffentlicht: (2023)
von: Mao, Jiageng, et al.
Veröffentlicht: (2023)
Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
OpenESS: Event-based Semantic Scene Understanding with Open Vocabularies
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
HomeRobot: Open-Vocabulary Mobile Manipulation
von: Yenamandra, Sriram, et al.
Veröffentlicht: (2023)
von: Yenamandra, Sriram, et al.
Veröffentlicht: (2023)
LINGO-Space: Language-Conditioned Incremental Grounding for Space
von: Kim, Dohyun, et al.
Veröffentlicht: (2024)
von: Kim, Dohyun, et al.
Veröffentlicht: (2024)
SEAL: Vision-Language Model-Based Safe End-to-End Cooperative Autonomous Driving with Adaptive Long-Tail Modeling
von: You, Junwei, et al.
Veröffentlicht: (2025)
von: You, Junwei, et al.
Veröffentlicht: (2025)
Enhancing End-to-End Autonomous Driving with Risk Semantic Distillaion from VLM
von: Qin, Jack, et al.
Veröffentlicht: (2025)
von: Qin, Jack, et al.
Veröffentlicht: (2025)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
von: Li, Zerui, et al.
Veröffentlicht: (2025)
von: Li, Zerui, et al.
Veröffentlicht: (2025)
SD-OVON: A Semantics-aware Dataset and Benchmark Generation Pipeline for Open-Vocabulary Object Navigation in Dynamic Scenes
von: Qiu, Dicong, et al.
Veröffentlicht: (2025)
von: Qiu, Dicong, et al.
Veröffentlicht: (2025)
DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving
von: Huang, Zilin, et al.
Veröffentlicht: (2026)
von: Huang, Zilin, et al.
Veröffentlicht: (2026)
SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving
von: Gao, Xiangbo, et al.
Veröffentlicht: (2025)
von: Gao, Xiangbo, et al.
Veröffentlicht: (2025)
Open-Vocabulary Action Localization with Iterative Visual Prompting
von: Wake, Naoki, et al.
Veröffentlicht: (2024)
von: Wake, Naoki, et al.
Veröffentlicht: (2024)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
CurricuVLM: Towards Safe Autonomous Driving via Personalized Safety-Critical Curriculum Learning with Vision-Language Models
von: Sheng, Zihao, et al.
Veröffentlicht: (2025)
von: Sheng, Zihao, et al.
Veröffentlicht: (2025)
SIMSplat: Predictive Driving Scene Editing with Language-aligned 4D Gaussian Splatting
von: Park, Sung-Yeon, et al.
Veröffentlicht: (2025)
von: Park, Sung-Yeon, et al.
Veröffentlicht: (2025)
From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning
von: Gado, Ahmed Y., et al.
Veröffentlicht: (2026)
von: Gado, Ahmed Y., et al.
Veröffentlicht: (2026)
Work Zones challenge VLM Trajectory Planning: Toward Mitigation and Robust Autonomous Driving
von: Liao, Yifan, et al.
Veröffentlicht: (2025)
von: Liao, Yifan, et al.
Veröffentlicht: (2025)
Break Out the Silverware -- Semantic Understanding of Stored Household Items
von: Levi-Richter, Michaela, et al.
Veröffentlicht: (2025)
von: Levi-Richter, Michaela, et al.
Veröffentlicht: (2025)
Scene-Graph ViT: End-to-End Open-Vocabulary Visual Relationship Detection
von: Salzmann, Tim, et al.
Veröffentlicht: (2024)
von: Salzmann, Tim, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024) -
Neuro-Symbolic Concepts
von: Mao, Jiayuan, et al.
Veröffentlicht: (2025) -
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022) -
Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
von: Werby, Abdelrhman, et al.
Veröffentlicht: (2024) -
Context-based Motion Retrieval using Open Vocabulary Methods for Autonomous Driving
von: Englmeier, Stefan, et al.
Veröffentlicht: (2025)