Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Lohner, Aaron, Compagno, Francesco, Francis, Jonathan, Oltramari, Alessandro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video Synthesis
por: Fang, Jianwu, et al.
Publicado: (2025)
por: Fang, Jianwu, et al.
Publicado: (2025)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
por: Zhang, Jiwen, et al.
Publicado: (2026)
por: Zhang, Jiwen, et al.
Publicado: (2026)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
por: Li, Haoyuan, et al.
Publicado: (2025)
por: Li, Haoyuan, et al.
Publicado: (2025)
Explainable Scene Understanding with Qualitative Representations and Graph Neural Networks
por: Belmecheri, Nassim, et al.
Publicado: (2025)
por: Belmecheri, Nassim, et al.
Publicado: (2025)
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
por: Gao, Haoxiang, et al.
Publicado: (2025)
por: Gao, Haoxiang, et al.
Publicado: (2025)
FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction
por: Rotondi, Dennis, et al.
Publicado: (2025)
por: Rotondi, Dennis, et al.
Publicado: (2025)
CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning
por: Gan, Rui, et al.
Publicado: (2026)
por: Gan, Rui, et al.
Publicado: (2026)
Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding
por: Longo, Antonello, et al.
Publicado: (2025)
por: Longo, Antonello, et al.
Publicado: (2025)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
por: Jia, Baoxiong, et al.
Publicado: (2024)
por: Jia, Baoxiong, et al.
Publicado: (2024)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
por: Song, Chan Hee, et al.
Publicado: (2024)
por: Song, Chan Hee, et al.
Publicado: (2024)
Unifying 2D and 3D Vision-Language Understanding
por: Jain, Ayush, et al.
Publicado: (2025)
por: Jain, Ayush, et al.
Publicado: (2025)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
por: Hao, Haihong, et al.
Publicado: (2026)
por: Hao, Haihong, et al.
Publicado: (2026)
NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models
por: Park, Sung-Yeon, et al.
Publicado: (2025)
por: Park, Sung-Yeon, et al.
Publicado: (2025)
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
por: Chang, Yue, et al.
Publicado: (2026)
por: Chang, Yue, et al.
Publicado: (2026)
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation
por: Wang, Sheng
Publicado: (2025)
por: Wang, Sheng
Publicado: (2025)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
por: Chow, Wei, et al.
Publicado: (2025)
por: Chow, Wei, et al.
Publicado: (2025)
Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding
por: Patratskiy, Maxim A., et al.
Publicado: (2025)
por: Patratskiy, Maxim A., et al.
Publicado: (2025)
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
por: Fan, Qingyu, et al.
Publicado: (2026)
por: Fan, Qingyu, et al.
Publicado: (2026)
Articulated 3D Scene Graphs for Open-World Mobile Manipulation
por: Büchner, Martin, et al.
Publicado: (2026)
por: Büchner, Martin, et al.
Publicado: (2026)
Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent
por: Yu, Che Rin, et al.
Publicado: (2025)
por: Yu, Che Rin, et al.
Publicado: (2025)
Physically Grounded Vision-Language Models for Robotic Manipulation
por: Gao, Jensen, et al.
Publicado: (2023)
por: Gao, Jensen, et al.
Publicado: (2023)
A Navigation Framework Utilizing Vision-Language Models
por: Duan, Yicheng, et al.
Publicado: (2025)
por: Duan, Yicheng, et al.
Publicado: (2025)
Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement
por: Feng, Zhicheng, et al.
Publicado: (2025)
por: Feng, Zhicheng, et al.
Publicado: (2025)
On Deep Learning for Geometric and Semantic Scene Understanding Using On-Vehicle 3D LiDAR
por: Li, Li
Publicado: (2024)
por: Li, Li
Publicado: (2024)
SLAG: Scalable Language-Augmented Gaussian Splatting
por: Szilagyi, Laszlo, et al.
Publicado: (2025)
por: Szilagyi, Laszlo, et al.
Publicado: (2025)
A Survey on Vision-Language-Action Models for Autonomous Driving
por: Jiang, Sicong, et al.
Publicado: (2025)
por: Jiang, Sicong, et al.
Publicado: (2025)
LIBERO-X: Robustness Litmus for Vision-Language-Action Models
por: Wang, Guodong, et al.
Publicado: (2026)
por: Wang, Guodong, et al.
Publicado: (2026)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
por: Tang, Zuojin, et al.
Publicado: (2026)
por: Tang, Zuojin, et al.
Publicado: (2026)
Semantic Enrichment of CAD-Based Industrial Environments via Scene Graphs for Simulation and Reasoning
por: Walus, Nathan Pascal, et al.
Publicado: (2026)
por: Walus, Nathan Pascal, et al.
Publicado: (2026)
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots
por: Modi, Giorgia, et al.
Publicado: (2026)
por: Modi, Giorgia, et al.
Publicado: (2026)
Point2Graph: An End-to-end Point Cloud-based 3D Open-Vocabulary Scene Graph for Robot Navigation
por: Xu, Yifan, et al.
Publicado: (2024)
por: Xu, Yifan, et al.
Publicado: (2024)
ChainFlow-VLA: Causal Flow Planning with Vision-Language Models
por: Wang, Xiyang, et al.
Publicado: (2026)
por: Wang, Xiyang, et al.
Publicado: (2026)
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
por: Li, Anqi, et al.
Publicado: (2025)
por: Li, Anqi, et al.
Publicado: (2025)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
por: Kim, Ju-Young, et al.
Publicado: (2025)
por: Kim, Ju-Young, et al.
Publicado: (2025)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
por: Guo, Jianing, et al.
Publicado: (2025)
por: Guo, Jianing, et al.
Publicado: (2025)
Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models
por: Martinez-Sanchez, Angel, et al.
Publicado: (2026)
por: Martinez-Sanchez, Angel, et al.
Publicado: (2026)
Fixed External Cameras as Common Prior Maps for Active 3D Scene Graph Generation
por: Modi, Giorgia, et al.
Publicado: (2026)
por: Modi, Giorgia, et al.
Publicado: (2026)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
por: Zhang, Yi, et al.
Publicado: (2025)
por: Zhang, Yi, et al.
Publicado: (2025)
DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving
por: Kirchner, Sven, et al.
Publicado: (2025)
por: Kirchner, Sven, et al.
Publicado: (2025)
A Brief Survey on Leveraging Large Scale Vision Models for Enhanced Robot Grasping
por: Kamboj, Abhi, et al.
Publicado: (2024)
por: Kamboj, Abhi, et al.
Publicado: (2024)
Ejemplares similares
-
EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video Synthesis
por: Fang, Jianwu, et al.
Publicado: (2025) -
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
por: Zhang, Jiwen, et al.
Publicado: (2026) -
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
por: Li, Haoyuan, et al.
Publicado: (2025) -
Explainable Scene Understanding with Qualitative Representations and Graph Neural Networks
por: Belmecheri, Nassim, et al.
Publicado: (2025) -
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
por: Gao, Haoxiang, et al.
Publicado: (2025)