A Comparative Evaluation of Large Vision-Language Models for 2D Object Detection under SOTIF Conditions
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Ji, Ding, Yilin, Zhao, Yongqi, Xu, Jiachen, Eichberger, Arno |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Chat2Scenario: Scenario Extraction From Dataset Through Utilization of Large Language Model
por: Zhao, Yongqi, et al.
Publicado: (2024)
por: Zhao, Yongqi, et al.
Publicado: (2024)
FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidanc
por: Tian, Yongqi, et al.
Publicado: (2025)
por: Tian, Yongqi, et al.
Publicado: (2025)
DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models
por: Huang, Shucheng, et al.
Publicado: (2025)
por: Huang, Shucheng, et al.
Publicado: (2025)
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
por: Zhang, Borong, et al.
Publicado: (2025)
por: Zhang, Borong, et al.
Publicado: (2025)
FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection
por: Yang, Anqi Joyce, et al.
Publicado: (2026)
por: Yang, Anqi Joyce, et al.
Publicado: (2026)
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
por: Zhou, Xueyang, et al.
Publicado: (2025)
por: Zhou, Xueyang, et al.
Publicado: (2025)
Cutting-Edge Detection of Fatigue in Drivers: A Comparative Study of Object Detection Models
por: Jones, Amelia
Publicado: (2024)
por: Jones, Amelia
Publicado: (2024)
Object-Scene-Camera Decomposition and Recomposition for Data-Efficient Monocular 3D Object Detection
por: Kuang, Zhaonian, et al.
Publicado: (2026)
por: Kuang, Zhaonian, et al.
Publicado: (2026)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
por: Bhat, Vineet, et al.
Publicado: (2025)
por: Bhat, Vineet, et al.
Publicado: (2025)
VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
por: Wu, Pengying, et al.
Publicado: (2024)
por: Wu, Pengying, et al.
Publicado: (2024)
Anyview: Generalizable Indoor 3D Object Detection with Variable Frames
por: Wu, Zhenyu, et al.
Publicado: (2023)
por: Wu, Zhenyu, et al.
Publicado: (2023)
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
por: Zhou, Jiaying, et al.
Publicado: (2026)
por: Zhou, Jiaying, et al.
Publicado: (2026)
WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation
por: Nie, Dujun, et al.
Publicado: (2025)
por: Nie, Dujun, et al.
Publicado: (2025)
Inverse++: Vision-Centric 3D Semantic Occupancy Prediction Assisted with 3D Object Detection
por: Ming, Zhenxing, et al.
Publicado: (2025)
por: Ming, Zhenxing, et al.
Publicado: (2025)
Perspective-Invariant 3D Object Detection
por: Liang, Ao, et al.
Publicado: (2025)
por: Liang, Ao, et al.
Publicado: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
por: Song, Wenxuan, et al.
Publicado: (2025)
por: Song, Wenxuan, et al.
Publicado: (2025)
Scalable Vision-Based 3D Object Detection and Monocular Depth Estimation for Autonomous Driving
por: Liu, Yuxuan
Publicado: (2024)
por: Liu, Yuxuan
Publicado: (2024)
CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
por: Kuang, Zhaonian, et al.
Publicado: (2026)
por: Kuang, Zhaonian, et al.
Publicado: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
por: Ding, Pengxiang, et al.
Publicado: (2023)
por: Ding, Pengxiang, et al.
Publicado: (2023)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
por: Chen, Jiayi, et al.
Publicado: (2025)
por: Chen, Jiayi, et al.
Publicado: (2025)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
por: Xie, Haozhe, et al.
Publicado: (2026)
por: Xie, Haozhe, et al.
Publicado: (2026)
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
por: Song, Zirui, et al.
Publicado: (2025)
por: Song, Zirui, et al.
Publicado: (2025)
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
por: Kim, Taewhan, et al.
Publicado: (2024)
por: Kim, Taewhan, et al.
Publicado: (2024)
TARGO: Benchmarking Target-driven Object Grasping under Occlusions
por: Xia, Yan, et al.
Publicado: (2024)
por: Xia, Yan, et al.
Publicado: (2024)
SDCM: Simulated Densifying and Compensatory Modeling Fusion for Radar-Vision 3-D Object Detection in Internet of Vehicles
por: Li, Shucong, et al.
Publicado: (2026)
por: Li, Shucong, et al.
Publicado: (2026)
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
por: Zantout, Nader, et al.
Publicado: (2025)
por: Zantout, Nader, et al.
Publicado: (2025)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
por: Song, Wenxuan, et al.
Publicado: (2026)
por: Song, Wenxuan, et al.
Publicado: (2026)
From Words to Poses: Enhancing Novel Object Pose Estimation with Vision Language Models
por: Pulli, Tessa, et al.
Publicado: (2024)
por: Pulli, Tessa, et al.
Publicado: (2024)
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
por: Fang, Yu, et al.
Publicado: (2026)
por: Fang, Yu, et al.
Publicado: (2026)
DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models
por: Li, Chenyang, et al.
Publicado: (2026)
por: Li, Chenyang, et al.
Publicado: (2026)
Towards Open-World Grasping with Large Vision-Language Models
por: Tziafas, Georgios, et al.
Publicado: (2024)
por: Tziafas, Georgios, et al.
Publicado: (2024)
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
por: Nguyen, Nghia, et al.
Publicado: (2024)
por: Nguyen, Nghia, et al.
Publicado: (2024)
Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models
por: Kapelyukh, Ivan, et al.
Publicado: (2023)
por: Kapelyukh, Ivan, et al.
Publicado: (2023)
Evaluation of Large Language Models for Anomaly Detection in Autonomous Vehicles
por: Loukas, Petros, et al.
Publicado: (2025)
por: Loukas, Petros, et al.
Publicado: (2025)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
por: Wang, Zhaowei, et al.
Publicado: (2024)
por: Wang, Zhaowei, et al.
Publicado: (2024)
PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory
por: Jin, Qunchao, et al.
Publicado: (2025)
por: Jin, Qunchao, et al.
Publicado: (2025)
DIO: Dataset of 3D Mesh Models of Indoor Objects for Robotics and Computer Vision Applications
por: Nimal, Nillan, et al.
Publicado: (2024)
por: Nimal, Nillan, et al.
Publicado: (2024)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
por: Han, Xiaofeng, et al.
Publicado: (2025)
por: Han, Xiaofeng, et al.
Publicado: (2025)
Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment
por: Liu, Kangcheng, et al.
Publicado: (2023)
por: Liu, Kangcheng, et al.
Publicado: (2023)
Query2Uncertainty: Robust Uncertainty Quantification and Calibration for 3D Object Detection under Distribution Shift
por: Beemelmanns, Till, et al.
Publicado: (2026)
por: Beemelmanns, Till, et al.
Publicado: (2026)
Ejemplares similares
-
Chat2Scenario: Scenario Extraction From Dataset Through Utilization of Large Language Model
por: Zhao, Yongqi, et al.
Publicado: (2024) -
FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidanc
por: Tian, Yongqi, et al.
Publicado: (2025) -
DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models
por: Huang, Shucheng, et al.
Publicado: (2025) -
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
por: Zhang, Borong, et al.
Publicado: (2025) -
FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection
por: Yang, Anqi Joyce, et al.
Publicado: (2026)