A Comparative Evaluation of Large Vision-Language Models for 2D Object Detection under SOTIF Conditions
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Ji, Ding, Yilin, Zhao, Yongqi, Xu, Jiachen, Eichberger, Arno |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chat2Scenario: Scenario Extraction From Dataset Through Utilization of Large Language Model
by: Zhao, Yongqi, et al.
Published: (2024)
by: Zhao, Yongqi, et al.
Published: (2024)
FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidanc
by: Tian, Yongqi, et al.
Published: (2025)
by: Tian, Yongqi, et al.
Published: (2025)
DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models
by: Huang, Shucheng, et al.
Published: (2025)
by: Huang, Shucheng, et al.
Published: (2025)
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
by: Zhang, Borong, et al.
Published: (2025)
by: Zhang, Borong, et al.
Published: (2025)
FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection
by: Yang, Anqi Joyce, et al.
Published: (2026)
by: Yang, Anqi Joyce, et al.
Published: (2026)
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
by: Zhou, Xueyang, et al.
Published: (2025)
by: Zhou, Xueyang, et al.
Published: (2025)
Cutting-Edge Detection of Fatigue in Drivers: A Comparative Study of Object Detection Models
by: Jones, Amelia
Published: (2024)
by: Jones, Amelia
Published: (2024)
Object-Scene-Camera Decomposition and Recomposition for Data-Efficient Monocular 3D Object Detection
by: Kuang, Zhaonian, et al.
Published: (2026)
by: Kuang, Zhaonian, et al.
Published: (2026)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
by: Wu, Pengying, et al.
Published: (2024)
by: Wu, Pengying, et al.
Published: (2024)
Anyview: Generalizable Indoor 3D Object Detection with Variable Frames
by: Wu, Zhenyu, et al.
Published: (2023)
by: Wu, Zhenyu, et al.
Published: (2023)
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
by: Zhou, Jiaying, et al.
Published: (2026)
by: Zhou, Jiaying, et al.
Published: (2026)
WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation
by: Nie, Dujun, et al.
Published: (2025)
by: Nie, Dujun, et al.
Published: (2025)
Inverse++: Vision-Centric 3D Semantic Occupancy Prediction Assisted with 3D Object Detection
by: Ming, Zhenxing, et al.
Published: (2025)
by: Ming, Zhenxing, et al.
Published: (2025)
Perspective-Invariant 3D Object Detection
by: Liang, Ao, et al.
Published: (2025)
by: Liang, Ao, et al.
Published: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Scalable Vision-Based 3D Object Detection and Monocular Depth Estimation for Autonomous Driving
by: Liu, Yuxuan
Published: (2024)
by: Liu, Yuxuan
Published: (2024)
CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
by: Kuang, Zhaonian, et al.
Published: (2026)
by: Kuang, Zhaonian, et al.
Published: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
by: Xie, Haozhe, et al.
Published: (2026)
by: Xie, Haozhe, et al.
Published: (2026)
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
by: Song, Zirui, et al.
Published: (2025)
by: Song, Zirui, et al.
Published: (2025)
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
by: Kim, Taewhan, et al.
Published: (2024)
by: Kim, Taewhan, et al.
Published: (2024)
TARGO: Benchmarking Target-driven Object Grasping under Occlusions
by: Xia, Yan, et al.
Published: (2024)
by: Xia, Yan, et al.
Published: (2024)
SDCM: Simulated Densifying and Compensatory Modeling Fusion for Radar-Vision 3-D Object Detection in Internet of Vehicles
by: Li, Shucong, et al.
Published: (2026)
by: Li, Shucong, et al.
Published: (2026)
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
by: Zantout, Nader, et al.
Published: (2025)
by: Zantout, Nader, et al.
Published: (2025)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
by: Song, Wenxuan, et al.
Published: (2026)
by: Song, Wenxuan, et al.
Published: (2026)
From Words to Poses: Enhancing Novel Object Pose Estimation with Vision Language Models
by: Pulli, Tessa, et al.
Published: (2024)
by: Pulli, Tessa, et al.
Published: (2024)
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
by: Fang, Yu, et al.
Published: (2026)
by: Fang, Yu, et al.
Published: (2026)
DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models
by: Li, Chenyang, et al.
Published: (2026)
by: Li, Chenyang, et al.
Published: (2026)
Towards Open-World Grasping with Large Vision-Language Models
by: Tziafas, Georgios, et al.
Published: (2024)
by: Tziafas, Georgios, et al.
Published: (2024)
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
by: Nguyen, Nghia, et al.
Published: (2024)
by: Nguyen, Nghia, et al.
Published: (2024)
Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models
by: Kapelyukh, Ivan, et al.
Published: (2023)
by: Kapelyukh, Ivan, et al.
Published: (2023)
Evaluation of Large Language Models for Anomaly Detection in Autonomous Vehicles
by: Loukas, Petros, et al.
Published: (2025)
by: Loukas, Petros, et al.
Published: (2025)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory
by: Jin, Qunchao, et al.
Published: (2025)
by: Jin, Qunchao, et al.
Published: (2025)
DIO: Dataset of 3D Mesh Models of Indoor Objects for Robotics and Computer Vision Applications
by: Nimal, Nillan, et al.
Published: (2024)
by: Nimal, Nillan, et al.
Published: (2024)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
by: Han, Xiaofeng, et al.
Published: (2025)
by: Han, Xiaofeng, et al.
Published: (2025)
Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment
by: Liu, Kangcheng, et al.
Published: (2023)
by: Liu, Kangcheng, et al.
Published: (2023)
Query2Uncertainty: Robust Uncertainty Quantification and Calibration for 3D Object Detection under Distribution Shift
by: Beemelmanns, Till, et al.
Published: (2026)
by: Beemelmanns, Till, et al.
Published: (2026)
Similar Items
-
Chat2Scenario: Scenario Extraction From Dataset Through Utilization of Large Language Model
by: Zhao, Yongqi, et al.
Published: (2024) -
FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidanc
by: Tian, Yongqi, et al.
Published: (2025) -
DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models
by: Huang, Shucheng, et al.
Published: (2025) -
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
by: Zhang, Borong, et al.
Published: (2025) -
FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection
by: Yang, Anqi Joyce, et al.
Published: (2026)