Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Robicheaux, Peter, Popov, Matvei, Madan, Anish, Robinson, Isaac, Nelson, Joseph, Ramanan, Deva, Peri, Neehar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
von: Robinson, Isaac, et al.
Veröffentlicht: (2025)
von: Robinson, Isaac, et al.
Veröffentlicht: (2025)
Revisiting Few-Shot Object Detection with Vision-Language Models
von: Madan, Anish, et al.
Veröffentlicht: (2023)
von: Madan, Anish, et al.
Veröffentlicht: (2023)
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2026)
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2026)
RefAV: Towards Planning-Centric Scenario Mining
von: Davidson, Cainan, et al.
Veröffentlicht: (2025)
von: Davidson, Cainan, et al.
Veröffentlicht: (2025)
Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection
von: Khurana, Mehar, et al.
Veröffentlicht: (2024)
von: Khurana, Mehar, et al.
Veröffentlicht: (2024)
SMORE: Simultaneous Map and Object REconstruction
von: Chodosh, Nathaniel, et al.
Veröffentlicht: (2024)
von: Chodosh, Nathaniel, et al.
Veröffentlicht: (2024)
Long-Tailed 3D Detection via Multi-Modal Fusion
von: Ma, Yechi, et al.
Veröffentlicht: (2023)
von: Ma, Yechi, et al.
Veröffentlicht: (2023)
MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
I Can't Believe It's Not Scene Flow!
von: Khatri, Ishan, et al.
Veröffentlicht: (2024)
von: Khatri, Ishan, et al.
Veröffentlicht: (2024)
Better Call SAL: Towards Learning to Segment Anything in Lidar
von: Ošep, Aljoša, et al.
Veröffentlicht: (2024)
von: Ošep, Aljoša, et al.
Veröffentlicht: (2024)
UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Vehicles
von: Li, Siyi, et al.
Veröffentlicht: (2025)
von: Li, Siyi, et al.
Veröffentlicht: (2025)
Planning with Adaptive World Models for Autonomous Driving
von: Vasudevan, Arun Balajee, et al.
Veröffentlicht: (2024)
von: Vasudevan, Arun Balajee, et al.
Veröffentlicht: (2024)
Revisiting the Role of Language Priors in Vision-Language Models
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
ZeroFlow: Scalable Scene Flow via Distillation
von: Vedder, Kyle, et al.
Veröffentlicht: (2023)
von: Vedder, Kyle, et al.
Veröffentlicht: (2023)
Neural Eulerian Scene Flow Fields
von: Vedder, Kyle, et al.
Veröffentlicht: (2024)
von: Vedder, Kyle, et al.
Veröffentlicht: (2024)
Predicting Long-horizon Futures by Conditioning on Geometry and Time
von: Khurana, Tarasha, et al.
Veröffentlicht: (2024)
von: Khurana, Tarasha, et al.
Veröffentlicht: (2024)
TAO-Amodal: A Benchmark for Tracking Any Object Amodally
von: Hsieh, Cheng-Yen, et al.
Veröffentlicht: (2023)
von: Hsieh, Cheng-Yen, et al.
Veröffentlicht: (2023)
Language Models as Black-Box Optimizers for Vision-Language Models
von: Liu, Shihong, et al.
Veröffentlicht: (2023)
von: Liu, Shihong, et al.
Veröffentlicht: (2023)
StableTrack: Stabilizing Multi-Object Tracking on Low-Frequency Detections
von: Shelukhan, Matvei, et al.
Veröffentlicht: (2025)
von: Shelukhan, Matvei, et al.
Veröffentlicht: (2025)
The Neglected Tails in Vision-Language Models
von: Parashar, Shubham, et al.
Veröffentlicht: (2024)
von: Parashar, Shubham, et al.
Veröffentlicht: (2024)
PAI-Bench: A Comprehensive Benchmark For Physical AI
von: Zhou, Fengzhe, et al.
Veröffentlicht: (2025)
von: Zhou, Fengzhe, et al.
Veröffentlicht: (2025)
Using Diffusion Priors for Video Amodal Segmentation
von: Chen, Kaihua, et al.
Veröffentlicht: (2024)
von: Chen, Kaihua, et al.
Veröffentlicht: (2024)
Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
von: Chen, Kaihua, et al.
Veröffentlicht: (2025)
von: Chen, Kaihua, et al.
Veröffentlicht: (2025)
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
von: Dong, Daxiang, et al.
Veröffentlicht: (2025)
von: Dong, Daxiang, et al.
Veröffentlicht: (2025)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models
von: Zhu, He, et al.
Veröffentlicht: (2025)
von: Zhu, He, et al.
Veröffentlicht: (2025)
VL-RouterBench: A Benchmark for Vision-Language Model Routing
von: Huang, Zhehao, et al.
Veröffentlicht: (2025)
von: Huang, Zhehao, et al.
Veröffentlicht: (2025)
B-SMALL: A Bayesian Neural Network approach to Sparse Model-Agnostic Meta-Learning
von: Madan, Anish, et al.
Veröffentlicht: (2021)
von: Madan, Anish, et al.
Veröffentlicht: (2021)
Evaluating a VR System for Collecting Safety-Critical Vehicle-Pedestrian Interactions
von: Weng, Erica, et al.
Veröffentlicht: (2023)
von: Weng, Erica, et al.
Veröffentlicht: (2023)
ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion
von: Duisterhof, Bardienus P., et al.
Veröffentlicht: (2025)
von: Duisterhof, Bardienus P., et al.
Veröffentlicht: (2025)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
von: Deria, Ankan, et al.
Veröffentlicht: (2026)
von: Deria, Ankan, et al.
Veröffentlicht: (2026)
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
AgriGPT-VL: Agricultural Vision-Language Understanding Suite
von: Yang, Bo, et al.
Veröffentlicht: (2025)
von: Yang, Bo, et al.
Veröffentlicht: (2025)
VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2024)
VL4AD: Vision-Language Models Improve Pixel-wise Anomaly Detection
von: Zhong, Liangyu, et al.
Veröffentlicht: (2024)
von: Zhong, Liangyu, et al.
Veröffentlicht: (2024)
Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association
von: Shelukhan, Matvei, et al.
Veröffentlicht: (2026)
von: Shelukhan, Matvei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
von: Robinson, Isaac, et al.
Veröffentlicht: (2025) -
Revisiting Few-Shot Object Detection with Vision-Language Models
von: Madan, Anish, et al.
Veröffentlicht: (2023) -
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2026) -
RefAV: Towards Planning-Centric Scenario Mining
von: Davidson, Cainan, et al.
Veröffentlicht: (2025) -
Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection
von: Khurana, Mehar, et al.
Veröffentlicht: (2024)