When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Author: | Al-Hamadani, Samer |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Intelligent Healthcare Imaging Platform: A VLM-Based Framework for Automated Medical Image Analysis and Clinical Report Generation
by: Al-Hamadani, Samer
Published: (2025)
by: Al-Hamadani, Samer
Published: (2025)
Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
by: Miyai, Atsuyuki, et al.
Published: (2024)
by: Miyai, Atsuyuki, et al.
Published: (2024)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
ClipGrader: Leveraging Vision-Language Models for Robust Label Quality Assessment in Object Detection
by: Lu, Hong, et al.
Published: (2025)
by: Lu, Hong, et al.
Published: (2025)
HASSOD: Hierarchical Adaptive Self-Supervised Object Detection
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
by: Seo, Soo Won, et al.
Published: (2026)
by: Seo, Soo Won, et al.
Published: (2026)
Unified Supervision For Vision-Language Modeling in 3D Computed Tomography
by: Lee, Hao-Chih, et al.
Published: (2025)
by: Lee, Hao-Chih, et al.
Published: (2025)
When Multi-Task Learning Meets Partial Supervision: A Computer Vision Review
by: Fontana, Maxime, et al.
Published: (2023)
by: Fontana, Maxime, et al.
Published: (2023)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
by: Li, Zhuowei, et al.
Published: (2025)
by: Li, Zhuowei, et al.
Published: (2025)
Edge AI: Evaluation of Model Compression Techniques for Convolutional Neural Networks
by: Francy, Samer, et al.
Published: (2024)
by: Francy, Samer, et al.
Published: (2024)
Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
by: Mann, Logan, et al.
Published: (2026)
by: Mann, Logan, et al.
Published: (2026)
Deep Vision-Based Framework for Coastal Flood Prediction Under Climate Change Impacts and Shoreline Adaptations
by: Karapetyan, Areg, et al.
Published: (2024)
by: Karapetyan, Areg, et al.
Published: (2024)
Hybrid Training for Vision-Language-Action Models
by: Mazzaglia, Pietro, et al.
Published: (2025)
by: Mazzaglia, Pietro, et al.
Published: (2025)
Hardness-Aware Scene Synthesis for Semi-Supervised 3D Object Detection
by: Zeng, Shuai, et al.
Published: (2024)
by: Zeng, Shuai, et al.
Published: (2024)
Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration
by: Xu, Yuanzhi, et al.
Published: (2026)
by: Xu, Yuanzhi, et al.
Published: (2026)
Evaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities
by: Datta, Shounak, et al.
Published: (2025)
by: Datta, Shounak, et al.
Published: (2025)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
by: Lee, Jihoon, et al.
Published: (2025)
by: Lee, Jihoon, et al.
Published: (2025)
When Training-Free NAS Meets Vision Transformer: A Neural Tangent Kernel Perspective
by: Zhou, Qiqi, et al.
Published: (2024)
by: Zhou, Qiqi, et al.
Published: (2024)
Prompting the Unseen: Detecting Hidden Backdoors in Black-Box Models
by: Huang, Zi-Xuan, et al.
Published: (2024)
by: Huang, Zi-Xuan, et al.
Published: (2024)
Harnessing Vision-Language Models for Time Series Anomaly Detection
by: He, Zelin, et al.
Published: (2025)
by: He, Zelin, et al.
Published: (2025)
Margin and Consistency Supervision for Calibrated and Robust Vision Models
by: Khazem, Salim
Published: (2026)
by: Khazem, Salim
Published: (2026)
Visual Modality Prompt for Adapting Vision-Language Object Detectors
by: Medeiros, Heitor R., et al.
Published: (2024)
by: Medeiros, Heitor R., et al.
Published: (2024)
Interactive Post-Training for Vision-Language-Action Models
by: Tan, Shuhan, et al.
Published: (2025)
by: Tan, Shuhan, et al.
Published: (2025)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
by: Kaul, Prannay, et al.
Published: (2024)
by: Kaul, Prannay, et al.
Published: (2024)
Segment Concealed Objects with Incomplete Supervision
by: He, Chunming, et al.
Published: (2025)
by: He, Chunming, et al.
Published: (2025)
Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection
by: Yu, Geng, et al.
Published: (2024)
by: Yu, Geng, et al.
Published: (2024)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025)
by: Lee, Heekyung, et al.
Published: (2025)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
by: Huang, Po-Hsuan, et al.
Published: (2024)
by: Huang, Po-Hsuan, et al.
Published: (2024)
Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
by: Li, Yihao, et al.
Published: (2025)
by: Li, Yihao, et al.
Published: (2025)
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
by: Roschmann, Simon, et al.
Published: (2025)
by: Roschmann, Simon, et al.
Published: (2025)
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
by: Li, Zhiqi, et al.
Published: (2025)
by: Li, Zhiqi, et al.
Published: (2025)
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
by: Chen, Yi-Chia, et al.
Published: (2024)
by: Chen, Yi-Chia, et al.
Published: (2024)
AdaNeg: Adaptive Negative Proxy Guided OOD Detection with Vision-Language Models
by: Zhang, Yabin, et al.
Published: (2024)
by: Zhang, Yabin, et al.
Published: (2024)
VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Models
by: Ye, Muchao, et al.
Published: (2024)
by: Ye, Muchao, et al.
Published: (2024)
LAPT: Label-driven Automated Prompt Tuning for OOD Detection with Vision-Language Models
by: Zhang, Yabin, et al.
Published: (2024)
by: Zhang, Yabin, et al.
Published: (2024)
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
by: Wu, Junfei, et al.
Published: (2024)
by: Wu, Junfei, et al.
Published: (2024)
AI-Powered Deepfake Detection Using CNN and Vision Transformer Architectures
by: Urmi, Sifatullah Sheikh, et al.
Published: (2026)
by: Urmi, Sifatullah Sheikh, et al.
Published: (2026)
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
by: Ramachandran, Rahul, et al.
Published: (2025)
by: Ramachandran, Rahul, et al.
Published: (2025)
SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
by: Yousaf, Adeel, et al.
Published: (2025)
by: Yousaf, Adeel, et al.
Published: (2025)
Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Spatial-Relation Data
by: Yamamoto, Takaki, et al.
Published: (2026)
by: Yamamoto, Takaki, et al.
Published: (2026)
Similar Items
-
Intelligent Healthcare Imaging Platform: A VLM-Based Framework for Automated Medical Image Analysis and Clinical Report Generation
by: Al-Hamadani, Samer
Published: (2025) -
Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
by: Miyai, Atsuyuki, et al.
Published: (2024) -
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025) -
ClipGrader: Leveraging Vision-Language Models for Robust Label Quality Assessment in Object Detection
by: Lu, Hong, et al.
Published: (2025) -
HASSOD: Hierarchical Adaptive Self-Supervised Object Detection
by: Cao, Shengcao, et al.
Published: (2024)