Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deka, Dawar Jyoti, Sethi, Amit, Ali, Syed Mohammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
von: Deng, Pei, et al.
Veröffentlicht: (2025)
von: Deng, Pei, et al.
Veröffentlicht: (2025)
Visible Iris Area as a Quality Metric for Reliable Iris Recognition Under Pupil Dilation and Eyelid Occlusion
von: Pessaud, Jack, et al.
Veröffentlicht: (2025)
von: Pessaud, Jack, et al.
Veröffentlicht: (2025)
Capacity Constraint Analysis Using Object Detection for Smart Manufacturing
von: Ahmad, Hafiz Mughees, et al.
Veröffentlicht: (2024)
von: Ahmad, Hafiz Mughees, et al.
Veröffentlicht: (2024)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
von: Huang, Yian, et al.
Veröffentlicht: (2026)
von: Huang, Yian, et al.
Veröffentlicht: (2026)
Habitat Classification from Ground-Level Imagery Using Deep Neural Networks
von: Shi, Hongrui, et al.
Veröffentlicht: (2025)
von: Shi, Hongrui, et al.
Veröffentlicht: (2025)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
von: Li, Yayuan, et al.
Veröffentlicht: (2025)
von: Li, Yayuan, et al.
Veröffentlicht: (2025)
SelvaBox: A high-resolution dataset for tropical tree crown detection
von: Baudchon, Hugo, et al.
Veröffentlicht: (2025)
von: Baudchon, Hugo, et al.
Veröffentlicht: (2025)
Image-Based Leopard Seal Recognition: Approaches and Challenges in Current Automated Systems
von: Salazar, Jorge Yero, et al.
Veröffentlicht: (2024)
von: Salazar, Jorge Yero, et al.
Veröffentlicht: (2024)
Dense Motion Captioning
von: Xu, Shiyao, et al.
Veröffentlicht: (2025)
von: Xu, Shiyao, et al.
Veröffentlicht: (2025)
TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
von: Lee, Byung Hoon, et al.
Veröffentlicht: (2025)
von: Lee, Byung Hoon, et al.
Veröffentlicht: (2025)
CoMatcher: Multi-View Collaborative Feature Matching
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
MoDE: Mixture of Diffusion Experts for Any Occluded Face Recognition
von: Fan, Qiannan, et al.
Veröffentlicht: (2025)
von: Fan, Qiannan, et al.
Veröffentlicht: (2025)
NeuroGaze-Distill: Brain-informed Distillation and Depression-Inspired Geometric Priors for Robust Facial Emotion Recognition
von: Li, Zilin, et al.
Veröffentlicht: (2025)
von: Li, Zilin, et al.
Veröffentlicht: (2025)
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
von: Diller, Christian, et al.
Veröffentlicht: (2023)
von: Diller, Christian, et al.
Veröffentlicht: (2023)
SelvaMask: Segmenting Trees in Tropical Forests and Beyond
von: Duguay, Simon-Olivier, et al.
Veröffentlicht: (2026)
von: Duguay, Simon-Olivier, et al.
Veröffentlicht: (2026)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
Do All Vision Transformers Need Registers? A Cross-Architectural Reassessment
von: Baxevanakis, Spiros, et al.
Veröffentlicht: (2026)
von: Baxevanakis, Spiros, et al.
Veröffentlicht: (2026)
PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
von: Yang, Zongyou, et al.
Veröffentlicht: (2025)
von: Yang, Zongyou, et al.
Veröffentlicht: (2025)
SH17: A Dataset for Human Safety and Personal Protective Equipment Detection in Manufacturing Industry
von: Ahmad, Hafiz Mughees, et al.
Veröffentlicht: (2024)
von: Ahmad, Hafiz Mughees, et al.
Veröffentlicht: (2024)
Efficient Temporally-Aware DeepFake Detection using H.264 Motion Vectors
von: Grönquist, Peter, et al.
Veröffentlicht: (2023)
von: Grönquist, Peter, et al.
Veröffentlicht: (2023)
WaveMix: A Resource-efficient Neural Network for Image Analysis
von: Jeevan, Pranav, et al.
Veröffentlicht: (2022)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2022)
FLD+: Data-efficient Evaluation Metric for Generative Models
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
Normalizing Flow-Based Metric for Image Generation
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
von: Dong, Haohua, et al.
Veröffentlicht: (2025)
von: Dong, Haohua, et al.
Veröffentlicht: (2025)
HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding
von: Jin, Haopeng, et al.
Veröffentlicht: (2026)
von: Jin, Haopeng, et al.
Veröffentlicht: (2026)
Towards Accurate and Efficient Waste Image Classification: A Hybrid Deep Learning and Machine Learning Approach
von: Nguyen, Ngoc-Bao-Quang, et al.
Veröffentlicht: (2025)
von: Nguyen, Ngoc-Bao-Quang, et al.
Veröffentlicht: (2025)
Dynamic Arthroscopic Navigation System for Anterior Cruciate Ligament Reconstruction Based on Multi-level Memory Architecture
von: Wang, Shuo, et al.
Veröffentlicht: (2025)
von: Wang, Shuo, et al.
Veröffentlicht: (2025)
Detecting AI-Generated Videos with Spiking Neural Networks
von: Jang, Minsuk, et al.
Veröffentlicht: (2026)
von: Jang, Minsuk, et al.
Veröffentlicht: (2026)
Distributed Intelligent System Architecture for UAV-Assisted Monitoring of Wind Energy Infrastructure
von: Svystun, Serhii, et al.
Veröffentlicht: (2024)
von: Svystun, Serhii, et al.
Veröffentlicht: (2024)
EgoTraj: Real-World Egocentric Human Trajectory Dataset for Multimodal Prediction
von: Yehia, Ahmad, et al.
Veröffentlicht: (2026)
von: Yehia, Ahmad, et al.
Veröffentlicht: (2026)
Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video Observations
von: Diller, Christian, et al.
Veröffentlicht: (2022)
von: Diller, Christian, et al.
Veröffentlicht: (2022)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
von: Brovko, D. V.
Veröffentlicht: (2025)
von: Brovko, D. V.
Veröffentlicht: (2025)
A Light Perspective for 3D Object Detection
von: Pederiva, Marcelo Eduardo, et al.
Veröffentlicht: (2025)
von: Pederiva, Marcelo Eduardo, et al.
Veröffentlicht: (2025)
Distance Estimation in Outdoor Driving Environments Using Phase-only Correlation Method with Event Cameras
von: Kobayashi, Masataka, et al.
Veröffentlicht: (2025)
von: Kobayashi, Masataka, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025) -
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
von: Wang, Yiming, et al.
Veröffentlicht: (2026) -
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
von: Deng, Pei, et al.
Veröffentlicht: (2025) -
Visible Iris Area as a Quality Metric for Reliable Iris Recognition Under Pupil Dilation and Eyelid Occlusion
von: Pessaud, Jack, et al.
Veröffentlicht: (2025) -
Capacity Constraint Analysis Using Object Detection for Smart Manufacturing
von: Ahmad, Hafiz Mughees, et al.
Veröffentlicht: (2024)