Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Ying, Hua, Yijing, Chai, Haojiang, Wang, Yanbo, Ye, TengQi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CerberusDet: Unified Multi-Dataset Object Detection
by: Tolstykh, Irina, et al.
Published: (2024)
by: Tolstykh, Irina, et al.
Published: (2024)
Proto-FG3D: Prototype-based Interpretable Fine-Grained 3D Shape Classification
by: Ma, Shuxian, et al.
Published: (2025)
by: Ma, Shuxian, et al.
Published: (2025)
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
by: Jin, Weiyang, et al.
Published: (2025)
by: Jin, Weiyang, et al.
Published: (2025)
MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
EncQA: Benchmarking Vision-Language Models on Visual Encodings for Charts
by: Mukherjee, Kushin, et al.
Published: (2025)
by: Mukherjee, Kushin, et al.
Published: (2025)
Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex
by: Grosbard, Idan Daniel, et al.
Published: (2026)
by: Grosbard, Idan Daniel, et al.
Published: (2026)
Efficient Diffusion Training through Parallelization with Truncated Karhunen-Loève Expansion
by: Ren, Yumeng, et al.
Published: (2025)
by: Ren, Yumeng, et al.
Published: (2025)
Synthetic Photography Detection: A Visual Guidance for Identifying Synthetic Images Created by AI
by: Mathys, Melanie, et al.
Published: (2024)
by: Mathys, Melanie, et al.
Published: (2024)
SeNeDiF-OOD: Semantic Nested Dichotomy Fusion for Out-of-Distribution Detection Methodology in Open-World Classification. A Case Study on Monument Style Classification
by: Antequera-Sánchez, Ignacio, et al.
Published: (2026)
by: Antequera-Sánchez, Ignacio, et al.
Published: (2026)
GTPBD: A Fine-Grained Global Terraced Parcel and Boundary Dataset
by: Zhang, Zhiwei, et al.
Published: (2025)
by: Zhang, Zhiwei, et al.
Published: (2025)
Measuring proximity to standard planes during fetal brain ultrasound scanning
by: Di Vece, Chiara, et al.
Published: (2024)
by: Di Vece, Chiara, et al.
Published: (2024)
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)
by: He, Mengqi, et al.
Published: (2025)
Context-Aware Full Body Anonymization using Text-to-Image Diffusion Models
by: Zwick, Pascal, et al.
Published: (2024)
by: Zwick, Pascal, et al.
Published: (2024)
MAR-MAER: Metric-Aware and Ambiguity-Adaptive Autoregressive Image Generation
by: Dong, Kai, et al.
Published: (2026)
by: Dong, Kai, et al.
Published: (2026)
BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos
by: Lin, Jiahao, et al.
Published: (2025)
by: Lin, Jiahao, et al.
Published: (2025)
Parameterizing Dataset Distillation via Gaussian Splatting
by: Jiang, Chenyang, et al.
Published: (2025)
by: Jiang, Chenyang, et al.
Published: (2025)
Exploring Transfer Learning for Deep Learning Polyp Detection in Colonoscopy Images Using YOLOv8
by: Vazquez, Fabian, et al.
Published: (2025)
by: Vazquez, Fabian, et al.
Published: (2025)
Domain Generalized Stereo Matching with Uncertainty-guided Data Augmentation
by: Du, Shuangli, et al.
Published: (2025)
by: Du, Shuangli, et al.
Published: (2025)
Quaternion Convolutional Neural Networks: Current Advances and Future Directions
by: Altamirano-Gomez, Gerardo, et al.
Published: (2023)
by: Altamirano-Gomez, Gerardo, et al.
Published: (2023)
A Review of Pseudo-Labeling for Computer Vision
by: Kage, Patrick, et al.
Published: (2024)
by: Kage, Patrick, et al.
Published: (2024)
Sparse vs Contiguous Adversarial Pixel Perturbations in Multimodal Models: An Empirical Analysis
by: Botocan, Cristian-Alexandru, et al.
Published: (2024)
by: Botocan, Cristian-Alexandru, et al.
Published: (2024)
GeoPos: A Minimal Positional Encoding for Enhanced Fine-Grained Details in Image Synthesis Using Convolutional Neural Networks
by: Hosseini, Mehran, et al.
Published: (2024)
by: Hosseini, Mehran, et al.
Published: (2024)
Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
by: Ma, Chong, et al.
Published: (2024)
by: Ma, Chong, et al.
Published: (2024)
FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare
by: Lekadir, Karim, et al.
Published: (2023)
by: Lekadir, Karim, et al.
Published: (2023)
Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image Generation
by: Lu, Zhuqiang, et al.
Published: (2023)
by: Lu, Zhuqiang, et al.
Published: (2023)
Towards Infusing Auxiliary Knowledge for Distracted Driver Detection
by: Balappanawar, Ishwar B, et al.
Published: (2024)
by: Balappanawar, Ishwar B, et al.
Published: (2024)
Beyond Specialization: Assessing the Capabilities of MLLMs in Age and Gender Estimation
by: Kuprashevich, Maksim, et al.
Published: (2024)
by: Kuprashevich, Maksim, et al.
Published: (2024)
Generative inpainting of incomplete Euclidean distance matrices of trajectories generated by a fractional Brownian motion
by: Lobashev, Alexander, et al.
Published: (2024)
by: Lobashev, Alexander, et al.
Published: (2024)
Generating Image Adversarial Examples by Embedding Digital Watermarks
by: Xiang, Yuexin, et al.
Published: (2020)
by: Xiang, Yuexin, et al.
Published: (2020)
Generation of Complex 3D Human Motion by Temporal and Spatial Composition of Diffusion Models
by: Mandelli, Lorenzo, et al.
Published: (2024)
by: Mandelli, Lorenzo, et al.
Published: (2024)
CoMA: Complementary Masking and Hierarchical Dynamic Multi-Window Self-Attention in a Unified Pre-training Framework
by: Li, Jiaxuan, et al.
Published: (2025)
by: Li, Jiaxuan, et al.
Published: (2025)
Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets
by: Tang, MingZe, et al.
Published: (2025)
by: Tang, MingZe, et al.
Published: (2025)
Deep EM with Hierarchical Latent Label Modelling for Multi-Site Prostate Lesion Segmentation
by: Yan, Wen, et al.
Published: (2026)
by: Yan, Wen, et al.
Published: (2026)
How to Choose Your Teacher for Fine Grained Image Recognition
by: Gosal, Oswin, et al.
Published: (2026)
by: Gosal, Oswin, et al.
Published: (2026)
Synthetic Image Generation in Cyber Influence Operations: An Emergent Threat?
by: Mathys, Melanie, et al.
Published: (2024)
by: Mathys, Melanie, et al.
Published: (2024)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
by: Rudman, William, et al.
Published: (2026)
by: Rudman, William, et al.
Published: (2026)
Comparing Zealous and Restrained AI Recommendations in a Real-World Human-AI Collaboration Task
by: Xu, Chengyuan, et al.
Published: (2024)
by: Xu, Chengyuan, et al.
Published: (2024)
Skeleton-based sign language recognition using a dual-stream spatio-temporal dynamic graph convolutional network
by: Liu, Liangjin, et al.
Published: (2025)
by: Liu, Liangjin, et al.
Published: (2025)
Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts
by: Tangtartharakul, Gene, et al.
Published: (2026)
by: Tangtartharakul, Gene, et al.
Published: (2026)
Towards Interpretable Visual Decoding with Attention to Brain Representations
by: Feng, Pinyuan, et al.
Published: (2025)
by: Feng, Pinyuan, et al.
Published: (2025)
Similar Items
-
CerberusDet: Unified Multi-Dataset Object Detection
by: Tolstykh, Irina, et al.
Published: (2024) -
Proto-FG3D: Prototype-based Interpretable Fine-Grained 3D Shape Classification
by: Ma, Shuxian, et al.
Published: (2025) -
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
by: Jin, Weiyang, et al.
Published: (2025) -
MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
by: Wang, Han, et al.
Published: (2024) -
EncQA: Benchmarking Vision-Language Models on Visual Encodings for Charts
by: Mukherjee, Kushin, et al.
Published: (2025)