Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
Fuente:
arXiv
Saved in:
| Main Authors: | Demidov, Dmitry, Zaheer, Zaigham, Han, Zongyan, Thawakar, Omkar, Anwer, Rao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
by: Demidov, Dmitry, et al.
Published: (2025)
by: Demidov, Dmitry, et al.
Published: (2025)
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
CoVR-R:Reason-Aware Composed Video Retrieval
by: Thawakar, Omkar, et al.
Published: (2026)
by: Thawakar, Omkar, et al.
Published: (2026)
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
by: Han, Zongyan, et al.
Published: (2025)
by: Han, Zongyan, et al.
Published: (2025)
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts
by: Ghaboura, Sara, et al.
Published: (2025)
by: Ghaboura, Sara, et al.
Published: (2025)
DiffuseMix: Label-Preserving Data Augmentation with Diffusion Models
by: Islam, Khawar, et al.
Published: (2024)
by: Islam, Khawar, et al.
Published: (2024)
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
by: Ghaboura, Sara, et al.
Published: (2025)
by: Ghaboura, Sara, et al.
Published: (2025)
Extract More from Less: Efficient Fine-Grained Visual Recognition in Low-Data Regimes
by: Demidov, Dmitry, et al.
Published: (2024)
by: Demidov, Dmitry, et al.
Published: (2024)
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
by: Thawakar, Omkar, et al.
Published: (2024)
by: Thawakar, Omkar, et al.
Published: (2024)
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding
by: Patle, Shubham, et al.
Published: (2026)
by: Patle, Shubham, et al.
Published: (2026)
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
by: Thawakar, Omkar, et al.
Published: (2023)
by: Thawakar, Omkar, et al.
Published: (2023)
MIRA: A Novel Framework for Fusing Modalities in Medical RAG
by: Wang, Jinhong, et al.
Published: (2025)
by: Wang, Jinhong, et al.
Published: (2025)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
by: Nawaz, Umair, et al.
Published: (2025)
by: Nawaz, Umair, et al.
Published: (2025)
GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing
by: Islam, Khawar, et al.
Published: (2024)
by: Islam, Khawar, et al.
Published: (2024)
Face Pyramid Vision Transformer
by: Islam, Khawar, et al.
Published: (2022)
by: Islam, Khawar, et al.
Published: (2022)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
by: Dissanayake, Dinura, et al.
Published: (2025)
by: Dissanayake, Dinura, et al.
Published: (2025)
Constricting Normal Latent Space for Anomaly Detection with Normal-only Training Data
by: Astrid, Marcella, et al.
Published: (2024)
by: Astrid, Marcella, et al.
Published: (2024)
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
by: Kumar, Komal, et al.
Published: (2025)
by: Kumar, Komal, et al.
Published: (2025)
All in One: Visual-Description-Guided Unified Point Cloud Segmentation
by: Han, Zongyan, et al.
Published: (2025)
by: Han, Zongyan, et al.
Published: (2025)
Salient Mask-Guided Vision Transformer for Fine-Grained Classification
by: Demidov, Dmitry, et al.
Published: (2023)
by: Demidov, Dmitry, et al.
Published: (2023)
DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
by: Li, Geng, et al.
Published: (2025)
by: Li, Geng, et al.
Published: (2025)
LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation
by: Ahn, Jinwoo, et al.
Published: (2024)
by: Ahn, Jinwoo, et al.
Published: (2024)
DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
by: Ishaq, Ayesha, et al.
Published: (2025)
by: Ishaq, Ayesha, et al.
Published: (2025)
AgriChain Visually Grounded Expert Verified Reasoning for Interpretable Agricultural Vision Language Models
by: Mahmood, Hazza, et al.
Published: (2026)
by: Mahmood, Hazza, et al.
Published: (2026)
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
Clustering Aided Weakly Supervised Training to Detect Anomalous Events in Surveillance Videos
by: Zaheer, Muhammad Zaigham, et al.
Published: (2022)
by: Zaheer, Muhammad Zaigham, et al.
Published: (2022)
Collaborative Learning of Anomalies with Privacy (CLAP) for Unsupervised Video Anomaly Detection: A New Baseline
by: Al-lahham, Anas, et al.
Published: (2024)
by: Al-lahham, Anas, et al.
Published: (2024)
Exploiting Autoencoder's Weakness to Generate Pseudo Anomalies
by: Astrid, Marcella, et al.
Published: (2024)
by: Astrid, Marcella, et al.
Published: (2024)
Stabilizing Adversarially Learned One-Class Novelty Detection Using Pseudo Anomalies
by: Zaheer, Muhammad Zaigham, et al.
Published: (2022)
by: Zaheer, Muhammad Zaigham, et al.
Published: (2022)
Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models
by: Rahman, Muhammad Atta ur, et al.
Published: (2025)
by: Rahman, Muhammad Atta ur, et al.
Published: (2025)
Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual Segmentation
by: Zhou, Jinxing, et al.
Published: (2026)
by: Zhou, Jinxing, et al.
Published: (2026)
DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition
by: Liu, Haijing, et al.
Published: (2025)
by: Liu, Haijing, et al.
Published: (2025)
Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration
by: Dudhane, Akshay, et al.
Published: (2024)
by: Dudhane, Akshay, et al.
Published: (2024)
FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation
by: Li, Bingyu, et al.
Published: (2025)
by: Li, Bingyu, et al.
Published: (2025)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
by: Kumar, Komal, et al.
Published: (2025)
by: Kumar, Komal, et al.
Published: (2025)
Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation
by: Boudjoghra, Mohamed El Amine, et al.
Published: (2024)
by: Boudjoghra, Mohamed El Amine, et al.
Published: (2024)
Open-Vocabulary Scene Text Recognition via Pseudo-Image Labeling and Margin Loss
by: Ren, Xuhua, et al.
Published: (2024)
by: Ren, Xuhua, et al.
Published: (2024)
Similar Items
-
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
by: Demidov, Dmitry, et al.
Published: (2025) -
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
by: Thawakar, Omkar, et al.
Published: (2025) -
CoVR-R:Reason-Aware Composed Video Retrieval
by: Thawakar, Omkar, et al.
Published: (2026) -
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
by: Han, Zongyan, et al.
Published: (2025) -
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts
by: Ghaboura, Sara, et al.
Published: (2025)