You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
Fuente:
arXiv
Saved in:
| Main Authors: | Lawrence, Logan, Saha, Oindrila, Wei, Megan, Sun, Chen, Maji, Subhransu, Van Horn, Grant |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generate, Transduct, Adapt: Iterative Transduction with VLMs
by: Saha, Oindrila, et al.
Published: (2025)
by: Saha, Oindrila, et al.
Published: (2025)
Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions
by: Saha, Oindrila, et al.
Published: (2024)
by: Saha, Oindrila, et al.
Published: (2024)
Not All Birds Look The Same: Identity-Preserving Generation For Birds
by: Sun, Aaron, et al.
Published: (2025)
by: Sun, Aaron, et al.
Published: (2025)
Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study
by: Liu, Wuao, et al.
Published: (2026)
by: Liu, Wuao, et al.
Published: (2026)
Merlin L48 Spectrogram Dataset
by: Sun, Aaron, et al.
Published: (2025)
by: Sun, Aaron, et al.
Published: (2025)
Human-in-the-Loop Visual Re-ID for Population Size Estimation
by: Perez, Gustavo, et al.
Published: (2023)
by: Perez, Gustavo, et al.
Published: (2023)
RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
by: Lawrence, Logan, et al.
Published: (2026)
by: Lawrence, Logan, et al.
Published: (2026)
3D Space as a Scratchpad for Editable Text-to-Image Generation
by: Saha, Oindrila, et al.
Published: (2026)
by: Saha, Oindrila, et al.
Published: (2026)
SIGMA-GEN: Structure and Identity Guided Multi-subject Assembly for Image Generation
by: Saha, Oindrila, et al.
Published: (2025)
by: Saha, Oindrila, et al.
Published: (2025)
Consensus-Driven Active Model Selection
by: Kay, Justin, et al.
Published: (2025)
by: Kay, Justin, et al.
Published: (2025)
Moment Sampling in Video LLMs for Long-Form Video QA
by: Chasmai, Mustafa, et al.
Published: (2025)
by: Chasmai, Mustafa, et al.
Published: (2025)
WildSAT: Learning Satellite Image Representations from Wildlife Observations
by: Daroya, Rangel, et al.
Published: (2024)
by: Daroya, Rangel, et al.
Published: (2024)
YouDream: Generating Anatomically Controllable Consistent Text-to-3D Animals
by: Mishra, Sandeep, et al.
Published: (2024)
by: Mishra, Sandeep, et al.
Published: (2024)
Audio Geolocation: A Natural Sounds Benchmark
by: Chasmai, Mustafa, et al.
Published: (2025)
by: Chasmai, Mustafa, et al.
Published: (2025)
The iNaturalist Sounds Dataset
by: Chasmai, Mustafa, et al.
Published: (2025)
by: Chasmai, Mustafa, et al.
Published: (2025)
Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships
by: Daroya, Rangel, et al.
Published: (2024)
by: Daroya, Rangel, et al.
Published: (2024)
Active Measurement of Two-Point Correlations
by: Hamilton, Max, et al.
Published: (2026)
by: Hamilton, Max, et al.
Published: (2026)
Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
by: Tam, Zhi Rui, et al.
Published: (2024)
by: Tam, Zhi Rui, et al.
Published: (2024)
Feedforward Few-shot Species Range Estimation
by: Lange, Christian, et al.
Published: (2025)
by: Lange, Christian, et al.
Published: (2025)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
TECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction
by: Nguyen, Quynh-Mai Thi, et al.
Published: (2024)
by: Nguyen, Quynh-Mai Thi, et al.
Published: (2024)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
by: Kim, Jeonghwan, et al.
Published: (2024)
by: Kim, Jeonghwan, et al.
Published: (2024)
SuperRivolution: Fine-Scale Rivers from Coarse Temporal Satellite Imagery
by: Daroya, Rangel, et al.
Published: (2025)
by: Daroya, Rangel, et al.
Published: (2025)
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
by: Zou, Xin, et al.
Published: (2024)
by: Zou, Xin, et al.
Published: (2024)
VIDMP3: Video Editing by Representing Motion with Pose and Position Priors
by: Mishra, Sandeep, et al.
Published: (2025)
by: Mishra, Sandeep, et al.
Published: (2025)
C3DAG: Controlled 3D Animal Generation using 3D pose guidance
by: Mishra, Sandeep, et al.
Published: (2024)
by: Mishra, Sandeep, et al.
Published: (2024)
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
by: He, Hulingxiao, et al.
Published: (2025)
by: He, Hulingxiao, et al.
Published: (2025)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
by: Lin, Junyan, et al.
Published: (2026)
by: Lin, Junyan, et al.
Published: (2026)
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
by: Just, Hoang Anh, et al.
Published: (2025)
by: Just, Hoang Anh, et al.
Published: (2025)
Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
by: Pang, Cong, et al.
Published: (2025)
by: Pang, Cong, et al.
Published: (2025)
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
by: Sun, Haoyuan, et al.
Published: (2025)
by: Sun, Haoyuan, et al.
Published: (2025)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
by: Nachshoni, Eviatar, et al.
Published: (2025)
by: Nachshoni, Eviatar, et al.
Published: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
by: Zhao, Zhixian, et al.
Published: (2026)
by: Zhao, Zhixian, et al.
Published: (2026)
Active Measurement: Efficient Estimation at Scale
by: Hamilton, Max, et al.
Published: (2025)
by: Hamilton, Max, et al.
Published: (2025)
Watch Before You Answer: Learning from Visually Grounded Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
Directional subset simulation method for reliability analysis
by: Kanjilal, Oindrila, et al.
Published: (2026)
by: Kanjilal, Oindrila, et al.
Published: (2026)
HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
by: Dönmez, Esra, et al.
Published: (2026)
by: Dönmez, Esra, et al.
Published: (2026)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Similar Items
-
Generate, Transduct, Adapt: Iterative Transduction with VLMs
by: Saha, Oindrila, et al.
Published: (2025) -
Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions
by: Saha, Oindrila, et al.
Published: (2024) -
Not All Birds Look The Same: Identity-Preserving Generation For Birds
by: Sun, Aaron, et al.
Published: (2025) -
Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study
by: Liu, Wuao, et al.
Published: (2026) -
Merlin L48 Spectrogram Dataset
by: Sun, Aaron, et al.
Published: (2025)