RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908918842654720 |
|---|---|
| author | Lawrence, Logan Chasmai, Mustafa Daroya, Rangel Liu, Wuao Jeong, Seoyun Sun, Aaron Hamilton, Max Delattre, Fabien Saha, Oindrila Maji, Subhransu Van Horn, Grant |
| author_facet | Lawrence, Logan Chasmai, Mustafa Daroya, Rangel Liu, Wuao Jeong, Seoyun Sun, Aaron Hamilton, Max Delattre, Fabien Saha, Oindrila Maji, Subhransu Van Horn, Grant |
| contents | Fine-grained bird species identification in the wild is frequently unanswerable from a single image: key cues may be non-visual (e.g. vocalization), or obscured due to occlusion, camera angle, or low resolution. Yet today's multimodal systems are typically judged on answerable, in-schema cases, encouraging confident guesses rather than principled abstention. We propose the RealBirdID benchmark: given an image of a bird, a system should either answer with a species or abstain with a concrete, evidence-based rationale: "requires vocalization," "low quality image," or "view obstructed". For each genus, the dataset includes a validation split composed of curated unanswerable examples with labeled rationales, paired with a companion set of clearly answerable instances. We find that (1) the species identification on the answerable set is challenging for a variety of open-source and proprietary models (less than 13% accuracy for MLLMs including GPT-5 and Gemini-2.5 Pro), (2) models with greater classification ability are not necessarily more calibrated to abstain from unanswerable examples, and (3) that MLLMs generally fail at providing correct reasons even when they do abstain. RealBirdID establishes a focused target for abstention-aware fine-grained recognition and a recipe for measuring progress. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_27033 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs Lawrence, Logan Chasmai, Mustafa Daroya, Rangel Liu, Wuao Jeong, Seoyun Sun, Aaron Hamilton, Max Delattre, Fabien Saha, Oindrila Maji, Subhransu Van Horn, Grant Computer Vision and Pattern Recognition Fine-grained bird species identification in the wild is frequently unanswerable from a single image: key cues may be non-visual (e.g. vocalization), or obscured due to occlusion, camera angle, or low resolution. Yet today's multimodal systems are typically judged on answerable, in-schema cases, encouraging confident guesses rather than principled abstention. We propose the RealBirdID benchmark: given an image of a bird, a system should either answer with a species or abstain with a concrete, evidence-based rationale: "requires vocalization," "low quality image," or "view obstructed". For each genus, the dataset includes a validation split composed of curated unanswerable examples with labeled rationales, paired with a companion set of clearly answerable instances. We find that (1) the species identification on the answerable set is challenging for a variety of open-source and proprietary models (less than 13% accuracy for MLLMs including GPT-5 and Gemini-2.5 Pro), (2) models with greater classification ability are not necessarily more calibrated to abstain from unanswerable examples, and (3) that MLLMs generally fail at providing correct reasons even when they do abstain. RealBirdID establishes a focused target for abstention-aware fine-grained recognition and a recipe for measuring progress. |
| title | RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.27033 |