RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lawrence, Logan, Chasmai, Mustafa, Daroya, Rangel, Liu, Wuao, Jeong, Seoyun, Sun, Aaron, Hamilton, Max, Delattre, Fabien, Saha, Oindrila, Maji, Subhransu, Van Horn, Grant
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908918842654720
author Lawrence, Logan
Chasmai, Mustafa
Daroya, Rangel
Liu, Wuao
Jeong, Seoyun
Sun, Aaron
Hamilton, Max
Delattre, Fabien
Saha, Oindrila
Maji, Subhransu
Van Horn, Grant
author_facet Lawrence, Logan
Chasmai, Mustafa
Daroya, Rangel
Liu, Wuao
Jeong, Seoyun
Sun, Aaron
Hamilton, Max
Delattre, Fabien
Saha, Oindrila
Maji, Subhransu
Van Horn, Grant
contents Fine-grained bird species identification in the wild is frequently unanswerable from a single image: key cues may be non-visual (e.g. vocalization), or obscured due to occlusion, camera angle, or low resolution. Yet today's multimodal systems are typically judged on answerable, in-schema cases, encouraging confident guesses rather than principled abstention. We propose the RealBirdID benchmark: given an image of a bird, a system should either answer with a species or abstain with a concrete, evidence-based rationale: "requires vocalization," "low quality image," or "view obstructed". For each genus, the dataset includes a validation split composed of curated unanswerable examples with labeled rationales, paired with a companion set of clearly answerable instances. We find that (1) the species identification on the answerable set is challenging for a variety of open-source and proprietary models (less than 13% accuracy for MLLMs including GPT-5 and Gemini-2.5 Pro), (2) models with greater classification ability are not necessarily more calibrated to abstain from unanswerable examples, and (3) that MLLMs generally fail at providing correct reasons even when they do abstain. RealBirdID establishes a focused target for abstention-aware fine-grained recognition and a recipe for measuring progress.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27033
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
Lawrence, Logan
Chasmai, Mustafa
Daroya, Rangel
Liu, Wuao
Jeong, Seoyun
Sun, Aaron
Hamilton, Max
Delattre, Fabien
Saha, Oindrila
Maji, Subhransu
Van Horn, Grant
Computer Vision and Pattern Recognition
Fine-grained bird species identification in the wild is frequently unanswerable from a single image: key cues may be non-visual (e.g. vocalization), or obscured due to occlusion, camera angle, or low resolution. Yet today's multimodal systems are typically judged on answerable, in-schema cases, encouraging confident guesses rather than principled abstention. We propose the RealBirdID benchmark: given an image of a bird, a system should either answer with a species or abstain with a concrete, evidence-based rationale: "requires vocalization," "low quality image," or "view obstructed". For each genus, the dataset includes a validation split composed of curated unanswerable examples with labeled rationales, paired with a companion set of clearly answerable instances. We find that (1) the species identification on the answerable set is challenging for a variety of open-source and proprietary models (less than 13% accuracy for MLLMs including GPT-5 and Gemini-2.5 Pro), (2) models with greater classification ability are not necessarily more calibrated to abstain from unanswerable examples, and (3) that MLLMs generally fail at providing correct reasons even when they do abstain. RealBirdID establishes a focused target for abstention-aware fine-grained recognition and a recipe for measuring progress.
title RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.27033