DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Jiawei, Lei, Ming, Yang, Yaning, Lin, Xinyan, Le, Yuquan, Ma, Qiwei, Xu, Zhiwei, Lv, Zheqi, Ang, Yuchen, Quan, Zhe, Chua, Tat-Seng
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915960641814528
author Wang, Jiawei
Lei, Ming
Yang, Yaning
Lin, Xinyan
Le, Yuquan
Ma, Qiwei
Xu, Zhiwei
Lv, Zheqi
Ang, Yuchen
Quan, Zhe
Chua, Tat-Seng
author_facet Wang, Jiawei
Lei, Ming
Yang, Yaning
Lin, Xinyan
Le, Yuquan
Ma, Qiwei
Xu, Zhiwei
Lv, Zheqi
Ang, Yuchen
Quan, Zhe
Chua, Tat-Seng
contents Identifying species in biology among tens of thousands of visually similar taxa while discovering unknown species in open-world environments remains a fundamental challenge in biodiversity research. Current methods treat identification and discovery as separate problems, with classification models assuming closed sets and discovery relying on threshold-based rejection. Here we present DeepTaxon, a retrieval-augmented multimodal framework that unifies species identification and discovery through interpretable reasoning over retrieved visual evidence. Given a query image, DeepTaxon retrieves the top-$k$ candidate species with $n$ exemplar images each from a retrieval index and performs chain-of-thought comparative reasoning. Critically, we redefine discovery as an explicit, retrieval-based decision problem rather than an implicit parametric memory problem. A sample is novel if and only if the retrieval index lacks sufficient evidence for identification, so each retrieval naturally yields a classification or discovery label without manual annotation, thereby providing automatic supervision for both tasks. We train the framework via supervised fine-tuning on synthetic retrieval-augmented data, followed by reinforcement learning on hard samples, converting high-recall retrieval into high-precision decisions that scale to massive taxonomic vocabularies. Extensive experiments on a large-scale in-distribution benchmark and six out-of-distribution datasets demonstrate consistent improvements in both identification and discovery. Ablation studies further reveal effective test-time scaling with candidate count $k$ and exemplar count $n$, strong zero-shot transfer to unseen domains, and consistent performance across retrieval encoders, establishing an interpretable solution for biodiversity research.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24029
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery
Wang, Jiawei
Lei, Ming
Yang, Yaning
Lin, Xinyan
Le, Yuquan
Ma, Qiwei
Xu, Zhiwei
Lv, Zheqi
Ang, Yuchen
Quan, Zhe
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Computation and Language
Information Retrieval
Multimedia
I.4.8; H.3.3; I.2.6
Identifying species in biology among tens of thousands of visually similar taxa while discovering unknown species in open-world environments remains a fundamental challenge in biodiversity research. Current methods treat identification and discovery as separate problems, with classification models assuming closed sets and discovery relying on threshold-based rejection. Here we present DeepTaxon, a retrieval-augmented multimodal framework that unifies species identification and discovery through interpretable reasoning over retrieved visual evidence. Given a query image, DeepTaxon retrieves the top-$k$ candidate species with $n$ exemplar images each from a retrieval index and performs chain-of-thought comparative reasoning. Critically, we redefine discovery as an explicit, retrieval-based decision problem rather than an implicit parametric memory problem. A sample is novel if and only if the retrieval index lacks sufficient evidence for identification, so each retrieval naturally yields a classification or discovery label without manual annotation, thereby providing automatic supervision for both tasks. We train the framework via supervised fine-tuning on synthetic retrieval-augmented data, followed by reinforcement learning on hard samples, converting high-recall retrieval into high-precision decisions that scale to massive taxonomic vocabularies. Extensive experiments on a large-scale in-distribution benchmark and six out-of-distribution datasets demonstrate consistent improvements in both identification and discovery. Ablation studies further reveal effective test-time scaling with candidate count $k$ and exemplar count $n$, strong zero-shot transfer to unseen domains, and consistent performance across retrieval encoders, establishing an interpretable solution for biodiversity research.
title DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery
topic Computer Vision and Pattern Recognition
Computation and Language
Information Retrieval
Multimedia
I.4.8; H.3.3; I.2.6
url https://arxiv.org/abs/2604.24029