BAGEL: Benchmarking Animal Knowledge Expertise in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Jiacheng, Hagiwara, Masato, Alizadeh, Milad, Gilsenan-McMahon, Ellen, Miron, Marius, Robinson, David, Chemla, Emmanuel, Keen, Sara, Narula, Gagan, Laurière, Mathieu, Geist, Matthieu, Pietquin, Olivier
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911602093064192
author Shen, Jiacheng
Hagiwara, Masato
Alizadeh, Milad
Gilsenan-McMahon, Ellen
Miron, Marius
Robinson, David
Chemla, Emmanuel
Keen, Sara
Narula, Gagan
Laurière, Mathieu
Geist, Matthieu
Pietquin, Olivier
author_facet Shen, Jiacheng
Hagiwara, Masato
Alizadeh, Milad
Gilsenan-McMahon, Ellen
Miron, Marius
Robinson, David
Chemla, Emmanuel
Keen, Sara
Narula, Gagan
Laurière, Mathieu
Geist, Matthieu
Pietquin, Olivier
contents Large language models have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language models handle specialized animal-related knowledge under a unified closed-book evaluation protocol. We introduce BAGEL, a benchmark for evaluating animal knowledge expertise in language models. BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, using a combination of curated examples and automatically generated closed-book question-answer pairs. The benchmark covers multiple aspects of animal knowledge, including taxonomy, morphology, habitat, behavior, vocalization, geographic distribution, and species interactions. By focusing on closed-book evaluation, BAGEL measures animal-related knowledge of models without external retrieval at inference time. BAGEL further supports fine-grained analysis across source domains, taxonomic groups, and knowledge categories, enabling a more precise characterization of model strengths and systematic failure modes. Our benchmark provides a new testbed for studying domain-specific knowledge generalization in language models and for improving their reliability in biodiversity-related applications.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16241
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BAGEL: Benchmarking Animal Knowledge Expertise in Language Models
Shen, Jiacheng
Hagiwara, Masato
Alizadeh, Milad
Gilsenan-McMahon, Ellen
Miron, Marius
Robinson, David
Chemla, Emmanuel
Keen, Sara
Narula, Gagan
Laurière, Mathieu
Geist, Matthieu
Pietquin, Olivier
Computation and Language
Artificial Intelligence
Large language models have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language models handle specialized animal-related knowledge under a unified closed-book evaluation protocol. We introduce BAGEL, a benchmark for evaluating animal knowledge expertise in language models. BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, using a combination of curated examples and automatically generated closed-book question-answer pairs. The benchmark covers multiple aspects of animal knowledge, including taxonomy, morphology, habitat, behavior, vocalization, geographic distribution, and species interactions. By focusing on closed-book evaluation, BAGEL measures animal-related knowledge of models without external retrieval at inference time. BAGEL further supports fine-grained analysis across source domains, taxonomic groups, and knowledge categories, enabling a more precise characterization of model strengths and systematic failure modes. Our benchmark provides a new testbed for studying domain-specific knowledge generalization in language models and for improving their reliability in biodiversity-related applications.
title BAGEL: Benchmarking Animal Knowledge Expertise in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.16241