BAGEL: Benchmarking Animal Knowledge Expertise in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911602093064192 |
|---|---|
| author | Shen, Jiacheng Hagiwara, Masato Alizadeh, Milad Gilsenan-McMahon, Ellen Miron, Marius Robinson, David Chemla, Emmanuel Keen, Sara Narula, Gagan Laurière, Mathieu Geist, Matthieu Pietquin, Olivier |
| author_facet | Shen, Jiacheng Hagiwara, Masato Alizadeh, Milad Gilsenan-McMahon, Ellen Miron, Marius Robinson, David Chemla, Emmanuel Keen, Sara Narula, Gagan Laurière, Mathieu Geist, Matthieu Pietquin, Olivier |
| contents | Large language models have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language models handle specialized animal-related knowledge under a unified closed-book evaluation protocol. We introduce BAGEL, a benchmark for evaluating animal knowledge expertise in language models. BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, using a combination of curated examples and automatically generated closed-book question-answer pairs. The benchmark covers multiple aspects of animal knowledge, including taxonomy, morphology, habitat, behavior, vocalization, geographic distribution, and species interactions. By focusing on closed-book evaluation, BAGEL measures animal-related knowledge of models without external retrieval at inference time. BAGEL further supports fine-grained analysis across source domains, taxonomic groups, and knowledge categories, enabling a more precise characterization of model strengths and systematic failure modes. Our benchmark provides a new testbed for studying domain-specific knowledge generalization in language models and for improving their reliability in biodiversity-related applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_16241 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | BAGEL: Benchmarking Animal Knowledge Expertise in Language Models Shen, Jiacheng Hagiwara, Masato Alizadeh, Milad Gilsenan-McMahon, Ellen Miron, Marius Robinson, David Chemla, Emmanuel Keen, Sara Narula, Gagan Laurière, Mathieu Geist, Matthieu Pietquin, Olivier Computation and Language Artificial Intelligence Large language models have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language models handle specialized animal-related knowledge under a unified closed-book evaluation protocol. We introduce BAGEL, a benchmark for evaluating animal knowledge expertise in language models. BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, using a combination of curated examples and automatically generated closed-book question-answer pairs. The benchmark covers multiple aspects of animal knowledge, including taxonomy, morphology, habitat, behavior, vocalization, geographic distribution, and species interactions. By focusing on closed-book evaluation, BAGEL measures animal-related knowledge of models without external retrieval at inference time. BAGEL further supports fine-grained analysis across source domains, taxonomic groups, and knowledge categories, enabling a more precise characterization of model strengths and systematic failure modes. Our benchmark provides a new testbed for studying domain-specific knowledge generalization in language models and for improving their reliability in biodiversity-related applications. |
| title | BAGEL: Benchmarking Animal Knowledge Expertise in Language Models |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2604.16241 |