Quantification of Biodiversity from Historical Survey Text with LLM-based Best-Worst Scaling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haider, Thomas, Perschl, Tobias, Rehbein, Malte
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917914997686272
author Haider, Thomas
Perschl, Tobias
Rehbein, Malte
author_facet Haider, Thomas
Perschl, Tobias
Rehbein, Malte
contents In this study, we evaluate methods to determine the frequency of species via quantity estimation from historical survey text. To that end, we formulate classification tasks and finally show that this problem can be adequately framed as a regression task using Best-Worst Scaling (BWS) with Large Language Models (LLMs). We test Ministral-8B, DeepSeek-V3, and GPT-4, finding that the latter two have reasonable agreement with humans and each other. We conclude that this approach is more cost-effective and similarly robust compared to a fine-grained multi-class approach, allowing automated quantity estimation across species.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04022
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quantification of Biodiversity from Historical Survey Text with LLM-based Best-Worst Scaling
Haider, Thomas
Perschl, Tobias
Rehbein, Malte
Computation and Language
In this study, we evaluate methods to determine the frequency of species via quantity estimation from historical survey text. To that end, we formulate classification tasks and finally show that this problem can be adequately framed as a regression task using Best-Worst Scaling (BWS) with Large Language Models (LLMs). We test Ministral-8B, DeepSeek-V3, and GPT-4, finding that the latter two have reasonable agreement with humans and each other. We conclude that this approach is more cost-effective and similarly robust compared to a fine-grained multi-class approach, allowing automated quantity estimation across species.
title Quantification of Biodiversity from Historical Survey Text with LLM-based Best-Worst Scaling
topic Computation and Language
url https://arxiv.org/abs/2502.04022