Anthropocentric bias in language model evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Millière, Raphaël, Rathkopf, Charles |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Philosophy of Cognitive Science in the Age of Deep Learning
by: Millière, Raphaël
Published: (2024)
by: Millière, Raphaël
Published: (2024)
Language Models as Models of Language
by: Millière, Raphaël
Published: (2024)
by: Millière, Raphaël
Published: (2024)
Normative Conflicts and Shallow AI Alignment
by: Millière, Raphaël
Published: (2025)
by: Millière, Raphaël
Published: (2025)
A Philosophical Introduction to Language Models - Part II: The Way Forward
by: Millière, Raphaël, et al.
Published: (2024)
by: Millière, Raphaël, et al.
Published: (2024)
The Vector Grounding Problem
by: Mollo, Dimitri Coelho, et al.
Published: (2023)
by: Mollo, Dimitri Coelho, et al.
Published: (2023)
A Philosophical Introduction to Language Models -- Part I: Continuity With Classic Debates
by: Millière, Raphaël, et al.
Published: (2024)
by: Millière, Raphaël, et al.
Published: (2024)
How Do Transformers Learn Variable Binding in Symbolic Programs?
by: Wu, Yiwei, et al.
Published: (2025)
by: Wu, Yiwei, et al.
Published: (2025)
LLMs as Models for Analogical Reasoning
by: Musker, Sam, et al.
Published: (2024)
by: Musker, Sam, et al.
Published: (2024)
Hallucination, reliability, and the role of generative AI in science
by: Rathkopf, Charles
Published: (2025)
by: Rathkopf, Charles
Published: (2025)
Decoding In-Context Learning: Neuroscience-inspired Analysis of Representations in Large Language Models
by: Yousefi, Safoora, et al.
Published: (2023)
by: Yousefi, Safoora, et al.
Published: (2023)
Addressing cognitive bias in medical language models
by: Schmidgall, Samuel, et al.
Published: (2024)
by: Schmidgall, Samuel, et al.
Published: (2024)
Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models
by: Guey, William, et al.
Published: (2026)
by: Guey, William, et al.
Published: (2026)
HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America
by: Ivetta, Guido, et al.
Published: (2025)
by: Ivetta, Guido, et al.
Published: (2025)
Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models
by: Salnikov, Mikhail, et al.
Published: (2025)
by: Salnikov, Mikhail, et al.
Published: (2025)
Re-evaluating Theory of Mind evaluation in large language models
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Factual consistency evaluation of summarization in the Era of large language models
by: Luo, Zheheng, et al.
Published: (2024)
by: Luo, Zheheng, et al.
Published: (2024)
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023)
by: Coda-Forno, Julian, et al.
Published: (2023)
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
by: Jelodar, Hamed, et al.
Published: (2025)
by: Jelodar, Hamed, et al.
Published: (2025)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
by: Jassim, Serwan, et al.
Published: (2023)
by: Jassim, Serwan, et al.
Published: (2023)
A closer look at how large language models trust humans: patterns and biases
by: Lerman, Valeria, et al.
Published: (2025)
by: Lerman, Valeria, et al.
Published: (2025)
Philosophy of cognitive science in the age of deep learning
by: Raphaël Millière
Published: (2024)
by: Raphaël Millière
Published: (2024)
Generating bilingual example sentences with large language models as lexicography assistants
by: Merx, Raphael, et al.
Published: (2024)
by: Merx, Raphael, et al.
Published: (2024)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
by: Wen, Zichen, et al.
Published: (2024)
by: Wen, Zichen, et al.
Published: (2024)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
by: Li, Kunning, et al.
Published: (2025)
by: Li, Kunning, et al.
Published: (2025)
MindScope: Exploring cognitive biases in large language models through Multi-Agent Systems
by: Xie, Zhentao, et al.
Published: (2024)
by: Xie, Zhentao, et al.
Published: (2024)
Morphological evaluation of subwords vocabulary used by BETO language model
by: García-Sierra, Óscar, et al.
Published: (2024)
by: García-Sierra, Óscar, et al.
Published: (2024)
BIPOLAR: Polarization-based granular framework for LLM bias evaluation
by: Pavlíček, Martin, et al.
Published: (2025)
by: Pavlíček, Martin, et al.
Published: (2025)
How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language
by: Das, Dipto, et al.
Published: (2025)
by: Das, Dipto, et al.
Published: (2025)
Source framing triggers systematic evaluation bias in Large Language Models
by: Germani, Federico, et al.
Published: (2025)
by: Germani, Federico, et al.
Published: (2025)
Leveraging language models for summarizing mental state examinations: A comprehensive evaluation and dataset release
by: Sahu, Nilesh Kumar, et al.
Published: (2024)
by: Sahu, Nilesh Kumar, et al.
Published: (2024)
KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations
by: Harbola, Chitranshu, et al.
Published: (2025)
by: Harbola, Chitranshu, et al.
Published: (2025)
LLMs left, right, and center: Assessing GPT's capabilities to label political bias from web domains
by: Hernandes, Raphael, et al.
Published: (2024)
by: Hernandes, Raphael, et al.
Published: (2024)
A database to support the evaluation of gender biases in GPT-4o output
by: Mehner, Luise, et al.
Published: (2025)
by: Mehner, Luise, et al.
Published: (2025)
Into the crossfire: evaluating the use of a language model to crowdsource gun violence reports
by: Belisario, Adriano, et al.
Published: (2024)
by: Belisario, Adriano, et al.
Published: (2024)
ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
by: Ahrens, Lara, et al.
Published: (2025)
by: Ahrens, Lara, et al.
Published: (2025)
Development and bilingual evaluation of Japanese medical large language model within reasonably low computational resources
by: Sukeda, Issey
Published: (2024)
by: Sukeda, Issey
Published: (2024)
'Neural howlround' in large language models: a self-reinforcing bias phenomenon, and a dynamic attenuation solution
by: Drake, Seth
Published: (2025)
by: Drake, Seth
Published: (2025)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
by: Matos, João, et al.
Published: (2024)
by: Matos, João, et al.
Published: (2024)
Similar Items
-
Philosophy of Cognitive Science in the Age of Deep Learning
by: Millière, Raphaël
Published: (2024) -
Language Models as Models of Language
by: Millière, Raphaël
Published: (2024) -
Normative Conflicts and Shallow AI Alignment
by: Millière, Raphaël
Published: (2025) -
A Philosophical Introduction to Language Models - Part II: The Way Forward
by: Millière, Raphaël, et al.
Published: (2024) -
The Vector Grounding Problem
by: Mollo, Dimitri Coelho, et al.
Published: (2023)