The SMeL Test: A simple benchmark for media literacy in language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahdritz, Gustaf, Kleiman, Anat
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915433102180352
author Ahdritz, Gustaf
Kleiman, Anat
author_facet Ahdritz, Gustaf
Kleiman, Anat
contents The internet is rife with unattributed, deliberately misleading, or otherwise untrustworthy content. Though large language models (LLMs) are often tasked with autonomous web browsing, the extent to which they have learned the simple heuristics human researchers use to navigate this noisy environment is not currently known. In this paper, we introduce the Synthetic Media Literacy Test (SMeL Test), a minimal benchmark that tests the ability of language models to actively filter out untrustworthy information in context. We benchmark a variety of commonly used instruction-tuned LLMs, including reasoning models, and find that no model consistently succeeds; while reasoning in particular is associated with higher scores, even the best API model we test hallucinates up to 70% of the time. Remarkably, larger and more capable models do not necessarily outperform their smaller counterparts. We hope our work sheds more light on this important form of hallucination and guides the development of new methods to combat it.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02074
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The SMeL Test: A simple benchmark for media literacy in language models
Ahdritz, Gustaf
Kleiman, Anat
Computation and Language
Machine Learning
The internet is rife with unattributed, deliberately misleading, or otherwise untrustworthy content. Though large language models (LLMs) are often tasked with autonomous web browsing, the extent to which they have learned the simple heuristics human researchers use to navigate this noisy environment is not currently known. In this paper, we introduce the Synthetic Media Literacy Test (SMeL Test), a minimal benchmark that tests the ability of language models to actively filter out untrustworthy information in context. We benchmark a variety of commonly used instruction-tuned LLMs, including reasoning models, and find that no model consistently succeeds; while reasoning in particular is associated with higher scores, even the best API model we test hallucinates up to 70% of the time. Remarkably, larger and more capable models do not necessarily outperform their smaller counterparts. We hope our work sheds more light on this important form of hallucination and guides the development of new methods to combat it.
title The SMeL Test: A simple benchmark for media literacy in language models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2508.02074