Phare: A Safety Probe for Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jeune, Pierre Le, Malézieux, Benoît, Xiao, Weixuan, Dora, Matteo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912394379264000
author Jeune, Pierre Le
Malézieux, Benoît
Xiao, Weixuan
Dora, Matteo
author_facet Jeune, Pierre Le
Malézieux, Benoît
Xiao, Weixuan
Dora, Matteo
contents Ensuring the safety of large language models (LLMs) is critical for responsible deployment, yet existing evaluations often prioritize performance over identifying failure modes. We introduce Phare, a multilingual diagnostic framework to probe and evaluate LLM behavior across three critical dimensions: hallucination and reliability, social biases, and harmful content generation. Our evaluation of 17 state-of-the-art LLMs reveals patterns of systematic vulnerabilities across all safety dimensions, including sycophancy, prompt sensitivity, and stereotype reproduction. By highlighting these specific failure modes rather than simply ranking models, Phare provides researchers and practitioners with actionable insights to build more robust, aligned, and trustworthy language systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11365
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Phare: A Safety Probe for Large Language Models
Jeune, Pierre Le
Malézieux, Benoît
Xiao, Weixuan
Dora, Matteo
Computers and Society
Artificial Intelligence
Computation and Language
Cryptography and Security
Ensuring the safety of large language models (LLMs) is critical for responsible deployment, yet existing evaluations often prioritize performance over identifying failure modes. We introduce Phare, a multilingual diagnostic framework to probe and evaluate LLM behavior across three critical dimensions: hallucination and reliability, social biases, and harmful content generation. Our evaluation of 17 state-of-the-art LLMs reveals patterns of systematic vulnerabilities across all safety dimensions, including sycophancy, prompt sensitivity, and stereotype reproduction. By highlighting these specific failure modes rather than simply ranking models, Phare provides researchers and practitioners with actionable insights to build more robust, aligned, and trustworthy language systems.
title Phare: A Safety Probe for Large Language Models
topic Computers and Society
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2505.11365