Phare: A Safety Probe for Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912394379264000 |
|---|---|
| author | Jeune, Pierre Le Malézieux, Benoît Xiao, Weixuan Dora, Matteo |
| author_facet | Jeune, Pierre Le Malézieux, Benoît Xiao, Weixuan Dora, Matteo |
| contents | Ensuring the safety of large language models (LLMs) is critical for responsible deployment, yet existing evaluations often prioritize performance over identifying failure modes. We introduce Phare, a multilingual diagnostic framework to probe and evaluate LLM behavior across three critical dimensions: hallucination and reliability, social biases, and harmful content generation. Our evaluation of 17 state-of-the-art LLMs reveals patterns of systematic vulnerabilities across all safety dimensions, including sycophancy, prompt sensitivity, and stereotype reproduction. By highlighting these specific failure modes rather than simply ranking models, Phare provides researchers and practitioners with actionable insights to build more robust, aligned, and trustworthy language systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_11365 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Phare: A Safety Probe for Large Language Models Jeune, Pierre Le Malézieux, Benoît Xiao, Weixuan Dora, Matteo Computers and Society Artificial Intelligence Computation and Language Cryptography and Security Ensuring the safety of large language models (LLMs) is critical for responsible deployment, yet existing evaluations often prioritize performance over identifying failure modes. We introduce Phare, a multilingual diagnostic framework to probe and evaluate LLM behavior across three critical dimensions: hallucination and reliability, social biases, and harmful content generation. Our evaluation of 17 state-of-the-art LLMs reveals patterns of systematic vulnerabilities across all safety dimensions, including sycophancy, prompt sensitivity, and stereotype reproduction. By highlighting these specific failure modes rather than simply ranking models, Phare provides researchers and practitioners with actionable insights to build more robust, aligned, and trustworthy language systems. |
| title | Phare: A Safety Probe for Large Language Models |
| topic | Computers and Society Artificial Intelligence Computation and Language Cryptography and Security |
| url | https://arxiv.org/abs/2505.11365 |