Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Banerjee, Somnath, Jha, Pranav, Hazra, Rima, Mukherjee, Animesh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917510125715456
author Banerjee, Somnath
Jha, Pranav
Hazra, Rima
Mukherjee, Animesh
author_facet Banerjee, Somnath
Jha, Pranav
Hazra, Rima
Mukherjee, Animesh
contents LLMs deployed multilingually are often audited via English explanations for non-English inputs. We evaluate extractive explanations ''where the model identifies input token spans as evidence alongside a generated rationale'' and uncover a systematic trade-off: English-pivot explanations can achieve higher span agreement with human rationales while their evidence becomes less causally grounded in the model's prediction, as measured by both comprehensiveness and sufficiency. Across 3 tasks, 5~languages, and 2~multilingual LLM families, we find that English explanations frequently produce fluent but loosely anchored rationales, with comprehensiveness degrading by up to 5.7x relative to native-language conditions - even as task accuracy remains stable across settings. For socially nuanced classification, English pivots also fail to preserve pragmatic cues, reducing both faithfulness and span agreement. We recommend auditing explanations in the input language, reporting multi-faceted faithfulness metrics beyond lexical overlap, and treating English rationales as communication summaries rather than faithful decision traces.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19274
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations
Banerjee, Somnath
Jha, Pranav
Hazra, Rima
Mukherjee, Animesh
Computation and Language
LLMs deployed multilingually are often audited via English explanations for non-English inputs. We evaluate extractive explanations ''where the model identifies input token spans as evidence alongside a generated rationale'' and uncover a systematic trade-off: English-pivot explanations can achieve higher span agreement with human rationales while their evidence becomes less causally grounded in the model's prediction, as measured by both comprehensiveness and sufficiency. Across 3 tasks, 5~languages, and 2~multilingual LLM families, we find that English explanations frequently produce fluent but loosely anchored rationales, with comprehensiveness degrading by up to 5.7x relative to native-language conditions - even as task accuracy remains stable across settings. For socially nuanced classification, English pivots also fail to preserve pragmatic cues, reducing both faithfulness and span agreement. We recommend auditing explanations in the input language, reporting multi-faceted faithfulness metrics beyond lexical overlap, and treating English rationales as communication summaries rather than faithful decision traces.
title Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations
topic Computation and Language
url https://arxiv.org/abs/2605.19274