Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Banerjee, Somnath, Jha, Pranav, Hazra, Rima, Mukherjee, Animesh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languages
by: Banerjee, Somnath, et al.
Published: (2026)
by: Banerjee, Somnath, et al.
Published: (2026)
How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
Evaluating the Ebb and Flow: An In-depth Analysis of Question-Answering Trends across Diverse Platforms
by: Hazra, Rima, et al.
Published: (2023)
by: Hazra, Rima, et al.
Published: (2023)
DistALANER: Distantly Supervised Active Learning Augmented Named Entity Recognition in the Open Source Software Ecosystem
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
by: Banerjee, Somnath, et al.
Published: (2025)
by: Banerjee, Somnath, et al.
Published: (2025)
Breaking Boundaries: Investigating the Effects of Model Editing on Cross-linguistic Performance
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models
by: Adak, Sayantan, et al.
Published: (2025)
by: Adak, Sayantan, et al.
Published: (2025)
ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
by: Banerjee, Somnath, et al.
Published: (2025)
by: Banerjee, Somnath, et al.
Published: (2025)
Context Matters: Pushing the Boundaries of Open-Ended Answer Generation with Graph-Structured Knowledge Context
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
SafeMath: Inference-time Safety improves Math Accuracy
by: Basu, Sagnik, et al.
Published: (2026)
by: Basu, Sagnik, et al.
Published: (2026)
MemeSense: An Adaptive In-Context Framework for Social Commonsense Driven Meme Moderation
by: Adak, Sayantan, et al.
Published: (2025)
by: Adak, Sayantan, et al.
Published: (2025)
Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
by: Banerjee, Somnath, et al.
Published: (2025)
by: Banerjee, Somnath, et al.
Published: (2025)
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models
by: Hazra, Rima, et al.
Published: (2024)
by: Hazra, Rima, et al.
Published: (2024)
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
by: Hazra, Rima, et al.
Published: (2024)
by: Hazra, Rima, et al.
Published: (2024)
Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
From Fluent to Verifiable: Claim-Level Auditability for Deep Research Agents
by: Rasheed, Razeen A, et al.
Published: (2026)
by: Rasheed, Razeen A, et al.
Published: (2026)
InfFeed: Influence Functions as a Feedback to Improve the Performance of Subjective Tasks
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
by: Agarwal, Chirag, et al.
Published: (2024)
by: Agarwal, Chirag, et al.
Published: (2024)
Duplicate Question Retrieval and Confirmation Time Prediction in Software Communities
by: Hazra, Rima, et al.
Published: (2023)
by: Hazra, Rima, et al.
Published: (2023)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
by: Wojciechowski, Adam, et al.
Published: (2024)
by: Wojciechowski, Adam, et al.
Published: (2024)
CLIX: Cross-Lingual Explanations of Idiomatic Expressions
by: Gluck, Aaron, et al.
Published: (2025)
by: Gluck, Aaron, et al.
Published: (2025)
Lost without translation -- Can transformer (language models) understand mood states?
by: Shivaprakash, Prakrithi, et al.
Published: (2025)
by: Shivaprakash, Prakrithi, et al.
Published: (2025)
Exploring the Trade-off Between Model Performance and Explanation Plausibility of Text Classifiers Using Human Rationales
by: Resck, Lucas E., et al.
Published: (2024)
by: Resck, Lucas E., et al.
Published: (2024)
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
by: Hazra, Rima, et al.
Published: (2026)
by: Hazra, Rima, et al.
Published: (2026)
Faithful Summarization of Consumer Health Queries: A Cross-Lingual Framework with LLMs
by: Abrar, Ajwad, et al.
Published: (2025)
by: Abrar, Ajwad, et al.
Published: (2025)
Tutoring Large Language Models to be Domain-adaptive, Precise, and Safe
by: Banerjee, Somnath
Published: (2026)
by: Banerjee, Somnath
Published: (2026)
Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms
by: Shukla, Vaibhav, et al.
Published: (2026)
by: Shukla, Vaibhav, et al.
Published: (2026)
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions
by: Sachdeva, Rachneet, et al.
Published: (2025)
by: Sachdeva, Rachneet, et al.
Published: (2025)
Towards Cross-Lingual Explanation of Artwork in Large-scale Vision Language Models
by: Ozaki, Shintaro, et al.
Published: (2024)
by: Ozaki, Shintaro, et al.
Published: (2024)
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
by: Salhan, Suchir, et al.
Published: (2024)
by: Salhan, Suchir, et al.
Published: (2024)
RA-MTR: A Retrieval Augmented Multi-Task Reader based Approach for Inspirational Quote Extraction from Long Documents
by: Adak, Sayantan, et al.
Published: (2025)
by: Adak, Sayantan, et al.
Published: (2025)
Regularization, Semi-supervision, and Supervision for a Plausible Attention-Based Explanation
by: Nguyen, Duc Hau, et al.
Published: (2025)
by: Nguyen, Duc Hau, et al.
Published: (2025)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
by: Bhan, Milan, et al.
Published: (2025)
by: Bhan, Milan, et al.
Published: (2025)
Towards Faithful Explanations for Text Classification with Robustness Improvement and Explanation Guided Training
by: Li, Dongfang, et al.
Published: (2023)
by: Li, Dongfang, et al.
Published: (2023)
Benchmarking Motivational Interviewing Competence of Large Language Models
by: Jha, Aishwariya, et al.
Published: (2026)
by: Jha, Aishwariya, et al.
Published: (2026)
Self-Critique and Refinement for Faithful Natural Language Explanations
by: Wang, Yingming, et al.
Published: (2025)
by: Wang, Yingming, et al.
Published: (2025)
Towards Faithful Model Explanation in NLP: A Survey
by: Lyu, Qing, et al.
Published: (2022)
by: Lyu, Qing, et al.
Published: (2022)
Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Prediction
by: Yuan, Mengying, et al.
Published: (2025)
by: Yuan, Mengying, et al.
Published: (2025)
On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations
by: Mehrpanah, Amir, et al.
Published: (2025)
by: Mehrpanah, Amir, et al.
Published: (2025)
Similar Items
-
Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languages
by: Banerjee, Somnath, et al.
Published: (2026) -
How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries
by: Banerjee, Somnath, et al.
Published: (2024) -
Evaluating the Ebb and Flow: An In-depth Analysis of Question-Answering Trends across Diverse Platforms
by: Hazra, Rima, et al.
Published: (2023) -
DistALANER: Distantly Supervised Active Learning Augmented Named Entity Recognition in the Open Source Software Ecosystem
by: Banerjee, Somnath, et al.
Published: (2024) -
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
by: Banerjee, Somnath, et al.
Published: (2025)