Gespeichert in:
| Hauptverfasser: | Gong, Xilin, Yang, Shu, Cao, Zehua, Billard, Lynne, Wang, Di |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.00300 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
von: Ghandeharioun, Asma, et al.
Veröffentlicht: (2024)
von: Ghandeharioun, Asma, et al.
Veröffentlicht: (2024)
FaithLM: Towards Faithful Explanations for Large Language Models
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
von: Zhang, Jiayi, et al.
Veröffentlicht: (2025)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2025)
Mitigating the Bias of Large Language Model Evaluation
von: Zhou, Hongli, et al.
Veröffentlicht: (2024)
von: Zhou, Hongli, et al.
Veröffentlicht: (2024)
Investigating CoT Monitorability in Large Reasoning Models
von: Yang, Shu, et al.
Veröffentlicht: (2025)
von: Yang, Shu, et al.
Veröffentlicht: (2025)
Mitigating Large Language Model Hallucination with Faithful Finetuning
von: Hu, Minda, et al.
Veröffentlicht: (2024)
von: Hu, Minda, et al.
Veröffentlicht: (2024)
MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
von: Hu, Jingyu, et al.
Veröffentlicht: (2025)
von: Hu, Jingyu, et al.
Veröffentlicht: (2025)
Mitigating Hidden Confounding by Progressive Confounder Imputation via Large Language Models
von: Yang, Hao, et al.
Veröffentlicht: (2025)
von: Yang, Hao, et al.
Veröffentlicht: (2025)
Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
von: Doi, Tomoki, et al.
Veröffentlicht: (2025)
von: Doi, Tomoki, et al.
Veröffentlicht: (2025)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
von: Matton, Katie, et al.
Veröffentlicht: (2025)
von: Matton, Katie, et al.
Veröffentlicht: (2025)
Locating and Mitigating Gender Bias in Large Language Models
von: Cai, Yuchen, et al.
Veröffentlicht: (2024)
von: Cai, Yuchen, et al.
Veröffentlicht: (2024)
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
von: Agarwal, Chirag, et al.
Veröffentlicht: (2024)
von: Agarwal, Chirag, et al.
Veröffentlicht: (2024)
Understanding and Mitigating Tokenization Bias in Language Models
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
von: Zhang, Zhenliang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenliang, et al.
Veröffentlicht: (2025)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
von: Siegel, Noah Y., et al.
Veröffentlicht: (2024)
von: Siegel, Noah Y., et al.
Veröffentlicht: (2024)
Bias in Large Language Models: Origin, Evaluation, and Mitigation
von: Guo, Yufei, et al.
Veröffentlicht: (2024)
von: Guo, Yufei, et al.
Veröffentlicht: (2024)
Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2024)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2024)
Mitigating Label Length Bias in Large Language Models
von: Sanz-Guerrero, Mario, et al.
Veröffentlicht: (2025)
von: Sanz-Guerrero, Mario, et al.
Veröffentlicht: (2025)
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
von: You, Liangliang, et al.
Veröffentlicht: (2025)
von: You, Liangliang, et al.
Veröffentlicht: (2025)
All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG
von: Wang, Dan, et al.
Veröffentlicht: (2026)
von: Wang, Dan, et al.
Veröffentlicht: (2026)
Understanding the Repeat Curse in Large Language Models from a Feature Perspective
von: Yao, Junchi, et al.
Veröffentlicht: (2025)
von: Yao, Junchi, et al.
Veröffentlicht: (2025)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
von: Bhan, Milan, et al.
Veröffentlicht: (2025)
von: Bhan, Milan, et al.
Veröffentlicht: (2025)
Self-Critique and Refinement for Faithful Natural Language Explanations
von: Wang, Yingming, et al.
Veröffentlicht: (2025)
von: Wang, Yingming, et al.
Veröffentlicht: (2025)
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
von: Sovrano, Francesco, et al.
Veröffentlicht: (2026)
von: Sovrano, Francesco, et al.
Veröffentlicht: (2026)
Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
Detection, Classification, and Mitigation of Gender Bias in Large Language Models
von: Cheng, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Cheng, Xiaoqing, et al.
Veröffentlicht: (2025)
Do Multilingual Large Language Models Mitigate Stereotype Bias?
von: Nie, Shangrui, et al.
Veröffentlicht: (2024)
von: Nie, Shangrui, et al.
Veröffentlicht: (2024)
Unveiling and Mitigating Bias in Mental Health Analysis with Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
Mitigating Bias in Queer Representation within Large Language Models: A Collaborative Agent Approach
von: Huang, Tianyi, et al.
Veröffentlicht: (2024)
von: Huang, Tianyi, et al.
Veröffentlicht: (2024)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
von: Alon, Bar, et al.
Veröffentlicht: (2026)
von: Alon, Bar, et al.
Veröffentlicht: (2026)
Towards Faithful Model Explanation in NLP: A Survey
von: Lyu, Qing, et al.
Veröffentlicht: (2022)
von: Lyu, Qing, et al.
Veröffentlicht: (2022)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
von: Huang, Yanwen, et al.
Veröffentlicht: (2025)
von: Huang, Yanwen, et al.
Veröffentlicht: (2025)
Large Language Model Agents Are Not Always Faithful Self-Evolvers
von: Zhao, Weixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2026)
Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models
von: Luo, Linhao, et al.
Veröffentlicht: (2024)
von: Luo, Linhao, et al.
Veröffentlicht: (2024)
Auto-Search and Refinement: An Automated Framework for Gender Bias Mitigation in Large Language Models
von: Xu, Yue, et al.
Veröffentlicht: (2025)
von: Xu, Yue, et al.
Veröffentlicht: (2025)
Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models
von: Tong, Schrasing, et al.
Veröffentlicht: (2024)
von: Tong, Schrasing, et al.
Veröffentlicht: (2024)
MBIAS: Mitigating Bias in Large Language Models While Retaining Context
von: Raza, Shaina, et al.
Veröffentlicht: (2024)
von: Raza, Shaina, et al.
Veröffentlicht: (2024)
From Critique to Clarity: A Pathway to Faithful and Personalized Code Explanations with Large Language Models
von: Xu, Zexing, et al.
Veröffentlicht: (2024)
von: Xu, Zexing, et al.
Veröffentlicht: (2024)
Multi-Persona Thinking for Bias Mitigation in Large Language Models
von: Chen, Yuxing, et al.
Veröffentlicht: (2026)
von: Chen, Yuxing, et al.
Veröffentlicht: (2026)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
von: Oi, Masanari, et al.
Veröffentlicht: (2024)
von: Oi, Masanari, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
von: Ghandeharioun, Asma, et al.
Veröffentlicht: (2024) -
FaithLM: Towards Faithful Explanations for Large Language Models
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024) -
Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
von: Zhang, Jiayi, et al.
Veröffentlicht: (2025) -
Mitigating the Bias of Large Language Model Evaluation
von: Zhou, Hongli, et al.
Veröffentlicht: (2024) -
Investigating CoT Monitorability in Large Reasoning Models
von: Yang, Shu, et al.
Veröffentlicht: (2025)