Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Doi, Tomoki, Isonuma, Masaru, Yanaka, Hitomi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Comprehensive Evaluation of Large Language Models for Topic Modeling
por: Doi, Tomoki, et al.
Publicado: (2024)
por: Doi, Tomoki, et al.
Publicado: (2024)
Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?
por: Shinozaki, Taiga, et al.
Publicado: (2025)
por: Shinozaki, Taiga, et al.
Publicado: (2025)
Unlearning Traces the Influential Training Data of Language Models
por: Isonuma, Masaru, et al.
Publicado: (2024)
por: Isonuma, Masaru, et al.
Publicado: (2024)
Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models
por: Kumon, Ryoma, et al.
Publicado: (2026)
por: Kumon, Ryoma, et al.
Publicado: (2026)
Can Large Language Models Robustly Perform Natural Language Inference for Japanese Comparatives?
por: Mikami, Yosuke, et al.
Publicado: (2025)
por: Mikami, Yosuke, et al.
Publicado: (2025)
What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?
por: Ryu, Koki, et al.
Publicado: (2026)
por: Ryu, Koki, et al.
Publicado: (2026)
Bridging Perception and Language: A Systematic Benchmark for LVLMs' Understanding of Amodal Completion Reports
por: Watahiki, Amane, et al.
Publicado: (2025)
por: Watahiki, Amane, et al.
Publicado: (2025)
Analyzing the Inner Workings of Transformers in Compositional Generalization
por: Kumon, Ryoma, et al.
Publicado: (2025)
por: Kumon, Ryoma, et al.
Publicado: (2025)
Neuron-Level Analysis of Cultural Understanding in Large Language Models
por: Yamamoto, Taisei, et al.
Publicado: (2025)
por: Yamamoto, Taisei, et al.
Publicado: (2025)
What's New in My Data? Novelty Exploration via Contrastive Generation
por: Isonuma, Masaru, et al.
Publicado: (2024)
por: Isonuma, Masaru, et al.
Publicado: (2024)
UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
por: Lu, Huimin, et al.
Publicado: (2025)
por: Lu, Huimin, et al.
Publicado: (2025)
Enhancing Rating Prediction with Off-the-Shelf LLMs Using In-Context User Reviews
por: Ryu, Koki, et al.
Publicado: (2025)
por: Ryu, Koki, et al.
Publicado: (2025)
NeuronMoE: Neuron-Guided Mixture-of-Experts for Efficient Multilingual LLM Extension
por: Li, Rongzhi, et al.
Publicado: (2026)
por: Li, Rongzhi, et al.
Publicado: (2026)
Evaluating Structural Generalization in Neural Machine Translation
por: Kumon, Ryoma, et al.
Publicado: (2024)
por: Kumon, Ryoma, et al.
Publicado: (2024)
Derivational Probing: Unveiling the Layer-wise Derivation of Syntactic Structures in Neural Language Models
por: Someya, Taiga, et al.
Publicado: (2025)
por: Someya, Taiga, et al.
Publicado: (2025)
Instability in Downstream Task Performance During LLM Pretraining
por: Nishida, Yuto, et al.
Publicado: (2025)
por: Nishida, Yuto, et al.
Publicado: (2025)
Developing a Guideline for the Labovian-Structural Analysis of Oral Narratives in Japanese
por: Watahiki, Amane, et al.
Publicado: (2026)
por: Watahiki, Amane, et al.
Publicado: (2026)
LLMs Struggle with NLI for Perfect Aspect: A Cross-Linguistic Study in Chinese and Japanese
por: Lu, Jie, et al.
Publicado: (2025)
por: Lu, Jie, et al.
Publicado: (2025)
Implementing a Logical Inference System for Japanese Comparatives
por: Mikami, Yosuke, et al.
Publicado: (2025)
por: Mikami, Yosuke, et al.
Publicado: (2025)
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
por: Kojima, Takeshi, et al.
Publicado: (2024)
por: Kojima, Takeshi, et al.
Publicado: (2024)
Exploring Intra and Inter-language Consistency in Embeddings with ICA
por: Li, Rongzhi, et al.
Publicado: (2024)
por: Li, Rongzhi, et al.
Publicado: (2024)
FaithLM: Towards Faithful Explanations for Large Language Models
por: Chuang, Yu-Neng, et al.
Publicado: (2024)
por: Chuang, Yu-Neng, et al.
Publicado: (2024)
Self-Critique and Refinement for Faithful Natural Language Explanations
por: Wang, Yingming, et al.
Publicado: (2025)
por: Wang, Yingming, et al.
Publicado: (2025)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
por: Nakata, Wataru, et al.
Publicado: (2024)
por: Nakata, Wataru, et al.
Publicado: (2024)
JBBQ: Japanese Bias Benchmark for Analyzing Social Biases in Large Language Models
por: Yanaka, Hitomi, et al.
Publicado: (2024)
por: Yanaka, Hitomi, et al.
Publicado: (2024)
Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
por: Yamamoto, Taisei, et al.
Publicado: (2025)
por: Yamamoto, Taisei, et al.
Publicado: (2025)
Towards Transfer Unlearning: Empirical Evidence of Cross-Domain Bias Mitigation
por: Lu, Huimin, et al.
Publicado: (2024)
por: Lu, Huimin, et al.
Publicado: (2024)
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
por: Agarwal, Chirag, et al.
Publicado: (2024)
por: Agarwal, Chirag, et al.
Publicado: (2024)
Exclusive Unlearning
por: Sasaki, Mutsumi, et al.
Publicado: (2026)
por: Sasaki, Mutsumi, et al.
Publicado: (2026)
Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries
por: Ide, Yusuke, et al.
Publicado: (2026)
por: Ide, Yusuke, et al.
Publicado: (2026)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
por: Matton, Katie, et al.
Publicado: (2025)
por: Matton, Katie, et al.
Publicado: (2025)
DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Models
por: Guo, YiQiu, et al.
Publicado: (2025)
por: Guo, YiQiu, et al.
Publicado: (2025)
Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models
por: Gong, Xilin, et al.
Publicado: (2026)
por: Gong, Xilin, et al.
Publicado: (2026)
Towards Faithful Explanations for Text Classification with Robustness Improvement and Explanation Guided Training
por: Li, Dongfang, et al.
Publicado: (2023)
por: Li, Dongfang, et al.
Publicado: (2023)
Intersectional Bias in Japanese Large Language Models from a Contextualized Perspective
por: Yanaka, Hitomi, et al.
Publicado: (2025)
por: Yanaka, Hitomi, et al.
Publicado: (2025)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
por: Siegel, Noah Y., et al.
Publicado: (2024)
por: Siegel, Noah Y., et al.
Publicado: (2024)
Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models
por: Yeo, Wei Jie, et al.
Publicado: (2024)
por: Yeo, Wei Jie, et al.
Publicado: (2024)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
por: Bhan, Milan, et al.
Publicado: (2025)
por: Bhan, Milan, et al.
Publicado: (2025)
Large Language Model Agents Are Not Always Faithful Self-Evolvers
por: Zhao, Weixiang, et al.
Publicado: (2026)
por: Zhao, Weixiang, et al.
Publicado: (2026)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
por: Fragkathoulas, Christos, et al.
Publicado: (2024)
por: Fragkathoulas, Christos, et al.
Publicado: (2024)
Ejemplares similares
-
Comprehensive Evaluation of Large Language Models for Topic Modeling
por: Doi, Tomoki, et al.
Publicado: (2024) -
Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?
por: Shinozaki, Taiga, et al.
Publicado: (2025) -
Unlearning Traces the Influential Training Data of Language Models
por: Isonuma, Masaru, et al.
Publicado: (2024) -
Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models
por: Kumon, Ryoma, et al.
Publicado: (2026) -
Can Large Language Models Robustly Perform Natural Language Inference for Japanese Comparatives?
por: Mikami, Yosuke, et al.
Publicado: (2025)