Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
Fuente:
arXiv
Salvato in:
| Autori principali: | Siegel, Noah Y., Heess, Nicolas, Perez-Ortiz, Maria, Camburu, Oana-Maria |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
di: Siegel, Noah Y., et al.
Pubblicazione: (2024)
di: Siegel, Noah Y., et al.
Pubblicazione: (2024)
Atomic Inference for NLI with Generated Facts as Atoms
di: Stacey, Joe, et al.
Pubblicazione: (2023)
di: Stacey, Joe, et al.
Pubblicazione: (2023)
The Sufficiency-Conciseness Trade-off in LLM Self-Explanation from an Information Bottleneck Perspective
di: Zahedzadeh, Ali, et al.
Pubblicazione: (2026)
di: Zahedzadeh, Ali, et al.
Pubblicazione: (2026)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2023)
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2023)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
di: Mahale, Ajay Pravin
Pubblicazione: (2026)
di: Mahale, Ajay Pravin
Pubblicazione: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
Inference to the Best Explanation in Large Language Models
di: Dalal, Dhairya, et al.
Pubblicazione: (2024)
di: Dalal, Dhairya, et al.
Pubblicazione: (2024)
Persuasiveness and Bias in LLM: Investigating the Impact of Persuasiveness and Reinforcement of Bias in Language Models
di: Roy, Saumya
Pubblicazione: (2025)
di: Roy, Saumya
Pubblicazione: (2025)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2025)
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2025)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
di: Wróbel, Krzysztof, et al.
Pubblicazione: (2026)
di: Wróbel, Krzysztof, et al.
Pubblicazione: (2026)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
di: Young, Richard J.
Pubblicazione: (2026)
di: Young, Richard J.
Pubblicazione: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
di: Oketunji, Abiodun Finbarrs, et al.
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs, et al.
Pubblicazione: (2023)
Latent Planning Emerges with Scale
di: Hanna, Michael, et al.
Pubblicazione: (2026)
di: Hanna, Michael, et al.
Pubblicazione: (2026)
Is Our Chatbot Telling Lies? Assessing Correctness of an LLM-based Dutch Support Chatbot
di: Lassche, Herman, et al.
Pubblicazione: (2024)
di: Lassche, Herman, et al.
Pubblicazione: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models
di: Chang, Edward Y.
Pubblicazione: (2024)
di: Chang, Edward Y.
Pubblicazione: (2024)
Internal Reasoning vs. External Control: A Thermodynamic Analysis of Sycophancy in Large Language Models
di: Chang, Edward Y.
Pubblicazione: (2025)
di: Chang, Edward Y.
Pubblicazione: (2025)
LASTIST: LArge-Scale Target-Independent STance dataset
di: Kim, DongJae, et al.
Pubblicazione: (2025)
di: Kim, DongJae, et al.
Pubblicazione: (2025)
SD$^2$: Self-Distilled Sparse Drafters
di: Lasby, Mike, et al.
Pubblicazione: (2025)
di: Lasby, Mike, et al.
Pubblicazione: (2025)
Analyzing LLM Reasoning to Uncover Mental Health Stigma
di: Sankar, Sreehari, et al.
Pubblicazione: (2026)
di: Sankar, Sreehari, et al.
Pubblicazione: (2026)
Xinyu: An Efficient LLM-based System for Commentary Generation
di: Wu, Yiquan, et al.
Pubblicazione: (2024)
di: Wu, Yiquan, et al.
Pubblicazione: (2024)
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
Presumed Cultural Identity: How Names Shape LLM Responses
di: Pawar, Siddhesh, et al.
Pubblicazione: (2025)
di: Pawar, Siddhesh, et al.
Pubblicazione: (2025)
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
Component-Aware Self-Speculative Decoding in Hybrid Language Models
di: Borobia, Hector, et al.
Pubblicazione: (2026)
di: Borobia, Hector, et al.
Pubblicazione: (2026)
More Is Not Always Better: Cross-Component Interference in LLM Agent Scaffolding
di: Liu, Ming
Pubblicazione: (2026)
di: Liu, Ming
Pubblicazione: (2026)
On the Effectiveness of LLM-Specific Fine-Tuning for Detecting AI-Generated Text
di: Gromadzki, Michał, et al.
Pubblicazione: (2026)
di: Gromadzki, Michał, et al.
Pubblicazione: (2026)
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
Persona Inconstancy in Multi-Agent LLM Collaboration: Conformity, Confabulation, and Impersonation
di: Baltaji, Razan, et al.
Pubblicazione: (2024)
di: Baltaji, Razan, et al.
Pubblicazione: (2024)
SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM
di: Wang, Quandong, et al.
Pubblicazione: (2024)
di: Wang, Quandong, et al.
Pubblicazione: (2024)
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
di: Chang, Edward Y.
Pubblicazione: (2026)
di: Chang, Edward Y.
Pubblicazione: (2026)
Red Teaming for Large Language Models At Scale: Tackling Hallucinations on Mathematics Tasks
di: Buszydlik, Aleksander, et al.
Pubblicazione: (2023)
di: Buszydlik, Aleksander, et al.
Pubblicazione: (2023)
Blessing or curse? A survey on the Impact of Generative AI on Fake News
di: Loth, Alexander, et al.
Pubblicazione: (2024)
di: Loth, Alexander, et al.
Pubblicazione: (2024)
The Perplexity Paradox: Why Code Compresses Better Than Math in LLM Prompts
di: Johnson, Warren
Pubblicazione: (2026)
di: Johnson, Warren
Pubblicazione: (2026)
Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance
di: Lu, Yaxi, et al.
Pubblicazione: (2024)
di: Lu, Yaxi, et al.
Pubblicazione: (2024)
Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2026)
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2026)
Unleashing the potential of prompt engineering for large language models
di: Chen, Banghao, et al.
Pubblicazione: (2023)
di: Chen, Banghao, et al.
Pubblicazione: (2023)
Serialisation Strategy Matters: How FHIR Data Format Affects LLM Medication Reconciliation
di: Pator, Sanjoy
Pubblicazione: (2026)
di: Pator, Sanjoy
Pubblicazione: (2026)
Assessment of Transformer-Based Encoder-Decoder Model for Human-Like Summarization
di: Nair, Sindhu, et al.
Pubblicazione: (2024)
di: Nair, Sindhu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
di: Siegel, Noah Y., et al.
Pubblicazione: (2024) -
Atomic Inference for NLI with Generated Facts as Atoms
di: Stacey, Joe, et al.
Pubblicazione: (2023) -
The Sufficiency-Conciseness Trade-off in LLM Self-Explanation from an Information Bottleneck Perspective
di: Zahedzadeh, Ali, et al.
Pubblicazione: (2026) -
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2023) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)