Enregistré dans:
| Auteurs principaux: | Mayne, Harry, Kang, Justin Singh, Gould, Dewi, Ramchandran, Kannan, Mahdi, Adam, Siegel, Noah Y. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2602.02639 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
par: Siegel, Noah Y., et autres
Publié: (2025)
par: Siegel, Noah Y., et autres
Publié: (2025)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
par: Siegel, Noah Y., et autres
Publié: (2024)
par: Siegel, Noah Y., et autres
Publié: (2024)
SPEX: Scaling Feature Interaction Explanations for LLMs
par: Kang, Justin Singh, et autres
Publié: (2025)
par: Kang, Justin Singh, et autres
Publié: (2025)
Quantifying Positional Biases in Text Embedding Models
par: Lee, Reagan J., et autres
Publié: (2024)
par: Lee, Reagan J., et autres
Publié: (2024)
Can sparse autoencoders be used to decompose and interpret steering vectors?
par: Mayne, Harry, et autres
Publié: (2024)
par: Mayne, Harry, et autres
Publié: (2024)
An Odd Estimator for Shapley Values
par: Fumagalli, Fabian, et autres
Publié: (2026)
par: Fumagalli, Fabian, et autres
Publié: (2026)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
par: Mayne, Harry, et autres
Publié: (2025)
par: Mayne, Harry, et autres
Publié: (2025)
The Fair Value of Data Under Heterogeneous Privacy Constraints in Federated Learning
par: Kang, Justin, et autres
Publié: (2023)
par: Kang, Justin, et autres
Publié: (2023)
ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
par: Butler, Landon, et autres
Publié: (2025)
par: Butler, Landon, et autres
Publié: (2025)
SAGE: A Realistic Benchmark for Semantic Understanding
par: Goel, Samarth, et autres
Publié: (2025)
par: Goel, Samarth, et autres
Publié: (2025)
EmbedLLM: Learning Compact Representations of Large Language Models
par: Zhuang, Richard, et autres
Publié: (2024)
par: Zhuang, Richard, et autres
Publié: (2024)
Adaptive Sparse Möbius Transforms for Learning Polynomials
par: Erginbas, Yigit Efe, et autres
Publié: (2026)
par: Erginbas, Yigit Efe, et autres
Publié: (2026)
Unsupervised Learning Approaches for Identifying ICU Patient Subgroups: Do Results Generalise?
par: Mayne, Harry, et autres
Publié: (2024)
par: Mayne, Harry, et autres
Publié: (2024)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
par: Fragkathoulas, Christos, et autres
Publié: (2024)
par: Fragkathoulas, Christos, et autres
Publié: (2024)
Small Language Model Helps Resolve Semantic Ambiguity of LLM Prompt
par: Huang, Zhenzhen, et autres
Publié: (2026)
par: Huang, Zhenzhen, et autres
Publié: (2026)
FaithLM: Towards Faithful Explanations for Large Language Models
par: Chuang, Yu-Neng, et autres
Publié: (2024)
par: Chuang, Yu-Neng, et autres
Publié: (2024)
Towards Anytime-Valid Statistical Watermarking
par: Huang, Baihe, et autres
Publié: (2026)
par: Huang, Baihe, et autres
Publié: (2026)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
par: Alon, Bar, et autres
Publié: (2026)
par: Alon, Bar, et autres
Publié: (2026)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
par: Quan, Xin, et autres
Publié: (2025)
par: Quan, Xin, et autres
Publié: (2025)
The Effect of Model Size on LLM Post-hoc Explainability via LIME
par: Heyen, Henning, et autres
Publié: (2024)
par: Heyen, Henning, et autres
Publié: (2024)
Learning to Understand: Identifying Interactions via the Möbius Transform
par: Kang, Justin S., et autres
Publié: (2024)
par: Kang, Justin S., et autres
Publié: (2024)
SKATE, a Scalable Tournament Eval: Weaker LLMs differentiate between stronger ones using verifiable challenges
par: Gould, Dewi S. W., et autres
Publié: (2025)
par: Gould, Dewi S. W., et autres
Publié: (2025)
Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
par: Ding, Sihao, et autres
Publié: (2025)
par: Ding, Sihao, et autres
Publié: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
par: Wojciechowski, Adam, et autres
Publié: (2024)
par: Wojciechowski, Adam, et autres
Publié: (2024)
Large language models can help boost food production, but be mindful of their risks
par: De Clercq, Djavan, et autres
Publié: (2024)
par: De Clercq, Djavan, et autres
Publié: (2024)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
par: Parcalabescu, Letitia, et autres
Publié: (2023)
par: Parcalabescu, Letitia, et autres
Publié: (2023)
Neuro-Argumentative Learning with Case-Based Reasoning
par: Gould, Adam, et autres
Publié: (2025)
par: Gould, Adam, et autres
Publié: (2025)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
par: Li, Chloe, et autres
Publié: (2025)
par: Li, Chloe, et autres
Publié: (2025)
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
par: Sovrano, Francesco, et autres
Publié: (2026)
par: Sovrano, Francesco, et autres
Publié: (2026)
DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
par: Guo, Yuhan, et autres
Publié: (2025)
par: Guo, Yuhan, et autres
Publié: (2025)
AirTrafficGen: Configurable Air Traffic Scenario Generation with Large Language Models
par: Gould, Dewi Sid William, et autres
Publié: (2025)
par: Gould, Dewi Sid William, et autres
Publié: (2025)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
par: Yang, Yushi, et autres
Publié: (2024)
par: Yang, Yushi, et autres
Publié: (2024)
A framework for assuring the accuracy and fidelity of an AI-enabled Digital Twin of en route UK airspace
par: Keane, Adam, et autres
Publié: (2026)
par: Keane, Adam, et autres
Publié: (2026)
Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace Adjustments
par: Zhang, Yifan, et autres
Publié: (2026)
par: Zhang, Yifan, et autres
Publié: (2026)
Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations
par: Manna, Supriya, et autres
Publié: (2024)
par: Manna, Supriya, et autres
Publié: (2024)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
par: Matton, Katie, et autres
Publié: (2025)
par: Matton, Katie, et autres
Publié: (2025)
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
par: Khouja, Jude, et autres
Publié: (2025)
par: Khouja, Jude, et autres
Publié: (2025)
Toward a Theory of Tokenization in LLMs
par: Rajaraman, Nived, et autres
Publié: (2024)
par: Rajaraman, Nived, et autres
Publié: (2024)
Negation Neglect: When models fail to learn negations in training
par: Mayne, Harry, et autres
Publié: (2026)
par: Mayne, Harry, et autres
Publié: (2026)
Evaluating Readability and Faithfulness of Concept-based Explanations
par: Li, Meng, et autres
Publié: (2024)
par: Li, Meng, et autres
Publié: (2024)
Documents similaires
-
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
par: Siegel, Noah Y., et autres
Publié: (2025) -
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
par: Siegel, Noah Y., et autres
Publié: (2024) -
SPEX: Scaling Feature Interaction Explanations for LLMs
par: Kang, Justin Singh, et autres
Publié: (2025) -
Quantifying Positional Biases in Text Embedding Models
par: Lee, Reagan J., et autres
Publié: (2024) -
Can sparse autoencoders be used to decompose and interpret steering vectors?
par: Mayne, Harry, et autres
Publié: (2024)