Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
Fuente:
arXiv
Salvato in:
| Autori principali: | Patel, Het, Chen, Tiejin, Wei, Hua, Papalexakis, Evangelos E., Chen, Jia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
di: McCann, Jordan F.
Pubblicazione: (2026)
di: McCann, Jordan F.
Pubblicazione: (2026)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
di: Ye, Hua, et al.
Pubblicazione: (2025)
di: Ye, Hua, et al.
Pubblicazione: (2025)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
di: Sun, Kangkang, et al.
Pubblicazione: (2026)
di: Sun, Kangkang, et al.
Pubblicazione: (2026)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
di: Mi, Zhendong, et al.
Pubblicazione: (2025)
di: Mi, Zhendong, et al.
Pubblicazione: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
di: Fadli, Samih
Pubblicazione: (2025)
di: Fadli, Samih
Pubblicazione: (2025)
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
di: Chen, Tiejin, et al.
Pubblicazione: (2026)
di: Chen, Tiejin, et al.
Pubblicazione: (2026)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
di: Cho, Seonglae, et al.
Pubblicazione: (2026)
di: Cho, Seonglae, et al.
Pubblicazione: (2026)
Node-Level Uncertainty Estimation in LLM-Generated SQL
di: Hasson, Hilaf, et al.
Pubblicazione: (2025)
di: Hasson, Hilaf, et al.
Pubblicazione: (2025)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
di: Mi, Zhendong, et al.
Pubblicazione: (2025)
di: Mi, Zhendong, et al.
Pubblicazione: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
di: Kugler, Kai
Pubblicazione: (2025)
di: Kugler, Kai
Pubblicazione: (2025)
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
di: Chen, Xinjie, et al.
Pubblicazione: (2026)
di: Chen, Xinjie, et al.
Pubblicazione: (2026)
Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
di: Keeman, Michael
Pubblicazione: (2026)
di: Keeman, Michael
Pubblicazione: (2026)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
di: Rozonoyer, Benjamin, et al.
Pubblicazione: (2026)
di: Rozonoyer, Benjamin, et al.
Pubblicazione: (2026)
Automated Bug Triaging using Instruction-Tuned Large Language Models
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
di: Schesch, Benedikt, et al.
Pubblicazione: (2026)
di: Schesch, Benedikt, et al.
Pubblicazione: (2026)
Survey Transfer Learning: Recycling Data with Silicon Responses
di: Amini, Ali
Pubblicazione: (2025)
di: Amini, Ali
Pubblicazione: (2025)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
di: Hashemi, Helia, et al.
Pubblicazione: (2024)
di: Hashemi, Helia, et al.
Pubblicazione: (2024)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
di: Zhu, Jiajun, et al.
Pubblicazione: (2025)
di: Zhu, Jiajun, et al.
Pubblicazione: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
di: Du, Bangde, et al.
Pubblicazione: (2025)
di: Du, Bangde, et al.
Pubblicazione: (2025)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
di: Yuan, Xin, et al.
Pubblicazione: (2025)
di: Yuan, Xin, et al.
Pubblicazione: (2025)
ContractBench: Can LLM Agents Preserve Observation Contracts?
di: Wang, Jicheng, et al.
Pubblicazione: (2026)
di: Wang, Jicheng, et al.
Pubblicazione: (2026)
Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
di: Bianchessi, Arthur S., et al.
Pubblicazione: (2025)
di: Bianchessi, Arthur S., et al.
Pubblicazione: (2025)
PersonalLLM: Tailoring LLMs to Individual Preferences
di: Zollo, Thomas P., et al.
Pubblicazione: (2024)
di: Zollo, Thomas P., et al.
Pubblicazione: (2024)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
di: Gwak, Jiho, et al.
Pubblicazione: (2025)
di: Gwak, Jiho, et al.
Pubblicazione: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
di: Liu, Zhongxin, et al.
Pubblicazione: (2025)
di: Liu, Zhongxin, et al.
Pubblicazione: (2025)
On the Influence of Discourse Relations in Persuasive Texts
di: Turk, Nawar, et al.
Pubblicazione: (2025)
di: Turk, Nawar, et al.
Pubblicazione: (2025)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
di: Easley, Eric, et al.
Pubblicazione: (2026)
di: Easley, Eric, et al.
Pubblicazione: (2026)
Calibrated Confidence Estimation for Tabular Question Answering
di: Voss, Lukas
Pubblicazione: (2026)
di: Voss, Lukas
Pubblicazione: (2026)
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
di: Liao, Jianxing, et al.
Pubblicazione: (2025)
di: Liao, Jianxing, et al.
Pubblicazione: (2025)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
di: Saxena, Udit
Pubblicazione: (2025)
di: Saxena, Udit
Pubblicazione: (2025)
Super Apriel: One Checkpoint, Many Speeds
di: Labs, SLAM, et al.
Pubblicazione: (2026)
di: Labs, SLAM, et al.
Pubblicazione: (2026)
OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind
di: Srishty, Sharmin Sultana, et al.
Pubblicazione: (2026)
di: Srishty, Sharmin Sultana, et al.
Pubblicazione: (2026)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
di: Han, Xudong, et al.
Pubblicazione: (2025)
di: Han, Xudong, et al.
Pubblicazione: (2025)
Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus
di: Arabov, Mullosharaf K.
Pubblicazione: (2026)
di: Arabov, Mullosharaf K.
Pubblicazione: (2026)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
di: Resck, Lucas, et al.
Pubblicazione: (2026)
di: Resck, Lucas, et al.
Pubblicazione: (2026)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
di: Patarlapalli, Sai Babu, et al.
Pubblicazione: (2026)
di: Patarlapalli, Sai Babu, et al.
Pubblicazione: (2026)
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation
di: Bertina, Abbas, et al.
Pubblicazione: (2025)
di: Bertina, Abbas, et al.
Pubblicazione: (2025)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
di: Bakish, Yarden, et al.
Pubblicazione: (2025)
di: Bakish, Yarden, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
di: McCann, Jordan F.
Pubblicazione: (2026) -
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
di: Ye, Hua, et al.
Pubblicazione: (2025) -
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
di: Sun, Kangkang, et al.
Pubblicazione: (2026) -
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
di: Cho, Seonglae, et al.
Pubblicazione: (2025) -
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
di: Mi, Zhendong, et al.
Pubblicazione: (2025)