The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cho, Seonglae, Wu, Zekun, Da Costa, Kleyton, Koshiyama, Adriano |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
Tool Calling is Linearly Readable and Steerable in Language Models
von: Wu, Zekun, et al.
Veröffentlicht: (2026)
von: Wu, Zekun, et al.
Veröffentlicht: (2026)
Evaluating Explainability in Machine Learning Predictions through Explainer-Agnostic Metrics
von: Munoz, Cristian, et al.
Veröffentlicht: (2023)
von: Munoz, Cristian, et al.
Veröffentlicht: (2023)
Eliciting Personality Traits in Large Language Models
von: Hilliard, Airlie, et al.
Veröffentlicht: (2024)
von: Hilliard, Airlie, et al.
Veröffentlicht: (2024)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
Confidence Regulation Neurons in Language Models
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
On the Geometric Structure of Layer Updates in Deep Language Models
von: Yoo, Jun-Sik
Veröffentlicht: (2026)
von: Yoo, Jun-Sik
Veröffentlicht: (2026)
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach
von: Wu, Zekun, et al.
Veröffentlicht: (2024)
von: Wu, Zekun, et al.
Veröffentlicht: (2024)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion
von: Guan, Xin, et al.
Veröffentlicht: (2025)
von: Guan, Xin, et al.
Veröffentlicht: (2025)
Confidence-Modulated Speculative Decoding for Large Language Models
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
Confidence Regularized Masked Language Modeling using Text Length
von: Ji, Seunghyun, et al.
Veröffentlicht: (2025)
von: Ji, Seunghyun, et al.
Veröffentlicht: (2025)
SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
von: Cheng, Wuxinlin, et al.
Veröffentlicht: (2025)
von: Cheng, Wuxinlin, et al.
Veröffentlicht: (2025)
LibVulnWatch: A Deep Assessment Agent System and Leaderboard for Uncovering Hidden Vulnerabilities in Open-Source AI Libraries
von: Wu, Zekun, et al.
Veröffentlicht: (2025)
von: Wu, Zekun, et al.
Veröffentlicht: (2025)
Large Language Model Confidence Estimation via Black-Box Access
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
von: Dawes, Cutter, et al.
Veröffentlicht: (2026)
von: Dawes, Cutter, et al.
Veröffentlicht: (2026)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
von: Galeone, Cosimo, et al.
Veröffentlicht: (2026)
von: Galeone, Cosimo, et al.
Veröffentlicht: (2026)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
von: Kyriakou, Athina, et al.
Veröffentlicht: (2026)
von: Kyriakou, Athina, et al.
Veröffentlicht: (2026)
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models
von: Wu, Da, et al.
Veröffentlicht: (2023)
von: Wu, Da, et al.
Veröffentlicht: (2023)
ReFT: Representation Finetuning for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
Manifold-based Sampling for In-Context Hallucination Detection in Large Language Models
von: Vamshi, Bodla Krishna, et al.
Veröffentlicht: (2026)
von: Vamshi, Bodla Krishna, et al.
Veröffentlicht: (2026)
Unveiling Imitation Learning: Exploring the Impact of Data Falsity to Large Language Model
von: Cho, Hyunsoo
Veröffentlicht: (2024)
von: Cho, Hyunsoo
Veröffentlicht: (2024)
CARE-RFT: Confidence-Anchored Reinforcement Finetuning for Reliable Reasoning in Large Language Models
von: Li, Shuozhe, et al.
Veröffentlicht: (2026)
von: Li, Shuozhe, et al.
Veröffentlicht: (2026)
EmbedLLM: Learning Compact Representations of Large Language Models
von: Zhuang, Richard, et al.
Veröffentlicht: (2024)
von: Zhuang, Richard, et al.
Veröffentlicht: (2024)
ProgCo: Program Helps Self-Correction of Large Language Models
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2025)
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2025)
A Survey of Quantized Graph Representation Learning: Connecting Graph Structures with Large Language Models
von: Lin, Qika, et al.
Veröffentlicht: (2025)
von: Lin, Qika, et al.
Veröffentlicht: (2025)
Synthesizing and Adapting Error Correction Data for Mobile Large Language Model Applications
von: Zhang, Yanxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Yanxiang, et al.
Veröffentlicht: (2025)
Representation Learning of Structured Data for Medical Foundation Models
von: Dwivedi, Vijay Prakash, et al.
Veröffentlicht: (2024)
von: Dwivedi, Vijay Prakash, et al.
Veröffentlicht: (2024)
Graph-based Confidence Calibration for Large Language Models
von: Li, Yukun, et al.
Veröffentlicht: (2024)
von: Li, Yukun, et al.
Veröffentlicht: (2024)
Clarify: Improving Model Robustness With Natural Language Corrections
von: Lee, Yoonho, et al.
Veröffentlicht: (2024)
von: Lee, Yoonho, et al.
Veröffentlicht: (2024)
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
von: Tsui, Ken
Veröffentlicht: (2025)
von: Tsui, Ken
Veröffentlicht: (2025)
CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models
von: Lv, Yaojia, et al.
Veröffentlicht: (2024)
von: Lv, Yaojia, et al.
Veröffentlicht: (2024)
Early Stopping for Large Reasoning Models via Confidence Dynamics
von: Hosseini, Parsa, et al.
Veröffentlicht: (2026)
von: Hosseini, Parsa, et al.
Veröffentlicht: (2026)
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
von: Xie, Tong, et al.
Veröffentlicht: (2025)
von: Xie, Tong, et al.
Veröffentlicht: (2025)
Leviathan: Decoupling Input and Output Representations in Language Models
von: Batley, Reza T., et al.
Veröffentlicht: (2026)
von: Batley, Reza T., et al.
Veröffentlicht: (2026)
Value-Aware Numerical Representations for Transformer Language Models
von: Dutulescu, Andreea, et al.
Veröffentlicht: (2026)
von: Dutulescu, Andreea, et al.
Veröffentlicht: (2026)
Layer by Layer: Uncovering Hidden Representations in Language Models
von: Skean, Oscar, et al.
Veröffentlicht: (2025)
von: Skean, Oscar, et al.
Veröffentlicht: (2025)
The Linear Representation Hypothesis and the Geometry of Large Language Models
von: Park, Kiho, et al.
Veröffentlicht: (2023)
von: Park, Kiho, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2026) -
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2025) -
Tool Calling is Linearly Readable and Steerable in Language Models
von: Wu, Zekun, et al.
Veröffentlicht: (2026) -
Evaluating Explainability in Machine Learning Predictions through Explainer-Agnostic Metrics
von: Munoz, Cristian, et al.
Veröffentlicht: (2023) -
Eliciting Personality Traits in Large Language Models
von: Hilliard, Airlie, et al.
Veröffentlicht: (2024)