SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
Fuente:
arXiv
Saved in:
| Main Authors: | Kirchhof, Michael, Füger, Luca, Goliński, Adam, Dhekane, Eeshan Gunesh, Blaas, Arno, Oh, Seong Joon, Williamson, Sinead |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
by: Santilli, Andrea, et al.
Published: (2025)
by: Santilli, Andrea, et al.
Published: (2025)
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026)
by: Ye, Zihuiwen, et al.
Published: (2026)
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
by: de Seyssel, Maureen, et al.
Published: (2025)
by: de Seyssel, Maureen, et al.
Published: (2025)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
by: Nakkiran, Preetum, et al.
Published: (2025)
by: Nakkiran, Preetum, et al.
Published: (2025)
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
by: Choudhury, Deepro, et al.
Published: (2025)
by: Choudhury, Deepro, et al.
Published: (2025)
Considerations for Distribution Shift Robustness of Diagnostic Models in Healthcare
by: Blaas, Arno, et al.
Published: (2024)
by: Blaas, Arno, et al.
Published: (2024)
Thou Shalt Not Prompt: Zero-Shot Human Activity Recognition in Smart Homes via Language Modeling of Sensor Data & Activities
by: Dhekane, Sourish Gunesh, et al.
Published: (2025)
by: Dhekane, Sourish Gunesh, et al.
Published: (2025)
Transfer Learning in Human Activity Recognition: A Survey
by: Dhekane, Sourish Gunesh, et al.
Published: (2024)
by: Dhekane, Sourish Gunesh, et al.
Published: (2024)
Poly-View Contrastive Learning
by: Shidani, Amitis, et al.
Published: (2024)
by: Shidani, Amitis, et al.
Published: (2024)
Benchmarking Uncertainty Disentanglement: Specialized Uncertainties for Specialized Tasks
by: Mucsányi, Bálint, et al.
Published: (2024)
by: Mucsányi, Bálint, et al.
Published: (2024)
LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs
by: Chen, Chacha, et al.
Published: (2026)
by: Chen, Chacha, et al.
Published: (2026)
Pretrained Visual Uncertainties
by: Kirchhof, Michael, et al.
Published: (2024)
by: Kirchhof, Michael, et al.
Published: (2024)
Scaling Properties of Continuous Diffusion Spoken Language Models
by: Ramapuram, Jason, et al.
Published: (2026)
by: Ramapuram, Jason, et al.
Published: (2026)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
by: Crabbé, Jonathan, et al.
Published: (2023)
by: Crabbé, Jonathan, et al.
Published: (2023)
Self-Supervised Learning with Gaussian Processes
by: Duan, Yunshan, et al.
Published: (2025)
by: Duan, Yunshan, et al.
Published: (2025)
DISCOVER: Identifying Patterns of Daily Living in Human Activities from Smart Home Data
by: Karpekov, Alexander, et al.
Published: (2025)
by: Karpekov, Alexander, et al.
Published: (2025)
Layout Agnostic Human Activity Recognition in Smart Homes through Textual Descriptions Of Sensor Triggers (TDOST)
by: Thukral, Megha, et al.
Published: (2024)
by: Thukral, Megha, et al.
Published: (2024)
Versuchsbasierte Entwicklung eines Bemessungskonzeptes zur Risssanierung in Massivlehmbauten (ERiMa)
by: Jörg Röder, et al.
Published: (2025)
by: Jörg Röder, et al.
Published: (2025)
Uncertainties of Latent Representations in Computer Vision
by: Kirchhof, Michael
Published: (2024)
by: Kirchhof, Michael
Published: (2024)
How PARTs assemble into wholes: Learning the relative composition of images
by: Ayoughi, Melika, et al.
Published: (2025)
by: Ayoughi, Melika, et al.
Published: (2025)
Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions
by: Plunkett, Dillon, et al.
Published: (2025)
by: Plunkett, Dillon, et al.
Published: (2025)
Posterior Uncertainty Quantification in Neural Networks using Data Augmentation
by: Wu, Luhuan, et al.
Published: (2024)
by: Wu, Luhuan, et al.
Published: (2024)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
by: Rodriguez, Pau, et al.
Published: (2025)
by: Rodriguez, Pau, et al.
Published: (2025)
Intermediate Layer Classifiers for OOD generalization
by: Uselis, Arnas, et al.
Published: (2025)
by: Uselis, Arnas, et al.
Published: (2025)
First Hallucination Tokens Are Different from Conditional Ones
by: Snel, Jakob, et al.
Published: (2025)
by: Snel, Jakob, et al.
Published: (2025)
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
by: Ramapuram, Jason, et al.
Published: (2024)
by: Ramapuram, Jason, et al.
Published: (2024)
A note on precotangent bundles: the example of Grassmannians
by: Goliński, Tomasz
Published: (2025)
by: Goliński, Tomasz
Published: (2025)
Scalable Ensemble Diversification for OOD Generalization and Detection
by: Rubinstein, Alexander, et al.
Published: (2024)
by: Rubinstein, Alexander, et al.
Published: (2024)
Dr.LLM: Dynamic Layer Routing in LLMs
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
NeuroAssist: Enhancing Cognitive-Computer Synergy with Adaptive AI and Advanced Neural Decoding for Efficient EEG Signal Classification
by: Dandamudi, Eeshan G.
Published: (2024)
by: Dandamudi, Eeshan G.
Published: (2024)
Controlling Language and Diffusion Models by Transporting Activations
by: Rodriguez, Pau, et al.
Published: (2024)
by: Rodriguez, Pau, et al.
Published: (2024)
Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models
by: Goel, Anmol, et al.
Published: (2026)
by: Goel, Anmol, et al.
Published: (2026)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
by: Do, Heejin, et al.
Published: (2025)
by: Do, Heejin, et al.
Published: (2025)
GenCtrl -- A Formal Controllability Toolkit for Generative Models
by: Cheng, Emily, et al.
Published: (2026)
by: Cheng, Emily, et al.
Published: (2026)
Incentivizing LLMs to Self-Verify Their Answers
by: Zhang, Fuxiang, et al.
Published: (2025)
by: Zhang, Fuxiang, et al.
Published: (2025)
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
by: Park, Sangwoo, et al.
Published: (2026)
by: Park, Sangwoo, et al.
Published: (2026)
Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
by: Ghasemabadi, Amirhosein, et al.
Published: (2025)
by: Ghasemabadi, Amirhosein, et al.
Published: (2025)
Does Data Scaling Lead to Visual Compositional Generalization?
by: Uselis, Arnas, et al.
Published: (2025)
by: Uselis, Arnas, et al.
Published: (2025)
Half-Truths Break Similarity-Based Retrieval
by: Kargi, Bora, et al.
Published: (2026)
by: Kargi, Bora, et al.
Published: (2026)
Similar Items
-
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
by: Santilli, Andrea, et al.
Published: (2025) -
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026) -
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
by: de Seyssel, Maureen, et al.
Published: (2025) -
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
by: Nakkiran, Preetum, et al.
Published: (2025) -
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)