Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yongjie, Wang, Yibo, Zhou, Xin, Shen, Zhiqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
by: Wang, Yongjie, et al.
Published: (2025)
by: Wang, Yongjie, et al.
Published: (2025)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
by: Huang, Yingbing, et al.
Published: (2025)
by: Huang, Yingbing, et al.
Published: (2025)
Large Language Models Report Subjective Experience Under Self-Referential Processing
by: Berg, Cameron, et al.
Published: (2025)
by: Berg, Cameron, et al.
Published: (2025)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Proceedings of the 20th International Conference on Knowledge, Information and Creativity Support Systems (KICSS 2025)
by: Hayama, Edited by Tessai, et al.
Published: (2025)
by: Hayama, Edited by Tessai, et al.
Published: (2025)
Vibe-Creation: The Epistemology of Human-AI Emergent Cognition
by: Levin, Ilya
Published: (2026)
by: Levin, Ilya
Published: (2026)
BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models
by: Abhishek, Alok, et al.
Published: (2025)
by: Abhishek, Alok, et al.
Published: (2025)
Stealth edits to large language models
by: Sutton, Oliver J., et al.
Published: (2024)
by: Sutton, Oliver J., et al.
Published: (2024)
SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
by: Abhishek, Alok, et al.
Published: (2026)
by: Abhishek, Alok, et al.
Published: (2026)
Towards Interpretable Soft Prompts
by: Patel, Oam, et al.
Published: (2025)
by: Patel, Oam, et al.
Published: (2025)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026)
by: Sarkar, Nilesh, et al.
Published: (2026)
Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
by: Rosales, Rafael, et al.
Published: (2025)
by: Rosales, Rafael, et al.
Published: (2025)
Data and AI governance: Promoting equity, ethics, and fairness in large language models
by: Abhishek, Alok, et al.
Published: (2025)
by: Abhishek, Alok, et al.
Published: (2025)
Reasoning Promotes Robustness in Theory of Mind Tasks
by: de Haan, Ian B., et al.
Published: (2026)
by: de Haan, Ian B., et al.
Published: (2026)
Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
by: Plevris, Vagelis, et al.
Published: (2023)
by: Plevris, Vagelis, et al.
Published: (2023)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
by: Buchner, Valentin Leonhard, et al.
Published: (2023)
by: Buchner, Valentin Leonhard, et al.
Published: (2023)
An Iterative Optimizing Framework for Radiology Report Summarization with ChatGPT
by: Ma, Chong, et al.
Published: (2023)
by: Ma, Chong, et al.
Published: (2023)
A Mixed User-Centered Approach to Enable Augmented Intelligence in Intelligent Tutoring Systems: The Case of MathAIde app
by: Guerino, Guilherme, et al.
Published: (2025)
by: Guerino, Guilherme, et al.
Published: (2025)
Developing Acoustic Models for Automatic Speech Recognition in Swedish
by: Salvi, Giampiero
Published: (2024)
by: Salvi, Giampiero
Published: (2024)
On measuring grounding and generalizing grounding problems
by: Quigley, Daniel, et al.
Published: (2025)
by: Quigley, Daniel, et al.
Published: (2025)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
by: Badshah, Sher, et al.
Published: (2025)
by: Badshah, Sher, et al.
Published: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
by: Yang, Yibo
Published: (2025)
by: Yang, Yibo
Published: (2025)
MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation
by: Lim, Ernest, et al.
Published: (2025)
by: Lim, Ernest, et al.
Published: (2025)
Enhancing Feature Selection and Interpretability in AI Regression Tasks Through Feature Attribution
by: Hinterleitner, Alexander, et al.
Published: (2024)
by: Hinterleitner, Alexander, et al.
Published: (2024)
From Language Models to Practical Self-Improving Computer Agents
by: Sheng, Alex
Published: (2024)
by: Sheng, Alex
Published: (2024)
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference
by: Jørgensen, Tollef Emil
Published: (2025)
by: Jørgensen, Tollef Emil
Published: (2025)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
by: Demir, M. Mikail, et al.
Published: (2025)
by: Demir, M. Mikail, et al.
Published: (2025)
Autonomous Editorial Systems and Computational Investigation with Artificial Intelligence
by: Banafea, Ahmed
Published: (2026)
by: Banafea, Ahmed
Published: (2026)
The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
Interpretability Can Be Actionable
by: Orgad, Hadas, et al.
Published: (2026)
by: Orgad, Hadas, et al.
Published: (2026)
CHORUS: An Agentic Framework for Generating Realistic Deliberation Data
by: Koursaris, A., et al.
Published: (2026)
by: Koursaris, A., et al.
Published: (2026)
FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
by: Mohammadzadeh, Saeed, et al.
Published: (2025)
by: Mohammadzadeh, Saeed, et al.
Published: (2025)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
by: Otal, Hakan T., et al.
Published: (2024)
by: Otal, Hakan T., et al.
Published: (2024)
Agentic AI Systems Applied to tasks in Financial Services: Modeling and model risk management crews
by: Okpala, Izunna, et al.
Published: (2025)
by: Okpala, Izunna, et al.
Published: (2025)
Feature Relevancy, Necessity and Usefulness: Complexity and Algorithms
by: Capdevielle, Tomás, et al.
Published: (2025)
by: Capdevielle, Tomás, et al.
Published: (2025)
A Theoretical Framework for Adaptive Utility-Weighted Benchmarking
by: Waggoner, Philip
Published: (2026)
by: Waggoner, Philip
Published: (2026)
A Taxonomy of Omnicidal Futures Involving Artificial Intelligence
by: Critch, Andrew, et al.
Published: (2025)
by: Critch, Andrew, et al.
Published: (2025)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
by: Anam, Rizal Khoirul
Published: (2025)
by: Anam, Rizal Khoirul
Published: (2025)
Achieving Distributive Justice in Federated Learning via Uncertainty Quantification
by: Carey, Alycia, et al.
Published: (2025)
by: Carey, Alycia, et al.
Published: (2025)
Similar Items
-
RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
by: Wang, Yongjie, et al.
Published: (2025) -
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
by: Huang, Yingbing, et al.
Published: (2025) -
Large Language Models Report Subjective Experience Under Self-Referential Processing
by: Berg, Cameron, et al.
Published: (2025) -
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024) -
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)