Self-Aware Knowledge Probing: Evaluating Language Models' Relational Knowledge through Confidence Calibration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kissling, Christopher, Merdjanovska, Elena, Akbik, Alan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
von: Merdjanovska, Elena, et al.
Veröffentlicht: (2024)
von: Merdjanovska, Elena, et al.
Veröffentlicht: (2024)
BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models
von: Wiland, Jacek, et al.
Veröffentlicht: (2024)
von: Wiland, Jacek, et al.
Veröffentlicht: (2024)
LM-PUB-QUIZ: A Comprehensive Framework for Zero-Shot Evaluation of Relational Knowledge in Language Models
von: Ploner, Max, et al.
Veröffentlicht: (2024)
von: Ploner, Max, et al.
Veröffentlicht: (2024)
Towards a Principled Evaluation of Knowledge Editors
von: Pohl, Sebastian, et al.
Veröffentlicht: (2025)
von: Pohl, Sebastian, et al.
Veröffentlicht: (2025)
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
von: Christoph, Daniel, et al.
Veröffentlicht: (2025)
von: Christoph, Daniel, et al.
Veröffentlicht: (2025)
MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration
von: He, Zhitao, et al.
Veröffentlicht: (2025)
von: He, Zhitao, et al.
Veröffentlicht: (2025)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2025)
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2025)
Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation
von: Rücker, Susanna, et al.
Veröffentlicht: (2025)
von: Rücker, Susanna, et al.
Veröffentlicht: (2025)
SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2024)
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2024)
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
von: Haller, Patrick, et al.
Veröffentlicht: (2024)
von: Haller, Patrick, et al.
Veröffentlicht: (2024)
TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification Tasks
von: Garbas, Lukas, et al.
Veröffentlicht: (2024)
von: Garbas, Lukas, et al.
Veröffentlicht: (2024)
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2026)
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2026)
What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
von: Toporkov, Olia, et al.
Veröffentlicht: (2025)
von: Toporkov, Olia, et al.
Veröffentlicht: (2025)
Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
von: Lu, Yuyin, et al.
Veröffentlicht: (2026)
von: Lu, Yuyin, et al.
Veröffentlicht: (2026)
Probing Language Models on Their Knowledge Source
von: Tighidet, Zineddine, et al.
Veröffentlicht: (2024)
von: Tighidet, Zineddine, et al.
Veröffentlicht: (2024)
What Matters When Building Universal Multilingual Named Entity Recognition Models?
von: Golde, Jonas, et al.
Veröffentlicht: (2026)
von: Golde, Jonas, et al.
Veröffentlicht: (2026)
Confidence Preservation Property in Knowledge Distillation Abstractions
von: Vengertsev, Dmitry, et al.
Veröffentlicht: (2024)
von: Vengertsev, Dmitry, et al.
Veröffentlicht: (2024)
Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs
von: Williams, Tristan, et al.
Veröffentlicht: (2026)
von: Williams, Tristan, et al.
Veröffentlicht: (2026)
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
von: Wei, Rongzhe, et al.
Veröffentlicht: (2025)
von: Wei, Rongzhe, et al.
Veröffentlicht: (2025)
Self-training Large Language Models through Knowledge Detection
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2024)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2024)
LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
Calibrating LLM Confidence by Probing Perturbed Representation Stability
von: Khanmohammadi, Reza, et al.
Veröffentlicht: (2025)
von: Khanmohammadi, Reza, et al.
Veröffentlicht: (2025)
CoCoA: Confidence and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models
von: Khandelwal, Anant, et al.
Veröffentlicht: (2025)
von: Khandelwal, Anant, et al.
Veröffentlicht: (2025)
Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
von: Golde, Jonas, et al.
Veröffentlicht: (2025)
von: Golde, Jonas, et al.
Veröffentlicht: (2025)
Large-Scale Label Interpretation Learning for Few-Shot Named Entity Recognition
von: Golde, Jonas, et al.
Veröffentlicht: (2024)
von: Golde, Jonas, et al.
Veröffentlicht: (2024)
Calibrating the Confidence of Large Language Models by Eliciting Fidelity
von: Zhang, Mozhi, et al.
Veröffentlicht: (2024)
von: Zhang, Mozhi, et al.
Veröffentlicht: (2024)
A Comprehensive Evaluation of Semantic Relation Knowledge of Pretrained Language Models and Humans
von: Cao, Zhihan, et al.
Veröffentlicht: (2024)
von: Cao, Zhihan, et al.
Veröffentlicht: (2024)
QA-Calibration of Language Model Confidence Scores
von: Manggala, Putra, et al.
Veröffentlicht: (2024)
von: Manggala, Putra, et al.
Veröffentlicht: (2024)
Knowledge-Aware Query Expansion with Large Language Models for Textual and Relational Retrieval
von: Xia, Yu, et al.
Veröffentlicht: (2024)
von: Xia, Yu, et al.
Veröffentlicht: (2024)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
von: Schulte, David, et al.
Veröffentlicht: (2024)
von: Schulte, David, et al.
Veröffentlicht: (2024)
Knowledge-Aware Self-Correction in Language Models via Structured Memory Graphs
von: Saha, Swayamjit
Veröffentlicht: (2025)
von: Saha, Swayamjit
Veröffentlicht: (2025)
Tracing Relational Knowledge Recall in Large Language Models
von: Popovič, Nicholas, et al.
Veröffentlicht: (2026)
von: Popovič, Nicholas, et al.
Veröffentlicht: (2026)
Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
von: Li, Alan, et al.
Veröffentlicht: (2025)
von: Li, Alan, et al.
Veröffentlicht: (2025)
Self-Consolidating Language Models: Continual Knowledge Incorporation from Context
von: Wang, Zekun, et al.
Veröffentlicht: (2026)
von: Wang, Zekun, et al.
Veröffentlicht: (2026)
Calibrating Verbalized Confidence with Self-Generated Distractors
von: Wang, Victor, et al.
Veröffentlicht: (2025)
von: Wang, Victor, et al.
Veröffentlicht: (2025)
Fact-Level Confidence Calibration and Self-Correction
von: Yuan, Yige, et al.
Veröffentlicht: (2024)
von: Yuan, Yige, et al.
Veröffentlicht: (2024)
Probing and Steering Evaluation Awareness of Language Models
von: Nguyen, Jord, et al.
Veröffentlicht: (2025)
von: Nguyen, Jord, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
von: Merdjanovska, Elena, et al.
Veröffentlicht: (2024) -
BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models
von: Wiland, Jacek, et al.
Veröffentlicht: (2024) -
LM-PUB-QUIZ: A Comprehensive Framework for Zero-Shot Evaluation of Relational Knowledge in Language Models
von: Ploner, Max, et al.
Veröffentlicht: (2024) -
Towards a Principled Evaluation of Knowledge Editors
von: Pohl, Sebastian, et al.
Veröffentlicht: (2025) -
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
von: Christoph, Daniel, et al.
Veröffentlicht: (2025)