Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Yuxi, de Araujo, Pedro Henrique Luz, Zaporojets, Klim, Roth, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Black-box Model Ensembling for Textual and Visual Question Answering via Information Fusion
by: Xia, Yuxi, et al.
Published: (2024)
by: Xia, Yuxi, et al.
Published: (2024)
Analysing zero-shot temporal relation extraction on clinical notes using temporal consistency
by: Kougia, Vasiliki, et al.
Published: (2024)
by: Kougia, Vasiliki, et al.
Published: (2024)
Functionality learning through specification instructions
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2023)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2023)
Helpful assistant or fruitful facilitator? Investigating how personas affect language model behavior
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2024)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2024)
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
Exploring prompts to elicit memorization in masked language model-based named entity recognition
by: Xia, Yuxi, et al.
Published: (2024)
by: Xia, Yuxi, et al.
Published: (2024)
EMERGE: A Benchmark for Updating Knowledge Graphs with Emerging Textual Knowledge
by: Zaporojets, Klim, et al.
Published: (2025)
by: Zaporojets, Klim, et al.
Published: (2025)
Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
by: Xia, Yuxi, et al.
Published: (2026)
by: Xia, Yuxi, et al.
Published: (2026)
Explaining Generalization of AI-Generated Text Detectors Through Linguistic Analysis
by: Xia, Yuxi, et al.
Published: (2026)
by: Xia, Yuxi, et al.
Published: (2026)
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
by: Xia, Yuxi, et al.
Published: (2026)
by: Xia, Yuxi, et al.
Published: (2026)
An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text
by: Schoenegger, Loris, et al.
Published: (2024)
by: Schoenegger, Loris, et al.
Published: (2024)
Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
Style-Compress: An LLM-Based Prompt Compression Framework Considering Task-Specific Styles
by: Pu, Xiao, et al.
Published: (2024)
by: Pu, Xiao, et al.
Published: (2024)
Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations
by: Hong, Pingjun, et al.
Published: (2026)
by: Hong, Pingjun, et al.
Published: (2026)
Low-Resource Machine Translation through Retrieval-Augmented LLM Prompting: A Study on the Mambai Language
by: Merx, Raphaël, et al.
Published: (2024)
by: Merx, Raphaël, et al.
Published: (2024)
Conversational User-AI Intervention: A Study on Prompt Rewriting for Improved LLM Response Generation
by: Sarkar, Rupak, et al.
Published: (2025)
by: Sarkar, Rupak, et al.
Published: (2025)
Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance
by: Yin, Ziqi, et al.
Published: (2024)
by: Yin, Ziqi, et al.
Published: (2024)
Are a Thousand Words Better Than a Single Picture? Beyond Images -- A Framework for Multi-Modal Knowledge Graph Dataset Enrichment
by: Zhang, Pengyu, et al.
Published: (2026)
by: Zhang, Pengyu, et al.
Published: (2026)
How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
by: Jung, Julie, et al.
Published: (2025)
by: Jung, Julie, et al.
Published: (2025)
Show and Tell: Prompt Strategies for Style Control in Multi-Turn LLM Code Generation
by: Bohr, Jeremiah
Published: (2025)
by: Bohr, Jeremiah
Published: (2025)
REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
by: Kong, Chuyi, et al.
Published: (2025)
by: Kong, Chuyi, et al.
Published: (2025)
Towards LLMs Robustness to Changes in Prompt Format Styles
by: Ngweta, Lilian, et al.
Published: (2025)
by: Ngweta, Lilian, et al.
Published: (2025)
Semantic Agreement Enables Efficient Open-Ended LLM Cascades
by: Soiffer, Duncan, et al.
Published: (2025)
by: Soiffer, Duncan, et al.
Published: (2025)
CYCLE: Cross-Year Contrastive Learning in Entity-Linking
by: Zhang, Pengyu, et al.
Published: (2024)
by: Zhang, Pengyu, et al.
Published: (2024)
Intuitive or Dependent? Investigating LLMs' Behavior Style to Conflicting Prompts
by: Ying, Jiahao, et al.
Published: (2023)
by: Ying, Jiahao, et al.
Published: (2023)
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
by: Le, Hieu Xuan, et al.
Published: (2026)
by: Le, Hieu Xuan, et al.
Published: (2026)
Disentangling Prompt Element Level Risk Factors for Hallucinations and Omissions in Mental Health LLM Responses
by: Ni, Congning, et al.
Published: (2026)
by: Ni, Congning, et al.
Published: (2026)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
by: Hinterleitner, Lukas, et al.
Published: (2026)
by: Hinterleitner, Lukas, et al.
Published: (2026)
StyleRec: A Benchmark Dataset for Prompt Recovery in Writing Style Transformation
by: Liu, Shenyang, et al.
Published: (2025)
by: Liu, Shenyang, et al.
Published: (2025)
Specification Overfitting in Artificial Intelligence
by: Roth, Benjamin, et al.
Published: (2024)
by: Roth, Benjamin, et al.
Published: (2024)
Process Supervision of Confidence Margin for Calibrated LLM Reasoning
by: Wang, Liaoyaqi, et al.
Published: (2026)
by: Wang, Liaoyaqi, et al.
Published: (2026)
On the Calibration of Multilingual Question Answering LLMs
by: Yang, Yahan, et al.
Published: (2023)
by: Yang, Yahan, et al.
Published: (2023)
Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales
by: Hong, Pingjun, et al.
Published: (2026)
by: Hong, Pingjun, et al.
Published: (2026)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
by: Wei, Hui, et al.
Published: (2024)
by: Wei, Hui, et al.
Published: (2024)
Enhancing LLM Agent Safety via Causal Influence Prompting
by: Hahm, Dongyoon, et al.
Published: (2025)
by: Hahm, Dongyoon, et al.
Published: (2025)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
by: Jung, Jaehun, et al.
Published: (2024)
by: Jung, Jaehun, et al.
Published: (2024)
Counterfactual LLM-based Framework for Measuring Rhetorical Style
by: Qiu, Jingyi, et al.
Published: (2025)
by: Qiu, Jingyi, et al.
Published: (2025)
Do language models accommodate their users? A study of linguistic convergence
by: Blevins, Terra, et al.
Published: (2025)
by: Blevins, Terra, et al.
Published: (2025)
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
by: Watts, Ishaan, et al.
Published: (2024)
by: Watts, Ishaan, et al.
Published: (2024)
TOAD: Task-Oriented Automatic Dialogs with Diverse Response Styles
by: Liu, Yinhong, et al.
Published: (2024)
by: Liu, Yinhong, et al.
Published: (2024)
Similar Items
-
Black-box Model Ensembling for Textual and Visual Question Answering via Information Fusion
by: Xia, Yuxi, et al.
Published: (2024) -
Analysing zero-shot temporal relation extraction on clinical notes using temporal consistency
by: Kougia, Vasiliki, et al.
Published: (2024) -
Functionality learning through specification instructions
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2023) -
Helpful assistant or fruitful facilitator? Investigating how personas affect language model behavior
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2024) -
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)