Calibrating Large Language Models Using Their Generations Only
Fuente:
arXiv
Salvato in:
| Autori principali: | Ulmer, Dennis, Gubri, Martin, Lee, Hwaran, Yun, Sangdoo, Oh, Seong Joon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
di: Gubri, Martin, et al.
Pubblicazione: (2024)
di: Gubri, Martin, et al.
Pubblicazione: (2024)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
di: Puerto, Haritz, et al.
Pubblicazione: (2024)
di: Puerto, Haritz, et al.
Pubblicazione: (2024)
Dr.LLM: Dynamic Layer Routing in LLMs
di: Heakl, Ahmed, et al.
Pubblicazione: (2025)
di: Heakl, Ahmed, et al.
Pubblicazione: (2025)
MASEval: Extending Multi-Agent Evaluation from Models to Systems
di: Emde, Cornelius, et al.
Pubblicazione: (2026)
di: Emde, Cornelius, et al.
Pubblicazione: (2026)
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
di: Green, Tommaso, et al.
Pubblicazione: (2025)
di: Green, Tommaso, et al.
Pubblicazione: (2025)
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
di: Yoo, Haneul, et al.
Pubblicazione: (2024)
di: Yoo, Haneul, et al.
Pubblicazione: (2024)
C-SEO Bench: Does Conversational SEO Work?
di: Puerto, Haritz, et al.
Pubblicazione: (2025)
di: Puerto, Haritz, et al.
Pubblicazione: (2025)
On Uncertainty In Natural Language Processing
di: Ulmer, Dennis
Pubblicazione: (2024)
di: Ulmer, Dennis
Pubblicazione: (2024)
DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
di: Rubinstein, Alexander, et al.
Pubblicazione: (2025)
di: Rubinstein, Alexander, et al.
Pubblicazione: (2025)
Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models
di: Goel, Anmol, et al.
Pubblicazione: (2026)
di: Goel, Anmol, et al.
Pubblicazione: (2026)
Non-Exchangeable Conformal Language Generation with Nearest Neighbors
di: Ulmer, Dennis, et al.
Pubblicazione: (2024)
di: Ulmer, Dennis, et al.
Pubblicazione: (2024)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
di: Kyriakou, Athina, et al.
Pubblicazione: (2026)
di: Kyriakou, Athina, et al.
Pubblicazione: (2026)
CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
di: Shim, Jung-Woo, et al.
Pubblicazione: (2025)
di: Shim, Jung-Woo, et al.
Pubblicazione: (2025)
Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
di: Shim, Jung-Woo, et al.
Pubblicazione: (2025)
di: Shim, Jung-Woo, et al.
Pubblicazione: (2025)
On Calibration of Large Language Models: From Response To Capability
di: Yang, Sin-Han, et al.
Pubblicazione: (2026)
di: Yang, Sin-Han, et al.
Pubblicazione: (2026)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
di: Lee, Isack, et al.
Pubblicazione: (2024)
di: Lee, Isack, et al.
Pubblicazione: (2024)
COPAL: Continual Pruning in Large Language Generative Models
di: Malla, Srikanth, et al.
Pubblicazione: (2024)
di: Malla, Srikanth, et al.
Pubblicazione: (2024)
AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
di: Lee, Sangjun, et al.
Pubblicazione: (2025)
di: Lee, Sangjun, et al.
Pubblicazione: (2025)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
di: Do, Heejin, et al.
Pubblicazione: (2025)
di: Do, Heejin, et al.
Pubblicazione: (2025)
Calibrating Long-form Generations from Large Language Models
di: Huang, Yukun, et al.
Pubblicazione: (2024)
di: Huang, Yukun, et al.
Pubblicazione: (2024)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
di: Chuang, Yung-Sung, et al.
Pubblicazione: (2024)
di: Chuang, Yung-Sung, et al.
Pubblicazione: (2024)
Explicit Diversity Conditions for Effective Question Answer Generation with Large Language Models
di: Yadav, Vikas, et al.
Pubblicazione: (2024)
di: Yadav, Vikas, et al.
Pubblicazione: (2024)
MEME: Multi-entity & Evolving Memory Evaluation
di: Jung, Seokwon, et al.
Pubblicazione: (2026)
di: Jung, Seokwon, et al.
Pubblicazione: (2026)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
di: Zhang, Qingru, et al.
Pubblicazione: (2025)
di: Zhang, Qingru, et al.
Pubblicazione: (2025)
LLM generation novelty through the lens of semantic similarity
di: Davydov, Philipp, et al.
Pubblicazione: (2025)
di: Davydov, Philipp, et al.
Pubblicazione: (2025)
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
di: Lee, Gihun, et al.
Pubblicazione: (2024)
di: Lee, Gihun, et al.
Pubblicazione: (2024)
Beware of Calibration Data for Pruning Large Language Models
di: Ji, Yixin, et al.
Pubblicazione: (2024)
di: Ji, Yixin, et al.
Pubblicazione: (2024)
Revisiting Uncertainty Estimation and Calibration of Large Language Models
di: Tao, Linwei, et al.
Pubblicazione: (2025)
di: Tao, Linwei, et al.
Pubblicazione: (2025)
Calibrating Language Models with Adaptive Temperature Scaling
di: Xie, Johnathan, et al.
Pubblicazione: (2024)
di: Xie, Johnathan, et al.
Pubblicazione: (2024)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
di: Hu, Xing, et al.
Pubblicazione: (2024)
di: Hu, Xing, et al.
Pubblicazione: (2024)
Uncertainty in Language Models: Assessment through Rank-Calibration
di: Huang, Xinmeng, et al.
Pubblicazione: (2024)
di: Huang, Xinmeng, et al.
Pubblicazione: (2024)
On the Entropy Calibration of Language Models
di: Cao, Steven, et al.
Pubblicazione: (2025)
di: Cao, Steven, et al.
Pubblicazione: (2025)
Query-Conditioned Test-Time Self-Training for Large Language Models
di: Song, Chaehee, et al.
Pubblicazione: (2026)
di: Song, Chaehee, et al.
Pubblicazione: (2026)
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
di: Krishnan, Ranganath, et al.
Pubblicazione: (2024)
di: Krishnan, Ranganath, et al.
Pubblicazione: (2024)
Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation
di: Mohamed, Asim, et al.
Pubblicazione: (2025)
di: Mohamed, Asim, et al.
Pubblicazione: (2025)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
di: Kirchhof, Michael, et al.
Pubblicazione: (2025)
di: Kirchhof, Michael, et al.
Pubblicazione: (2025)
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation
di: Wang, Ziyu, et al.
Pubblicazione: (2024)
di: Wang, Ziyu, et al.
Pubblicazione: (2024)
Personalizing Large Language Models using Retrieval Augmented Generation and Knowledge Graph
di: Prahlad, Deeksha, et al.
Pubblicazione: (2025)
di: Prahlad, Deeksha, et al.
Pubblicazione: (2025)
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT
di: Shakil, Hassan, et al.
Pubblicazione: (2024)
di: Shakil, Hassan, et al.
Pubblicazione: (2024)
Graph-based Confidence Calibration for Large Language Models
di: Li, Yukun, et al.
Pubblicazione: (2024)
di: Li, Yukun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
di: Gubri, Martin, et al.
Pubblicazione: (2024) -
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
di: Puerto, Haritz, et al.
Pubblicazione: (2024) -
Dr.LLM: Dynamic Layer Routing in LLMs
di: Heakl, Ahmed, et al.
Pubblicazione: (2025) -
MASEval: Extending Multi-Agent Evaluation from Models to Systems
di: Emde, Cornelius, et al.
Pubblicazione: (2026) -
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
di: Green, Tommaso, et al.
Pubblicazione: (2025)