Probabilistic Calibration Is a Trainable Capability in Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Baldelli, Davide, Kuriakose, Sruthi, Hashemzadeh, Maryam, Zouaq, Amal, Chandar, Sarath |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
di: Baldelli, Davide, et al.
Pubblicazione: (2026)
di: Baldelli, Davide, et al.
Pubblicazione: (2026)
A Deep Dive into the Trade-Offs of Parameter-Efficient Preference Alignment Techniques
di: Thakkar, Megh, et al.
Pubblicazione: (2024)
di: Thakkar, Megh, et al.
Pubblicazione: (2024)
What is the Best Process Model Representation? A Comparative Analysis for Process Modeling with Large Language Models
di: Brissard, Alexis, et al.
Pubblicazione: (2025)
di: Brissard, Alexis, et al.
Pubblicazione: (2025)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Reducing Hallucinations in Language Model-based SPARQL Query Generation Using Post-Generation Memory Retrieval
di: Sharma, Aditya, et al.
Pubblicazione: (2025)
di: Sharma, Aditya, et al.
Pubblicazione: (2025)
Faithfulness Measurable Masked Language Models
di: Madsen, Andreas, et al.
Pubblicazione: (2023)
di: Madsen, Andreas, et al.
Pubblicazione: (2023)
Ontology-Constrained Generation of Domain-Specific Clinical Summaries
di: Mehenni, Gaya, et al.
Pubblicazione: (2024)
di: Mehenni, Gaya, et al.
Pubblicazione: (2024)
CADmium: Fine-Tuning Code Language Models for Text-Driven Sequential CAD Design
di: Govindarajan, Prashant, et al.
Pubblicazione: (2025)
di: Govindarajan, Prashant, et al.
Pubblicazione: (2025)
Are self-explanations from Large Language Models faithful?
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
Investigating the Multilingual Calibration Effects of Language Model Instruction-Tuning
di: Huang, Jerry, et al.
Pubblicazione: (2026)
di: Huang, Jerry, et al.
Pubblicazione: (2026)
A Comprehensive Evaluation of Neural SPARQL Query Generation from Natural Language Questions
di: Diallo, Papa Abdou Karim Karou, et al.
Pubblicazione: (2023)
di: Diallo, Papa Abdou Karim Karou, et al.
Pubblicazione: (2023)
FRASE: Structured Representations for Generalizable SPARQL Query Generation
di: Diallo, Papa Abdou Karim Karou, et al.
Pubblicazione: (2025)
di: Diallo, Papa Abdou Karim Karou, et al.
Pubblicazione: (2025)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
di: Prato, Gabriele, et al.
Pubblicazione: (2023)
di: Prato, Gabriele, et al.
Pubblicazione: (2023)
GeoCoder: Solving Geometry Problems by Generating Modular Code through Vision-Language Models
di: Sharma, Aditya, et al.
Pubblicazione: (2024)
di: Sharma, Aditya, et al.
Pubblicazione: (2024)
Sub-goal Distillation: A Method to Improve Small Language Agents
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2024)
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2024)
Enhancing Frame Detection with Retrieval Augmented Generation
di: Diallo, Papa Abdou Karim Karou, et al.
Pubblicazione: (2025)
di: Diallo, Papa Abdou Karim Karou, et al.
Pubblicazione: (2025)
Do Large Language Models Know How Much They Know?
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Interpretability Needs a New Paradigm
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
DeSQ: Decomposition-based SPARQL Query Generation
di: Diallo, Papa Abdou Karim Karou, et al.
Pubblicazione: (2026)
di: Diallo, Papa Abdou Karim Karou, et al.
Pubblicazione: (2026)
Why Don't Prompt-Based Fairness Metrics Correlate?
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
What Are the Odds? Language Models Are Capable of Probabilistic Reasoning
di: Paruchuri, Akshay, et al.
Pubblicazione: (2024)
di: Paruchuri, Akshay, et al.
Pubblicazione: (2024)
Towards Practical Tool Usage for Continually Learning LLMs
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Too Big to Fool: Resisting Deception in Language Models
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2026)
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2026)
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
di: Pihlakas, Roland, et al.
Pubblicazione: (2025)
di: Pihlakas, Roland, et al.
Pubblicazione: (2025)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
di: Bhan, Milan, et al.
Pubblicazione: (2025)
di: Bhan, Milan, et al.
Pubblicazione: (2025)
MedHal: An Evaluation Dataset for Medical Hallucination Detection
di: Mehenni, Gaya, et al.
Pubblicazione: (2025)
di: Mehenni, Gaya, et al.
Pubblicazione: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Reasoning in Large Language Models: A Geometric Perspective
di: Cosentino, Romain, et al.
Pubblicazione: (2024)
di: Cosentino, Romain, et al.
Pubblicazione: (2024)
On Calibration of Large Language Models: From Response To Capability
di: Yang, Sin-Han, et al.
Pubblicazione: (2026)
di: Yang, Sin-Han, et al.
Pubblicazione: (2026)
Punctuation-aware Hybrid Trainable Sparse Attention for Large Language Models
di: Qiu, Junxiang, et al.
Pubblicazione: (2026)
di: Qiu, Junxiang, et al.
Pubblicazione: (2026)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
di: Abbes, Istabrak, et al.
Pubblicazione: (2025)
di: Abbes, Istabrak, et al.
Pubblicazione: (2025)
NeoBERT: A Next-Generation BERT
di: Breton, Lola Le, et al.
Pubblicazione: (2025)
di: Breton, Lola Le, et al.
Pubblicazione: (2025)
ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure
di: Yoon, Hee Suk, et al.
Pubblicazione: (2023)
di: Yoon, Hee Suk, et al.
Pubblicazione: (2023)
Fact-Checking with Large Language Models via Probabilistic Certainty and Consistency
di: Wang, Haoran, et al.
Pubblicazione: (2026)
di: Wang, Haoran, et al.
Pubblicazione: (2026)
Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs
di: Thakkar, Megh, et al.
Pubblicazione: (2024)
di: Thakkar, Megh, et al.
Pubblicazione: (2024)
Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
di: Chebbi, Amal, et al.
Pubblicazione: (2025)
di: Chebbi, Amal, et al.
Pubblicazione: (2025)
Trainable Transformer in Transformer
di: Panigrahi, Abhishek, et al.
Pubblicazione: (2023)
di: Panigrahi, Abhishek, et al.
Pubblicazione: (2023)
Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes
di: Bochkov, A.
Pubblicazione: (2026)
di: Bochkov, A.
Pubblicazione: (2026)
Documenti analoghi
-
LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
di: Baldelli, Davide, et al.
Pubblicazione: (2026) -
A Deep Dive into the Trade-Offs of Parameter-Efficient Preference Alignment Techniques
di: Thakkar, Megh, et al.
Pubblicazione: (2024) -
What is the Best Process Model Representation? A Comparative Analysis for Process Modeling with Large Language Models
di: Brissard, Alexis, et al.
Pubblicazione: (2025) -
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
di: Prato, Gabriele, et al.
Pubblicazione: (2025) -
Reducing Hallucinations in Language Model-based SPARQL Query Generation Using Post-Generation Memory Retrieval
di: Sharma, Aditya, et al.
Pubblicazione: (2025)