Language Model Probabilities are Not Calibrated in Numeric Contexts
Fuente:
arXiv
Salvato in:
| Autori principali: | Lovering, Charles, Krumdick, Michael, Lai, Viet Dac, Ebner, Seth, Kumar, Nilesh, Reddy, Varshini, Koncel-Kedziorski, Rik, Tanner, Chris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DocFinQA: A Long-Context Financial Reasoning Dataset
di: Reddy, Varshini, et al.
Pubblicazione: (2024)
di: Reddy, Varshini, et al.
Pubblicazione: (2024)
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
di: Koncel-Kedziorski, Rik, et al.
Pubblicazione: (2023)
di: Koncel-Kedziorski, Rik, et al.
Pubblicazione: (2023)
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
di: Krumdick, Michael, et al.
Pubblicazione: (2025)
di: Krumdick, Michael, et al.
Pubblicazione: (2025)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
di: Lai, Viet Dac, et al.
Pubblicazione: (2024)
di: Lai, Viet Dac, et al.
Pubblicazione: (2024)
Cost-Efficient Estimation of General Abilities Across Benchmarks
di: Krumdick, Michael, et al.
Pubblicazione: (2026)
di: Krumdick, Michael, et al.
Pubblicazione: (2026)
On Finding Inconsistencies in Documents
di: Lovering, Charles J., et al.
Pubblicazione: (2025)
di: Lovering, Charles J., et al.
Pubblicazione: (2025)
An Analysis of Multilingual FActScore
di: Vu, Kim Trong, et al.
Pubblicazione: (2024)
di: Vu, Kim Trong, et al.
Pubblicazione: (2024)
Tokenization with Split Trees
di: Schmidt, Craig W., et al.
Pubblicazione: (2026)
di: Schmidt, Craig W., et al.
Pubblicazione: (2026)
PrimeX: A Dataset of Worldview, Opinion, and Explanation
di: Koncel-Kedziorski, Rik, et al.
Pubblicazione: (2025)
di: Koncel-Kedziorski, Rik, et al.
Pubblicazione: (2025)
Improving Language Model Personas via Rationalization with Psychological Scaffolds
di: Joshi, Brihi, et al.
Pubblicazione: (2025)
di: Joshi, Brihi, et al.
Pubblicazione: (2025)
The Effect of Scripts and Formats on LLM Numeracy
di: Reddy, Varshini, et al.
Pubblicazione: (2026)
di: Reddy, Varshini, et al.
Pubblicazione: (2026)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
di: Schmidt, Craig W., et al.
Pubblicazione: (2025)
di: Schmidt, Craig W., et al.
Pubblicazione: (2025)
Not How Many, But Which: Parameter Placement in Low-Rank Adaptation
di: Sehanobish, Arijit, et al.
Pubblicazione: (2026)
di: Sehanobish, Arijit, et al.
Pubblicazione: (2026)
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Partial Policy Gradients for RL in LLMs
di: Mathur, Puneet, et al.
Pubblicazione: (2026)
di: Mathur, Puneet, et al.
Pubblicazione: (2026)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
di: Chang, Yapei, et al.
Pubblicazione: (2025)
di: Chang, Yapei, et al.
Pubblicazione: (2025)
Drift No More? Context Equilibria in Multi-Turn LLM Interactions
di: Dongre, Vardhan, et al.
Pubblicazione: (2025)
di: Dongre, Vardhan, et al.
Pubblicazione: (2025)
Tokenization Is More Than Compression
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
di: Reddy, Varshini, et al.
Pubblicazione: (2025)
di: Reddy, Varshini, et al.
Pubblicazione: (2025)
Calibrating Verbalized Probabilities for Large Language Models
di: Wang, Cheng, et al.
Pubblicazione: (2024)
di: Wang, Cheng, et al.
Pubblicazione: (2024)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
di: Dongre, Vardhan, et al.
Pubblicazione: (2026)
di: Dongre, Vardhan, et al.
Pubblicazione: (2026)
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
di: Krumdick, Michael, et al.
Pubblicazione: (2026)
di: Krumdick, Michael, et al.
Pubblicazione: (2026)
In-Context Unlearning: Language Models as Few Shot Unlearners
di: Pawelczyk, Martin, et al.
Pubblicazione: (2023)
di: Pawelczyk, Martin, et al.
Pubblicazione: (2023)
Considering Fundamental Rights in the European Standardisation of Artificial Intelligence: Nonsense or Strategic Alliance?
di: Ho-Dac, Marion
Pubblicazione: (2024)
di: Ho-Dac, Marion
Pubblicazione: (2024)
First Analysis of the EU Artifical Intelligence Act: Towards a Global Standard for Trustworthy AI?
di: Ho-Dac, Marion
Pubblicazione: (2024)
di: Ho-Dac, Marion
Pubblicazione: (2024)
Post-Training Statistical Calibration for Higher Activation Sparsity
di: Chua, Vui Seng, et al.
Pubblicazione: (2024)
di: Chua, Vui Seng, et al.
Pubblicazione: (2024)
DySTAN: Joint Modeling of Sedentary Activity and Social Context from Smartphone Sensors
di: Sneh, Aditya, et al.
Pubblicazione: (2025)
di: Sneh, Aditya, et al.
Pubblicazione: (2025)
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
di: Taesiri, Mohammad Reza, et al.
Pubblicazione: (2025)
di: Taesiri, Mohammad Reza, et al.
Pubblicazione: (2025)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
di: Van Nguyen, Chien, et al.
Pubblicazione: (2024)
di: Van Nguyen, Chien, et al.
Pubblicazione: (2024)
Ensembling Tabular Foundation Models - A Diversity Ceiling And A Calibration Trap
di: Tanna, Aditya, et al.
Pubblicazione: (2026)
di: Tanna, Aditya, et al.
Pubblicazione: (2026)
Text-to-SQL Calibration: No Need to Ask -- Just Rescale Model Probabilities
di: Ramachandran, Ashwin, et al.
Pubblicazione: (2024)
di: Ramachandran, Ashwin, et al.
Pubblicazione: (2024)
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation
di: Wang, Ziyu, et al.
Pubblicazione: (2024)
di: Wang, Ziyu, et al.
Pubblicazione: (2024)
Confidence Calibration in Large Language Models
di: Michael, Noam, et al.
Pubblicazione: (2026)
di: Michael, Noam, et al.
Pubblicazione: (2026)
PSC: Extending Context Window of Large Language Models via Phase Shift Calibration
di: Zhu, Wenqiao, et al.
Pubblicazione: (2025)
di: Zhu, Wenqiao, et al.
Pubblicazione: (2025)
SlimLM: An Efficient Small Language Model for On-Device Document Assistance
di: Pham, Thang M., et al.
Pubblicazione: (2024)
di: Pham, Thang M., et al.
Pubblicazione: (2024)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
di: Hu, Yifan, et al.
Pubblicazione: (2025)
di: Hu, Yifan, et al.
Pubblicazione: (2025)
Frontier AI Ethics: Anticipating and Evaluating the Societal Impacts of Language Model Agents
di: Lazar, Seth
Pubblicazione: (2024)
di: Lazar, Seth
Pubblicazione: (2024)
Orion-Bix: Bi-Axial Attention for Tabular In-Context Learning
di: Bouadi, Mohamed, et al.
Pubblicazione: (2025)
di: Bouadi, Mohamed, et al.
Pubblicazione: (2025)
Privacy Issues in Large Language Models: A Survey
di: Neel, Seth, et al.
Pubblicazione: (2023)
di: Neel, Seth, et al.
Pubblicazione: (2023)
Documenti analoghi
-
DocFinQA: A Long-Context Financial Reasoning Dataset
di: Reddy, Varshini, et al.
Pubblicazione: (2024) -
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
di: Koncel-Kedziorski, Rik, et al.
Pubblicazione: (2023) -
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
di: Krumdick, Michael, et al.
Pubblicazione: (2025) -
SEC-QA: A Systematic Evaluation Corpus for Financial QA
di: Lai, Viet Dac, et al.
Pubblicazione: (2024) -
Cost-Efficient Estimation of General Abilities Across Benchmarks
di: Krumdick, Michael, et al.
Pubblicazione: (2026)