Calibrating Expressions of Certainty
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Peiqi, Lam, Barbara D., Liu, Yingcheng, Asgari-Targhi, Ameneh, Panda, Rameswar, Wells, William M., Kapur, Tina, Golland, Polina |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Diversity Measurement and Subset Selection for Instruction Tuning Datasets
di: Wang, Peiqi, et al.
Pubblicazione: (2024)
di: Wang, Peiqi, et al.
Pubblicazione: (2024)
Connecting Jensen-Shannon and Kullback-Leibler Divergences: A New Bound for Representation Learning
di: Dorent, Reuben, et al.
Pubblicazione: (2025)
di: Dorent, Reuben, et al.
Pubblicazione: (2025)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
di: Brandon, William, et al.
Pubblicazione: (2024)
di: Brandon, William, et al.
Pubblicazione: (2024)
Gated Linear Attention Transformers with Hardware-Efficient Training
di: Yang, Songlin, et al.
Pubblicazione: (2023)
di: Yang, Songlin, et al.
Pubblicazione: (2023)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
di: Nrusimha, Aniruddha, et al.
Pubblicazione: (2024)
di: Nrusimha, Aniruddha, et al.
Pubblicazione: (2024)
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
di: Nrusimha, Aniruddha, et al.
Pubblicazione: (2025)
di: Nrusimha, Aniruddha, et al.
Pubblicazione: (2025)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
di: Tan, Shawn, et al.
Pubblicazione: (2024)
di: Tan, Shawn, et al.
Pubblicazione: (2024)
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
di: Guo, Zhen, et al.
Pubblicazione: (2024)
di: Guo, Zhen, et al.
Pubblicazione: (2024)
The Confidence Trap: Gender Bias and Predictive Certainty in LLMs
di: Sabir, Ahmed, et al.
Pubblicazione: (2026)
di: Sabir, Ahmed, et al.
Pubblicazione: (2026)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
di: Yang, Songlin, et al.
Pubblicazione: (2025)
di: Yang, Songlin, et al.
Pubblicazione: (2025)
The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
di: Tomov, Tim, et al.
Pubblicazione: (2025)
di: Tomov, Tim, et al.
Pubblicazione: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
di: Pan, Bowen, et al.
Pubblicazione: (2024)
di: Pan, Bowen, et al.
Pubblicazione: (2024)
Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
di: Kang, Junmo, et al.
Pubblicazione: (2024)
di: Kang, Junmo, et al.
Pubblicazione: (2024)
TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
di: Xu, Zhangchen, et al.
Pubblicazione: (2025)
di: Xu, Zhangchen, et al.
Pubblicazione: (2025)
Causality and Scientific Inquiry: Lessons from Space Physics and Medical Sciences
di: Asgari-Targhi, Marzieh, et al.
Pubblicazione: (2026)
di: Asgari-Targhi, Marzieh, et al.
Pubblicazione: (2026)
Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
di: Taghanaki, Saeid Asgari, et al.
Pubblicazione: (2025)
di: Taghanaki, Saeid Asgari, et al.
Pubblicazione: (2025)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
di: Jain, Neel, et al.
Pubblicazione: (2024)
di: Jain, Neel, et al.
Pubblicazione: (2024)
Process Supervision of Confidence Margin for Calibrated LLM Reasoning
di: Wang, Liaoyaqi, et al.
Pubblicazione: (2026)
di: Wang, Liaoyaqi, et al.
Pubblicazione: (2026)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
di: Shen, Yikang, et al.
Pubblicazione: (2024)
di: Shen, Yikang, et al.
Pubblicazione: (2024)
Beyond Simple Averaging: Improving NLP Ensemble Performance with Topological-Data-Analysis-Based Weighting
di: Proskura, Polina, et al.
Pubblicazione: (2024)
di: Proskura, Polina, et al.
Pubblicazione: (2024)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
di: Kang, Zhewei, et al.
Pubblicazione: (2025)
di: Kang, Zhewei, et al.
Pubblicazione: (2025)
Efficient Post-Training Pruning of Large Language Models with Statistical Correction
di: Yu, Peiqi, et al.
Pubblicazione: (2026)
di: Yu, Peiqi, et al.
Pubblicazione: (2026)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
di: Taghanaki, Saeid Asgari, et al.
Pubblicazione: (2024)
di: Taghanaki, Saeid Asgari, et al.
Pubblicazione: (2024)
PRISM: Demystifying Retention and Interaction in Mid-Training
di: Runwal, Bharat, et al.
Pubblicazione: (2026)
di: Runwal, Bharat, et al.
Pubblicazione: (2026)
Less is More: Rethinking Few-Shot Learning and Recurrent Neural Nets
di: Pereg, Deborah, et al.
Pubblicazione: (2022)
di: Pereg, Deborah, et al.
Pubblicazione: (2022)
Unified Cross-Modal Medical Image Synthesis with Hierarchical Mixture of Product-of-Experts
di: Dorent, Reuben, et al.
Pubblicazione: (2024)
di: Dorent, Reuben, et al.
Pubblicazione: (2024)
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
di: Huang, Jiameng, et al.
Pubblicazione: (2025)
di: Huang, Jiameng, et al.
Pubblicazione: (2025)
The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
di: Zhang, Jiaxin, et al.
Pubblicazione: (2026)
di: Zhang, Jiaxin, et al.
Pubblicazione: (2026)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
di: Nakkiran, Preetum, et al.
Pubblicazione: (2025)
di: Nakkiran, Preetum, et al.
Pubblicazione: (2025)
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
di: Daheim, Nico, et al.
Pubblicazione: (2024)
di: Daheim, Nico, et al.
Pubblicazione: (2024)
Scattered Mixture-of-Experts Implementation
di: Tan, Shawn, et al.
Pubblicazione: (2024)
di: Tan, Shawn, et al.
Pubblicazione: (2024)
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
di: Zhou, Xuwen, et al.
Pubblicazione: (2026)
di: Zhou, Xuwen, et al.
Pubblicazione: (2026)
Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
di: Jiang, Zhengping, et al.
Pubblicazione: (2025)
di: Jiang, Zhengping, et al.
Pubblicazione: (2025)
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment
di: Karim, Ahmed, et al.
Pubblicazione: (2025)
di: Karim, Ahmed, et al.
Pubblicazione: (2025)
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings
di: Hallee, Logan, et al.
Pubblicazione: (2024)
di: Hallee, Logan, et al.
Pubblicazione: (2024)
Fetuses Made Simple: Modeling and Tracking of Fetal Shape and Pose
di: Liu, Yingcheng, et al.
Pubblicazione: (2025)
di: Liu, Yingcheng, et al.
Pubblicazione: (2025)
From the Inside Out: Progressive Distribution Refinement for Confidence Calibration
di: Yang, Xizhong, et al.
Pubblicazione: (2026)
di: Yang, Xizhong, et al.
Pubblicazione: (2026)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
di: Liu, Hongfu, et al.
Pubblicazione: (2024)
di: Liu, Hongfu, et al.
Pubblicazione: (2024)
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
di: Behnia, Tina, et al.
Pubblicazione: (2025)
di: Behnia, Tina, et al.
Pubblicazione: (2025)
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2025)
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Diversity Measurement and Subset Selection for Instruction Tuning Datasets
di: Wang, Peiqi, et al.
Pubblicazione: (2024) -
Connecting Jensen-Shannon and Kullback-Leibler Divergences: A New Bound for Representation Learning
di: Dorent, Reuben, et al.
Pubblicazione: (2025) -
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
di: Brandon, William, et al.
Pubblicazione: (2024) -
Gated Linear Attention Transformers with Hardware-Efficient Training
di: Yang, Songlin, et al.
Pubblicazione: (2023) -
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
di: Nrusimha, Aniruddha, et al.
Pubblicazione: (2024)