Uncertainty in Language Models: Assessment through Rank-Calibration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Xinmeng, Li, Shuo, Yu, Mengxin, Sesia, Matteo, Hassani, Hamed, Lee, Insup, Bastani, Osbert, Dobriban, Edgar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
von: Huang, Xinmeng, et al.
Veröffentlicht: (2024)
von: Huang, Xinmeng, et al.
Veröffentlicht: (2024)
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction
von: Li, Shuo, et al.
Veröffentlicht: (2023)
von: Li, Shuo, et al.
Veröffentlicht: (2023)
Evaluating the Performance of Large Language Models via Debates
von: Moniri, Behrad, et al.
Veröffentlicht: (2024)
von: Moniri, Behrad, et al.
Veröffentlicht: (2024)
Effective Reinforcement Learning for Reasoning in Language Models
von: Huang, Lianghuan, et al.
Veröffentlicht: (2025)
von: Huang, Lianghuan, et al.
Veröffentlicht: (2025)
Policy-Guided Stepwise Model Routing for Cost-Effective Reasoning
von: Si, Wenwen, et al.
Veröffentlicht: (2026)
von: Si, Wenwen, et al.
Veröffentlicht: (2026)
Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
von: Si, Wenwen, et al.
Veröffentlicht: (2025)
von: Si, Wenwen, et al.
Veröffentlicht: (2025)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
von: Anupam, Sagnik, et al.
Veröffentlicht: (2025)
von: Anupam, Sagnik, et al.
Veröffentlicht: (2025)
Evaluating the Diversity and Quality of LLM Generated Content
von: Shypula, Alexander, et al.
Veröffentlicht: (2025)
von: Shypula, Alexander, et al.
Veröffentlicht: (2025)
Optimal Multitask Linear Regression and Contextual Bandits under Sparse Heterogeneity
von: Huang, Xinmeng, et al.
Veröffentlicht: (2023)
von: Huang, Xinmeng, et al.
Veröffentlicht: (2023)
Are AI Capabilities Increasing Exponentially? A Competing Hypothesis
von: Ge, Haosen, et al.
Veröffentlicht: (2026)
von: Ge, Haosen, et al.
Veröffentlicht: (2026)
Conformal Information Pursuit for Interactively Guiding Large Language Models
von: Chan, Kwan Ho Ryan, et al.
Veröffentlicht: (2025)
von: Chan, Kwan Ho Ryan, et al.
Veröffentlicht: (2025)
A Fast, Reliable, and Secure Programming Language for LLM Agents with Code Actions
von: Mell, Stephen, et al.
Veröffentlicht: (2025)
von: Mell, Stephen, et al.
Veröffentlicht: (2025)
Watermarking Language Models with Error Correcting Codes
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
Jailbreaking Black Box Large Language Models in Twenty Queries
von: Chao, Patrick, et al.
Veröffentlicht: (2023)
von: Chao, Patrick, et al.
Veröffentlicht: (2023)
RAPID: An Efficient Reinforcement Learning Algorithm for Small Language Models
von: Huang, Lianghuan, et al.
Veröffentlicht: (2025)
von: Huang, Lianghuan, et al.
Veröffentlicht: (2025)
Conformal Inference under High-Dimensional Covariate Shifts via Likelihood-Ratio Regularization
von: Joshi, Sunay, et al.
Veröffentlicht: (2025)
von: Joshi, Sunay, et al.
Veröffentlicht: (2025)
PopPy: Opportunistically Exploiting Parallelism in Python Compound AI Applications
von: Mell, Stephen, et al.
Veröffentlicht: (2026)
von: Mell, Stephen, et al.
Veröffentlicht: (2026)
Revisiting Uncertainty Estimation and Calibration of Large Language Models
von: Tao, Linwei, et al.
Veröffentlicht: (2025)
von: Tao, Linwei, et al.
Veröffentlicht: (2025)
Diversity By Design: Leveraging Distribution Matching for Offline Model-Based Optimization
von: Yao, Michael S., et al.
Veröffentlicht: (2025)
von: Yao, Michael S., et al.
Veröffentlicht: (2025)
Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
von: Moore, Kyle, et al.
Veröffentlicht: (2025)
von: Moore, Kyle, et al.
Veröffentlicht: (2025)
Credence Calibration Game? Calibrating Large Language Models through Structured Play
von: Fang, Ke, et al.
Veröffentlicht: (2025)
von: Fang, Ke, et al.
Veröffentlicht: (2025)
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
von: Pavlovic, Maja, et al.
Veröffentlicht: (2026)
von: Pavlovic, Maja, et al.
Veröffentlicht: (2026)
SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
von: Lu, Yu-Chen, et al.
Veröffentlicht: (2025)
von: Lu, Yu-Chen, et al.
Veröffentlicht: (2025)
Uncertainty Quantification for Neurosymbolic Programs via Compositional Conformal Prediction
von: Ramalingam, Ramya, et al.
Veröffentlicht: (2024)
von: Ramalingam, Ramya, et al.
Veröffentlicht: (2024)
Statistical Methods in Generative AI
von: Dobriban, Edgar
Veröffentlicht: (2025)
von: Dobriban, Edgar
Veröffentlicht: (2025)
Decaf: Improving Neural Decompilation with Automatic Feedback and Search
von: Shypula, Alexander, et al.
Veröffentlicht: (2026)
von: Shypula, Alexander, et al.
Veröffentlicht: (2026)
Solving a Research Problem in Mathematical Statistics with AI Assistance
von: Dobriban, Edgar
Veröffentlicht: (2025)
von: Dobriban, Edgar
Veröffentlicht: (2025)
Generative Adversarial Model-Based Optimization via Source Critic Regularization
von: Yao, Michael S., et al.
Veröffentlicht: (2024)
von: Yao, Michael S., et al.
Veröffentlicht: (2024)
Foundations of Top-$k$ Decoding For Language Models
von: Noarov, Georgy, et al.
Veröffentlicht: (2025)
von: Noarov, Georgy, et al.
Veröffentlicht: (2025)
Detecting Safety Violations Across Many Agent Traces
von: Stein, Adam, et al.
Veröffentlicht: (2026)
von: Stein, Adam, et al.
Veröffentlicht: (2026)
Calibrating Reasoning in Language Models with Internal Consistency
von: Xie, Zhihui, et al.
Veröffentlicht: (2024)
von: Xie, Zhihui, et al.
Veröffentlicht: (2024)
Knowledgeable Language Models as Black-Box Optimizers for Personalized Medicine
von: Yao, Michael S., et al.
Veröffentlicht: (2025)
von: Yao, Michael S., et al.
Veröffentlicht: (2025)
Detoxification of Large Language Models through Output-layer Fusion with a Calibration Model
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
von: Krishnan, Ranganath, et al.
Veröffentlicht: (2024)
von: Krishnan, Ranganath, et al.
Veröffentlicht: (2024)
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation
von: Wang, Ziyu, et al.
Veröffentlicht: (2024)
von: Wang, Ziyu, et al.
Veröffentlicht: (2024)
How Confident Is the First Token? An Uncertainty-Calibrated Prompt Optimization Framework for Large Language Model Classification and Understanding
von: Chen, Wei, et al.
Veröffentlicht: (2026)
von: Chen, Wei, et al.
Veröffentlicht: (2026)
Uncertainty Quantification of Large Language Models through Multi-Dimensional Responses
von: Chen, Tiejin, et al.
Veröffentlicht: (2025)
von: Chen, Tiejin, et al.
Veröffentlicht: (2025)
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
von: Park, Jean, et al.
Veröffentlicht: (2024)
von: Park, Jean, et al.
Veröffentlicht: (2024)
Asymptotic Normality of Generalized Low-Rank Matrix Sensing via Riemannian Geometry
von: Bastani, Osbert
Veröffentlicht: (2024)
von: Bastani, Osbert
Veröffentlicht: (2024)
Uncertainty Estimation of Large Language Models in Medical Question Answering
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
von: Huang, Xinmeng, et al.
Veröffentlicht: (2024) -
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction
von: Li, Shuo, et al.
Veröffentlicht: (2023) -
Evaluating the Performance of Large Language Models via Debates
von: Moniri, Behrad, et al.
Veröffentlicht: (2024) -
Effective Reinforcement Learning for Reasoning in Language Models
von: Huang, Lianghuan, et al.
Veröffentlicht: (2025) -
Policy-Guided Stepwise Model Routing for Cost-Effective Reasoning
von: Si, Wenwen, et al.
Veröffentlicht: (2026)