Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Liaoyaqi, Jiang, Zhengping, Liu, Anqi, Van Durme, Benjamin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
von: Jiang, Zhengping, et al.
Veröffentlicht: (2025)
von: Jiang, Zhengping, et al.
Veröffentlicht: (2025)
Process Supervision of Confidence Margin for Calibrated LLM Reasoning
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026)
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026)
SEQR: Secure and Efficient QR-based LoRA Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
SpectR: Dynamically Composing LM Experts with Spectral Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
RE-Adapt: Reverse Engineered Adaptation of Large Language Models
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
Perplexity Cannot Always Tell Right from Wrong
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
von: Subramani, Nishant, et al.
Veröffentlicht: (2025)
von: Subramani, Nishant, et al.
Veröffentlicht: (2025)
Do Androids Know They're Only Dreaming of Electric Sheep?
von: CH-Wang, Sky, et al.
Veröffentlicht: (2023)
von: CH-Wang, Sky, et al.
Veröffentlicht: (2023)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
von: Wang, Boshi, et al.
Veröffentlicht: (2024)
von: Wang, Boshi, et al.
Veröffentlicht: (2024)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time
von: Hu, Michael Y., et al.
Veröffentlicht: (2026)
von: Hu, Michael Y., et al.
Veröffentlicht: (2026)
Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
von: Chen, Tong, et al.
Veröffentlicht: (2024)
von: Chen, Tong, et al.
Veröffentlicht: (2024)
DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
MIKE: A New Benchmark for Fine-grained Multimodal Entity Knowledge Editing
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning
von: Liu, Yong, et al.
Veröffentlicht: (2024)
von: Liu, Yong, et al.
Veröffentlicht: (2024)
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
Show Me How It's Done: The Role of Explanations in Fine-Tuning Language Models
von: Ballout, Mohamad, et al.
Veröffentlicht: (2024)
von: Ballout, Mohamad, et al.
Veröffentlicht: (2024)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
Aligning Large Language Models via Fine-grained Supervision
von: Xu, Dehong, et al.
Veröffentlicht: (2024)
von: Xu, Dehong, et al.
Veröffentlicht: (2024)
Reasoning Models Don't Always Say What They Think
von: Chen, Yanda, et al.
Veröffentlicht: (2025)
von: Chen, Yanda, et al.
Veröffentlicht: (2025)
PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
von: Guo, Kevin H., et al.
Veröffentlicht: (2026)
von: Guo, Kevin H., et al.
Veröffentlicht: (2026)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
von: Nikitin, Alexander, et al.
Veröffentlicht: (2024)
von: Nikitin, Alexander, et al.
Veröffentlicht: (2024)
Mini-Giants: "Small" Language Models and Open Source Win-Win
von: Zhou, Zhengping, et al.
Veröffentlicht: (2023)
von: Zhou, Zhengping, et al.
Veröffentlicht: (2023)
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
von: Xia, Yuchen, et al.
Veröffentlicht: (2024)
von: Xia, Yuchen, et al.
Veröffentlicht: (2024)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
von: Weng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Weng, Jiaqi, et al.
Veröffentlicht: (2025)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
von: Kim, Taeho, et al.
Veröffentlicht: (2024)
von: Kim, Taeho, et al.
Veröffentlicht: (2024)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
von: Yoon, Junsang, et al.
Veröffentlicht: (2024)
von: Yoon, Junsang, et al.
Veröffentlicht: (2024)
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
Proximal Supervised Fine-Tuning
von: Zhu, Wenhong, et al.
Veröffentlicht: (2025)
von: Zhu, Wenhong, et al.
Veröffentlicht: (2025)
MeTHanol: Modularized Thinking Language Models with Intermediate Layer Thinking, Decoding and Bootstrapping Reasoning
von: Xi, Ningyuan, et al.
Veröffentlicht: (2024)
von: Xi, Ningyuan, et al.
Veröffentlicht: (2024)
MeMo: Memory as a Model
von: Quek, Ryan Wei Heng, et al.
Veröffentlicht: (2026)
von: Quek, Ryan Wei Heng, et al.
Veröffentlicht: (2026)
When Numbers Tell Half the Story: Human-Metric Alignment in Topic Model Evaluation
von: Prouteau, Thibault, et al.
Veröffentlicht: (2026)
von: Prouteau, Thibault, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
von: Jiang, Zhengping, et al.
Veröffentlicht: (2025) -
Process Supervision of Confidence Margin for Calibrated LLM Reasoning
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026) -
SEQR: Secure and Efficient QR-based LoRA Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025) -
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
von: Fleshman, William, et al.
Veröffentlicht: (2025) -
SpectR: Dynamically Composing LM Experts with Spectral Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)