Capturing LLM Capabilities via Evidence-Calibrated Query Clustering
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Fangzhou, Silwal, Sandeep, Zhang, Qiuyi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Randomization Boosts KV Caching, Learning Balances Query Load: A Joint Perspective
di: Wu, Fangzhou, et al.
Pubblicazione: (2026)
di: Wu, Fangzhou, et al.
Pubblicazione: (2026)
DynMuon: A Dynamic Spectral Shaping View of Muon
di: Wu, Fangzhou, et al.
Pubblicazione: (2026)
di: Wu, Fangzhou, et al.
Pubblicazione: (2026)
Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
di: Wu, Fangzhou, et al.
Pubblicazione: (2025)
di: Wu, Fangzhou, et al.
Pubblicazione: (2025)
The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
di: Zhang, Jiaxin, et al.
Pubblicazione: (2026)
di: Zhang, Jiaxin, et al.
Pubblicazione: (2026)
Beyond Worst-Case Dimensionality Reduction for Sparse Vectors
di: Silwal, Sandeep, et al.
Pubblicazione: (2025)
di: Silwal, Sandeep, et al.
Pubblicazione: (2025)
Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
di: Wu, Jiayun, et al.
Pubblicazione: (2025)
di: Wu, Jiayun, et al.
Pubblicazione: (2025)
Prompts Generalize with Low Data: Non-vacuous Generalization Bounds for Optimizing Prompts with More Informative Priors
di: Madras, David, et al.
Pubblicazione: (2025)
di: Madras, David, et al.
Pubblicazione: (2025)
Advancing Long-Term Multi-Energy Load Forecasting with Patchformer: A Patch and Transformer-Based Approach
di: Hong, Qiuyi, et al.
Pubblicazione: (2024)
di: Hong, Qiuyi, et al.
Pubblicazione: (2024)
On Calibration of Large Language Models: From Response To Capability
di: Yang, Sin-Han, et al.
Pubblicazione: (2026)
di: Yang, Sin-Han, et al.
Pubblicazione: (2026)
Complex Logical Query Answering by Calibrating Knowledge Graph Completion Models
di: Xiao, Changyi, et al.
Pubblicazione: (2024)
di: Xiao, Changyi, et al.
Pubblicazione: (2024)
Learn to Think: Bootstrapping LLM Reasoning Capability Through Graph Representation Learning
di: Gao, Hang, et al.
Pubblicazione: (2025)
di: Gao, Hang, et al.
Pubblicazione: (2025)
ModelGPT: Unleashing LLM's Capabilities for Tailored Model Generation
di: Tang, Zihao, et al.
Pubblicazione: (2024)
di: Tang, Zihao, et al.
Pubblicazione: (2024)
Graph-R1: Incentivizing the Zero-Shot Graph Learning Capability in LLMs via Explicit Reasoning
di: Wu, Yicong, et al.
Pubblicazione: (2025)
di: Wu, Yicong, et al.
Pubblicazione: (2025)
Ontology of Belief Diversity: A Community-Based Epistemological Approach
di: Fischella, Tyler, et al.
Pubblicazione: (2024)
di: Fischella, Tyler, et al.
Pubblicazione: (2024)
A Modular Dataset to Demonstrate LLM Abstraction Capability
di: Atanas, Adam, et al.
Pubblicazione: (2025)
di: Atanas, Adam, et al.
Pubblicazione: (2025)
QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill
di: Jones, Dalton, et al.
Pubblicazione: (2026)
di: Jones, Dalton, et al.
Pubblicazione: (2026)
[Re] Benchmarking LLM Capabilities in Negotiation through Scoreable Games
di: Pollo, Jorge Carrasco, et al.
Pubblicazione: (2026)
di: Pollo, Jorge Carrasco, et al.
Pubblicazione: (2026)
Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
di: Bohnet, Bernd, et al.
Pubblicazione: (2025)
di: Bohnet, Bernd, et al.
Pubblicazione: (2025)
CluCERT: Certifying LLM Robustness via Clustering-Guided Denoising Smoothing
di: Wang, Zixia, et al.
Pubblicazione: (2025)
di: Wang, Zixia, et al.
Pubblicazione: (2025)
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
di: Zhan, Weixiao, et al.
Pubblicazione: (2026)
di: Zhan, Weixiao, et al.
Pubblicazione: (2026)
Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits
di: Pershin, Maksim, et al.
Pubblicazione: (2026)
di: Pershin, Maksim, et al.
Pubblicazione: (2026)
Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
di: Luo, Beier, et al.
Pubblicazione: (2025)
di: Luo, Beier, et al.
Pubblicazione: (2025)
Increasing LLM Coding Capabilities through Diverse Synthetic Coding Tasks
di: Abed, Amal, et al.
Pubblicazione: (2025)
di: Abed, Amal, et al.
Pubblicazione: (2025)
ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
di: Patel, Shivam, et al.
Pubblicazione: (2025)
di: Patel, Shivam, et al.
Pubblicazione: (2025)
Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
di: Zhong, Qimin, et al.
Pubblicazione: (2025)
di: Zhong, Qimin, et al.
Pubblicazione: (2025)
Calibration-Aware Policy Optimization for Reasoning LLMs
di: Wang, Ziqi, et al.
Pubblicazione: (2026)
di: Wang, Ziqi, et al.
Pubblicazione: (2026)
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts
di: Zhang, Rui, et al.
Pubblicazione: (2026)
di: Zhang, Rui, et al.
Pubblicazione: (2026)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
di: Dadkhahi, Hamid, et al.
Pubblicazione: (2025)
di: Dadkhahi, Hamid, et al.
Pubblicazione: (2025)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2025)
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2025)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
AIMM: An AI-Driven Multimodal Framework for Detecting Social-Media-Influenced Stock Market Manipulation
di: Neela, Sandeep
Pubblicazione: (2025)
di: Neela, Sandeep
Pubblicazione: (2025)
Machine Unlearning via Null Space Calibration
di: Chen, Huiqiang, et al.
Pubblicazione: (2024)
di: Chen, Huiqiang, et al.
Pubblicazione: (2024)
Understanding Transformer Reasoning Capabilities via Graph Algorithms
di: Sanford, Clayton, et al.
Pubblicazione: (2024)
di: Sanford, Clayton, et al.
Pubblicazione: (2024)
Is More Context Always Better? Examining LLM Reasoning Capability for Time Interval Prediction
di: Cao, Yanan, et al.
Pubblicazione: (2026)
di: Cao, Yanan, et al.
Pubblicazione: (2026)
Quantifying Calibration Error in Neural Networks Through Evidence-Based Theory
di: Ouattara, Koffi Ismael, et al.
Pubblicazione: (2024)
di: Ouattara, Koffi Ismael, et al.
Pubblicazione: (2024)
Plausibility Is Not Prediction: Contrastive Evidence for LLM-Based Cellular Perturbation Reasoning
di: Yuan, Xinyu, et al.
Pubblicazione: (2026)
di: Yuan, Xinyu, et al.
Pubblicazione: (2026)
Self-Reinforcing Controllable Synthesis of Rare Relational Data via Bayesian Calibration
di: Zhang, Chongsheng, et al.
Pubblicazione: (2026)
di: Zhang, Chongsheng, et al.
Pubblicazione: (2026)
DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QA
di: Yin, Jianing, et al.
Pubblicazione: (2026)
di: Yin, Jianing, et al.
Pubblicazione: (2026)
DDT: A Dual-Masking Dual-Expert Transformer for Energy Time-Series Forecasting
di: Zhu, Mingnan, et al.
Pubblicazione: (2026)
di: Zhu, Mingnan, et al.
Pubblicazione: (2026)
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2025)
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Randomization Boosts KV Caching, Learning Balances Query Load: A Joint Perspective
di: Wu, Fangzhou, et al.
Pubblicazione: (2026) -
DynMuon: A Dynamic Spectral Shaping View of Muon
di: Wu, Fangzhou, et al.
Pubblicazione: (2026) -
Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
di: Wu, Fangzhou, et al.
Pubblicazione: (2025) -
The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
di: Zhang, Jiaxin, et al.
Pubblicazione: (2026) -
Beyond Worst-Case Dimensionality Reduction for Sparse Vectors
di: Silwal, Sandeep, et al.
Pubblicazione: (2025)