Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability
Fuente:
arXiv
Saved in:
| Main Author: | Naik, Ninad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CIRCUS: Circuit Consensus under Uncertainty via Stability Ensembles
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
Less is More for Improving Automatic Evaluation of Factual Consistency
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
Enhancing Annotated Bibliography Generation with LLM Ensembles
by: Bermejo, Sergio
Published: (2024)
by: Bermejo, Sergio
Published: (2024)
DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance
by: Cohen, Seffi, et al.
Published: (2025)
by: Cohen, Seffi, et al.
Published: (2025)
DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning
by: Chan, Chi-Min, et al.
Published: (2026)
by: Chan, Chi-Min, et al.
Published: (2026)
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
by: Jia, Furong, et al.
Published: (2025)
by: Jia, Furong, et al.
Published: (2025)
Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
by: Zhao, Zirui, et al.
Published: (2024)
by: Zhao, Zirui, et al.
Published: (2024)
A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability
by: Omidvar, Hamed, et al.
Published: (2026)
by: Omidvar, Hamed, et al.
Published: (2026)
Pyramid MoA: A Probabilistic Framework for Cost-Optimized Anytime Inference
by: Khaled, Arindam
Published: (2026)
by: Khaled, Arindam
Published: (2026)
Do We Need Frontier Models to Verify Mathematical Proofs?
by: Naik, Aaditya, et al.
Published: (2026)
by: Naik, Aaditya, et al.
Published: (2026)
LLM-Powered Ensemble Learning for Paper Source Tracing: A GPU-Free Approach
by: Chen, Kunlong, et al.
Published: (2024)
by: Chen, Kunlong, et al.
Published: (2024)
TourPlanner: A Competitive Consensus Framework with Constraint-Gated Reinforcement Learning for Travel Planning
by: Wang, Yinuo, et al.
Published: (2026)
by: Wang, Yinuo, et al.
Published: (2026)
Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling
by: Zheng, Hang, et al.
Published: (2025)
by: Zheng, Hang, et al.
Published: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
Concurrent Criterion Validation of a Validity Screen for LLM Confidence Signals via Selective Prediction
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
by: Sturman, Olivia, et al.
Published: (2024)
by: Sturman, Olivia, et al.
Published: (2024)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
PromptMind Team at EHRSQL-2024: Improving Reliability of SQL Generation using Ensemble LLMs
by: Gundabathula, Satya K, et al.
Published: (2024)
by: Gundabathula, Satya K, et al.
Published: (2024)
EMORL: Ensemble Multi-Objective Reinforcement Learning for Efficient and Flexible LLM Fine-Tuning
by: Kong, Lingxiao, et al.
Published: (2025)
by: Kong, Lingxiao, et al.
Published: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
by: Hong, Yihan, et al.
Published: (2026)
by: Hong, Yihan, et al.
Published: (2026)
Uncertainty-Aware Fusion: An Ensemble Framework for Mitigating Hallucinations in Large Language Models
by: Dey, Prasenjit, et al.
Published: (2025)
by: Dey, Prasenjit, et al.
Published: (2025)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
by: Bouchard, Dylan, et al.
Published: (2025)
by: Bouchard, Dylan, et al.
Published: (2025)
Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Reliability
by: Guo, Kevin H., et al.
Published: (2026)
by: Guo, Kevin H., et al.
Published: (2026)
Enhancing NLP Robustness and Generalization through LLM-Generated Contrast Sets: A Scalable Framework for Systematic Evaluation and Adversarial Training
by: Lin, Hender
Published: (2025)
by: Lin, Hender
Published: (2025)
Factual Inconsistency in Data-to-Text Generation Scales Exponentially with LLM Size: A Statistical Validation
by: Mahapatra, Joy, et al.
Published: (2025)
by: Mahapatra, Joy, et al.
Published: (2025)
PromptMind Team at MEDIQA-CORR 2024: Improving Clinical Text Correction with Error Categorization and LLM Ensembles
by: Gundabathula, Satya Kesav, et al.
Published: (2024)
by: Gundabathula, Satya Kesav, et al.
Published: (2024)
Vidur: A Large-Scale Simulation Framework For LLM Inference
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
A Lightweight LLM Framework for Disaster Humanitarian Information Classification
by: Jinzhen, Han, et al.
Published: (2026)
by: Jinzhen, Han, et al.
Published: (2026)
Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments
by: Mannekote, Amogh, et al.
Published: (2024)
by: Mannekote, Amogh, et al.
Published: (2024)
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024)
by: Schoenegger, Philipp, et al.
Published: (2024)
LLM generation novelty through the lens of semantic similarity
by: Davydov, Philipp, et al.
Published: (2025)
by: Davydov, Philipp, et al.
Published: (2025)
Reinforce LLM Reasoning through Multi-Agent Reflection
by: Yuan, Yurun, et al.
Published: (2025)
by: Yuan, Yurun, et al.
Published: (2025)
TaeBench: Improving Quality of Toxic Adversarial Examples
by: Zhu, Xuan, et al.
Published: (2024)
by: Zhu, Xuan, et al.
Published: (2024)
REQUAL-LM: Reliability and Equity through Aggregation in Large Language Models
by: Ebrahimi, Sana, et al.
Published: (2024)
by: Ebrahimi, Sana, et al.
Published: (2024)
Evaluating Fine-Tuned LLM Model For Medical Transcription With Small Low-Resource Languages Validated Dataset
by: Chowdhury, Mohammed Nowshad Ruhani, et al.
Published: (2026)
by: Chowdhury, Mohammed Nowshad Ruhani, et al.
Published: (2026)
Chem-FINESE: Validating Fine-Grained Few-shot Entity Extraction through Text Reconstruction
by: Wang, Qingyun, et al.
Published: (2024)
by: Wang, Qingyun, et al.
Published: (2024)
A Multi-LLM Debiasing Framework
by: Owens, Deonna M., et al.
Published: (2024)
by: Owens, Deonna M., et al.
Published: (2024)
An LLM Feature-based Framework for Dialogue Constructiveness Assessment
by: Zhou, Lexin, et al.
Published: (2024)
by: Zhou, Lexin, et al.
Published: (2024)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
by: Li, Changming, et al.
Published: (2026)
by: Li, Changming, et al.
Published: (2026)
Evaluating Cooperation in LLM Social Groups through Elected Leadership
by: Faulkner, Ryan, et al.
Published: (2026)
by: Faulkner, Ryan, et al.
Published: (2026)
Similar Items
-
CIRCUS: Circuit Consensus under Uncertainty via Stability Ensembles
by: Parekh, Swapnil
Published: (2026) -
Less is More for Improving Automatic Evaluation of Factual Consistency
by: Wang, Tong, et al.
Published: (2024) -
Enhancing Annotated Bibliography Generation with LLM Ensembles
by: Bermejo, Sergio
Published: (2024) -
DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance
by: Cohen, Seffi, et al.
Published: (2025) -
DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning
by: Chan, Chi-Min, et al.
Published: (2026)