VERDI: Single-Call Confidence Estimation for Verification-Based LLM Judges via Decomposed Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Jasmine, Dantsev, Danylo, Sun, Muyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FGTR: Fine-Grained Multi-Table Retrieval via Hierarchical LLM Reasoning
by: Sun, Chaojie, et al.
Published: (2026)
by: Sun, Chaojie, et al.
Published: (2026)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Judging with Personality and Confidence: A Study on Personality-Conditioned LLM Relevance Assessment
by: Chen, Nuo, et al.
Published: (2026)
by: Chen, Nuo, et al.
Published: (2026)
Task-Adaptive Embedding Refinement via Test-time LLM Guidance
by: Gera, Ariel, et al.
Published: (2026)
by: Gera, Ariel, et al.
Published: (2026)
Advancing Academic Knowledge Retrieval via LLM-enhanced Representation Similarity Fusion
by: Dai, Wei, et al.
Published: (2024)
by: Dai, Wei, et al.
Published: (2024)
Navigating Ideation Space: Decomposed Conceptual Representations for Positioning Scientific Ideas
by: Shen, Yuexi, et al.
Published: (2026)
by: Shen, Yuexi, et al.
Published: (2026)
Drowning in Documents: Consequences of Scaling Reranker Inference
by: Jacob, Mathew, et al.
Published: (2024)
by: Jacob, Mathew, et al.
Published: (2024)
Atomic Information Flow: A Network Flow Model for Tool Attributions in RAG Systems
by: Gao, James, et al.
Published: (2026)
by: Gao, James, et al.
Published: (2026)
Graph-based Confidence Calibration for Large Language Models
by: Li, Yukun, et al.
Published: (2024)
by: Li, Yukun, et al.
Published: (2024)
Diagnosing LLM Reranker Behavior Under Fixed Evidence Pools
by: Arat, Baris, et al.
Published: (2026)
by: Arat, Baris, et al.
Published: (2026)
Scaling Up LLM Reviews for Google Ads Content Moderation
by: Qiao, Wei, et al.
Published: (2024)
by: Qiao, Wei, et al.
Published: (2024)
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
by: LeVine, Will, et al.
Published: (2025)
by: LeVine, Will, et al.
Published: (2025)
LawLLM: Law Large Language Model for the US Legal System
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification
by: MiroMind Team, et al.
Published: (2026)
by: MiroMind Team, et al.
Published: (2026)
Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators
by: Su, Zhengyang, et al.
Published: (2026)
by: Su, Zhengyang, et al.
Published: (2026)
Transforming User Defined Criteria into Explainable Indicators with an Integrated LLM AHP System
by: Bang, Geonwoo, et al.
Published: (2025)
by: Bang, Geonwoo, et al.
Published: (2025)
LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
by: Zhang, Yunfan, et al.
Published: (2026)
by: Zhang, Yunfan, et al.
Published: (2026)
Evaluation of LLM-based Strategies for the Extraction of Food Product Information from Online Shops
by: Brosch, Christoph, et al.
Published: (2025)
by: Brosch, Christoph, et al.
Published: (2025)
FIRE: Fact-checking with Iterative Retrieval and Verification
by: Xie, Zhuohan, et al.
Published: (2024)
by: Xie, Zhuohan, et al.
Published: (2024)
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
PASC: Pipeline-Aware Conformal Prediction with Joint Coverage Guarantees for Multi-Stage NLP and LLM Pipelines
by: Kotte, Varun
Published: (2026)
by: Kotte, Varun
Published: (2026)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
LLM vs. Lawyers: Identifying a Subset of Summary Judgments in a Large UK Case Law Dataset
by: Izzidien, Ahmed, et al.
Published: (2024)
by: Izzidien, Ahmed, et al.
Published: (2024)
Description-Based Text Similarity
by: Ravfogel, Shauli, et al.
Published: (2023)
by: Ravfogel, Shauli, et al.
Published: (2023)
On the Theoretical Limitations of Embedding-Based Retrieval
by: Weller, Orion, et al.
Published: (2025)
by: Weller, Orion, et al.
Published: (2025)
MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
by: Li, Guoyao, et al.
Published: (2025)
by: Li, Guoyao, et al.
Published: (2025)
Integrating Large Language Models with Graphical Session-Based Recommendation
by: Guo, Naicheng, et al.
Published: (2024)
by: Guo, Naicheng, et al.
Published: (2024)
A comparison of latent semantic analysis and correspondence analysis of document-term matrices
by: Qi, Qianqian, et al.
Published: (2021)
by: Qi, Qianqian, et al.
Published: (2021)
Case-Grounded Evidence Verification: A Framework for Constructing Evidence-Sensitive Supervision
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
StealthRank: LLM Ranking Manipulation via Stealthy Prompt Optimization
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Large Language Model Can Be a Foundation for Hidden Rationale-Based Retrieval
by: Ji, Luo, et al.
Published: (2024)
by: Ji, Luo, et al.
Published: (2024)
GraphER: An Efficient Graph-Based Enrichment and Reranking Method for Retrieval-Augmented Generation
by: Miao, Ruizhong, et al.
Published: (2026)
by: Miao, Ruizhong, et al.
Published: (2026)
ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging
by: Verma, Neha, et al.
Published: (2026)
by: Verma, Neha, et al.
Published: (2026)
ELMO: Efficiency via Low-precision and Peak Memory Optimization in Large Output Spaces
by: Zhang, Jinbin, et al.
Published: (2025)
by: Zhang, Jinbin, et al.
Published: (2025)
Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts
by: Neumann, Julius, et al.
Published: (2025)
by: Neumann, Julius, et al.
Published: (2025)
CURE:Circuit-Aware Unlearning for LLM-based Recommendation
by: Chen, Ziheng, et al.
Published: (2026)
by: Chen, Ziheng, et al.
Published: (2026)
Scaling Test-Time Inference with Policy-Optimized, Dynamic Retrieval-Augmented Generation via KV Caching and Decoding
by: Srinivas, Sakhinana Sagar, et al.
Published: (2025)
by: Srinivas, Sakhinana Sagar, et al.
Published: (2025)
Membership Inference Attacks on LLM-based Recommender Systems
by: He, Jiajie, et al.
Published: (2025)
by: He, Jiajie, et al.
Published: (2025)
Simple Is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation
by: Li, Mufei, et al.
Published: (2024)
by: Li, Mufei, et al.
Published: (2024)
An Integrated Data Processing Framework for Pretraining Foundation Models
by: Sun, Yiding, et al.
Published: (2024)
by: Sun, Yiding, et al.
Published: (2024)
Similar Items
-
FGTR: Fine-Grained Multi-Table Retrieval via Hierarchical LLM Reasoning
by: Sun, Chaojie, et al.
Published: (2026) -
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
by: Xi, Zhiheng, et al.
Published: (2025) -
Judging with Personality and Confidence: A Study on Personality-Conditioned LLM Relevance Assessment
by: Chen, Nuo, et al.
Published: (2026) -
Task-Adaptive Embedding Refinement via Test-time LLM Guidance
by: Gera, Ariel, et al.
Published: (2026) -
Advancing Academic Knowledge Retrieval via LLM-enhanced Representation Similarity Fusion
by: Dai, Wei, et al.
Published: (2024)