Trust but Verify! A Survey on Verification Design for Test-time Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Venktesh, V, Rathee, Mandeep, Anand, Avishek |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test-time Corpus Feedback: From Retrieval to RAG
by: Rathee, Mandeep, et al.
Published: (2025)
by: Rathee, Mandeep, et al.
Published: (2025)
QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims
by: V, Venktesh, et al.
Published: (2024)
by: V, Venktesh, et al.
Published: (2024)
Evaluating List Construction and Temporal Understanding capabilities of Large Language Models
by: Dumitru, Alexandru, et al.
Published: (2025)
by: Dumitru, Alexandru, et al.
Published: (2025)
Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
by: Chungkham, Primakov, et al.
Published: (2025)
by: Chungkham, Primakov, et al.
Published: (2025)
SUNAR: Semantic Uncertainty based Neighborhood Aware Retrieval for Complex QA
by: Venktesh, V, et al.
Published: (2025)
by: Venktesh, V, et al.
Published: (2025)
When More Reformulations Hurt: Avoiding Drift using Ranker Feedback
by: Venktesh, V, et al.
Published: (2026)
by: Venktesh, V, et al.
Published: (2026)
Guiding Retrieval using LLM-based Listwise Rankers
by: Rathee, Mandeep, et al.
Published: (2025)
by: Rathee, Mandeep, et al.
Published: (2025)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
by: Liu, Xiaoyuan, et al.
Published: (2025)
by: Liu, Xiaoyuan, et al.
Published: (2025)
LiveFC: A System for Live Fact-Checking of Audio Streams
by: V, Venktesh, et al.
Published: (2024)
by: V, Venktesh, et al.
Published: (2024)
A Benchmark for Open-Domain Numerical Fact-Checking Enhanced by Claim Decomposition
by: Venktesh, V, et al.
Published: (2025)
by: Venktesh, V, et al.
Published: (2025)
DEXTER: A Benchmark for open-domain Complex Question Answering using LLMs
by: Prabhu, Venktesh V. Deepali, et al.
Published: (2024)
by: Prabhu, Venktesh V. Deepali, et al.
Published: (2024)
Breaking the Lens of the Telescope: Online Relevance Estimation over Large Retrieval Sets
by: Rathee, Mandeep, et al.
Published: (2025)
by: Rathee, Mandeep, et al.
Published: (2025)
Reproducing Adaptive Reranking for Reasoning-Intensive IR
by: Rathee, Mandeep, et al.
Published: (2026)
by: Rathee, Mandeep, et al.
Published: (2026)
The Surprising Effectiveness of Rankers Trained on Expanded Queries
by: Anand, Abhijit, et al.
Published: (2024)
by: Anand, Abhijit, et al.
Published: (2024)
Understanding the User: An Intent-Based Ranking Dataset
by: Anand, Abhijit, et al.
Published: (2024)
by: Anand, Abhijit, et al.
Published: (2024)
T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models
by: Kang, Minki, et al.
Published: (2025)
by: Kang, Minki, et al.
Published: (2025)
Budget-aware Test-time Scaling via Discriminative Verification
by: Montgomery, Kyle, et al.
Published: (2025)
by: Montgomery, Kyle, et al.
Published: (2025)
Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
by: Yang, Zhuoyi, et al.
Published: (2025)
by: Yang, Zhuoyi, et al.
Published: (2025)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
by: Su, Xin, et al.
Published: (2026)
by: Su, Xin, et al.
Published: (2026)
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
by: Chen, Jiefeng, et al.
Published: (2025)
by: Chen, Jiefeng, et al.
Published: (2025)
VerifiAgent: a Unified Verification Agent in Language Model Reasoning
by: Han, Jiuzhou, et al.
Published: (2025)
by: Han, Jiuzhou, et al.
Published: (2025)
Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning
by: Pronesti, Massimiliano, et al.
Published: (2026)
by: Pronesti, Massimiliano, et al.
Published: (2026)
HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification
by: Paudel, Bibek, et al.
Published: (2025)
by: Paudel, Bibek, et al.
Published: (2025)
Sleep-time Compute: Beyond Inference Scaling at Test-time
by: Lin, Kevin, et al.
Published: (2025)
by: Lin, Kevin, et al.
Published: (2025)
MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering
by: Ning, Yingpeng, et al.
Published: (2025)
by: Ning, Yingpeng, et al.
Published: (2025)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
Claim Verification in the Age of Large Language Models: A Survey
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
by: Fu, Daocheng, et al.
Published: (2025)
by: Fu, Daocheng, et al.
Published: (2025)
Tool Verification for Test-Time Reinforcement Learning
by: Liao, Ruotong, et al.
Published: (2026)
by: Liao, Ruotong, et al.
Published: (2026)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
by: Zhang, Jiazheng, et al.
Published: (2026)
by: Zhang, Jiazheng, et al.
Published: (2026)
LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
by: Lin, Weizhe, et al.
Published: (2025)
by: Lin, Weizhe, et al.
Published: (2025)
AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Verifiable Mathematical Reasoning
by: Bruno, Alessio
Published: (2026)
by: Bruno, Alessio
Published: (2026)
A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
by: Zhang, Qiyuan, et al.
Published: (2025)
by: Zhang, Qiyuan, et al.
Published: (2025)
VerifAI: A Verifiable Open-Source Search Engine for Biomedical Question Answering
by: Košprdić, Miloš, et al.
Published: (2026)
by: Košprdić, Miloš, et al.
Published: (2026)
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
by: Schnabl, Christoph, et al.
Published: (2025)
by: Schnabl, Christoph, et al.
Published: (2025)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
by: Chen, Jiangjie, et al.
Published: (2025)
by: Chen, Jiangjie, et al.
Published: (2025)
Infinite Problem Generator: Verifiably Scaling Physics Reasoning Data with Agentic Workflows
by: Sharan, Aditya, et al.
Published: (2026)
by: Sharan, Aditya, et al.
Published: (2026)
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
by: Ficek, Aleksander, et al.
Published: (2025)
by: Ficek, Aleksander, et al.
Published: (2025)
Similar Items
-
Test-time Corpus Feedback: From Retrieval to RAG
by: Rathee, Mandeep, et al.
Published: (2025) -
QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims
by: V, Venktesh, et al.
Published: (2024) -
Evaluating List Construction and Temporal Understanding capabilities of Large Language Models
by: Dumitru, Alexandru, et al.
Published: (2025) -
Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
by: Chungkham, Primakov, et al.
Published: (2025) -
SUNAR: Semantic Uncertainty based Neighborhood Aware Retrieval for Complex QA
by: Venktesh, V, et al.
Published: (2025)