When Does Verification Pay Off? A Closer Look at LLMs as Solution Verifiers
Fuente:
arXiv
Guardado en:
| Autores principales: | Lu, Jack, Teehan, Ryan, Jin, Jinran, Ren, Mengye |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
por: Dai, Hui, et al.
Publicado: (2024)
por: Dai, Hui, et al.
Publicado: (2024)
Context Tuning for In-Context Optimization
por: Lu, Jack, et al.
Publicado: (2025)
por: Lu, Jack, et al.
Publicado: (2025)
CoLLEGe: Concept Embedding Generation for Large Language Models
por: Teehan, Ryan, et al.
Publicado: (2024)
por: Teehan, Ryan, et al.
Publicado: (2024)
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
por: Dai, Hui, et al.
Publicado: (2026)
por: Dai, Hui, et al.
Publicado: (2026)
ProCreate, Don't Reproduce! Propulsive Energy Diffusion for Creative Generation
por: Lu, Jack, et al.
Publicado: (2024)
por: Lu, Jack, et al.
Publicado: (2024)
A Closer Look into LLMs for Table Understanding
por: Wang, Jia, et al.
Publicado: (2026)
por: Wang, Jia, et al.
Publicado: (2026)
A General Framework for Inference-time Scaling and Steering of Diffusion Models
por: Singhal, Raghav, et al.
Publicado: (2025)
por: Singhal, Raghav, et al.
Publicado: (2025)
A Closer Look at Logical Reasoning with LLMs: The Choice of Tool Matters
por: Lam, Long Hei Matthew, et al.
Publicado: (2024)
por: Lam, Long Hei Matthew, et al.
Publicado: (2024)
A Closer Look at System Prompt Robustness
por: Mu, Norman, et al.
Publicado: (2025)
por: Mu, Norman, et al.
Publicado: (2025)
A Closer Look at Claim Decomposition
por: Wanner, Miriam, et al.
Publicado: (2024)
por: Wanner, Miriam, et al.
Publicado: (2024)
A Closer Look at the Self-Verification Abilities of Large Language Models in Logical Reasoning
por: Hong, Ruixin, et al.
Publicado: (2023)
por: Hong, Ruixin, et al.
Publicado: (2023)
A Closer Look at the Limitations of Instruction Tuning
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMs
por: Sanz-Guerrero, Mario, et al.
Publicado: (2025)
por: Sanz-Guerrero, Mario, et al.
Publicado: (2025)
A Closer Look into Mixture-of-Experts in Large Language Models
por: Lo, Ka Man, et al.
Publicado: (2024)
por: Lo, Ka Man, et al.
Publicado: (2024)
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
por: Sprague, Zayne, et al.
Publicado: (2025)
por: Sprague, Zayne, et al.
Publicado: (2025)
Is This a Bad Table? A Closer Look at the Evaluation of Table Generation from Text
por: Ramu, Pritika, et al.
Publicado: (2024)
por: Ramu, Pritika, et al.
Publicado: (2024)
A Closer Look at Machine Unlearning for Large Language Models
por: Yuan, Xiaojian, et al.
Publicado: (2024)
por: Yuan, Xiaojian, et al.
Publicado: (2024)
Scaling Behaviors of Evolutionary Algorithms on GPUs: When Does Parallelism Pay Off?
por: Yu, Xinmeng, et al.
Publicado: (2026)
por: Yu, Xinmeng, et al.
Publicado: (2026)
When Does Sparsity Mitigate the Curse of Depth in LLMs
por: Muhtar, Dilxat, et al.
Publicado: (2026)
por: Muhtar, Dilxat, et al.
Publicado: (2026)
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
por: Ryu, Hyun, et al.
Publicado: (2024)
por: Ryu, Hyun, et al.
Publicado: (2024)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
por: Singhi, Nishad, et al.
Publicado: (2025)
por: Singhi, Nishad, et al.
Publicado: (2025)
A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs
por: Merchant, Humzah, et al.
Publicado: (2025)
por: Merchant, Humzah, et al.
Publicado: (2025)
A Closer Look at Classification Evaluation Metrics and a Critical Reflection of Common Evaluation Practice
por: Opitz, Juri
Publicado: (2024)
por: Opitz, Juri
Publicado: (2024)
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
por: Du, Wenyu, et al.
Publicado: (2024)
por: Du, Wenyu, et al.
Publicado: (2024)
VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues
por: Zhang, Jianshu, et al.
Publicado: (2025)
por: Zhang, Jianshu, et al.
Publicado: (2025)
Shrinking the Generation-Verification Gap with Weak Verifiers
por: Saad-Falcon, Jon, et al.
Publicado: (2025)
por: Saad-Falcon, Jon, et al.
Publicado: (2025)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
por: Liu, Xiaoyuan, et al.
Publicado: (2025)
por: Liu, Xiaoyuan, et al.
Publicado: (2025)
When to Think and When to Look: Uncertainty-Guided Lookback
por: Bi, Jing, et al.
Publicado: (2025)
por: Bi, Jing, et al.
Publicado: (2025)
ptt5-v2: A Closer Look at Continued Pretraining of T5 Models for the Portuguese Language
por: Piau, Marcos, et al.
Publicado: (2024)
por: Piau, Marcos, et al.
Publicado: (2024)
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
por: Zhai, Weiqi, et al.
Publicado: (2026)
por: Zhai, Weiqi, et al.
Publicado: (2026)
AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models
por: Yang, Chih-Kai, et al.
Publicado: (2025)
por: Yang, Chih-Kai, et al.
Publicado: (2025)
A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
por: Liu, Alexander H., et al.
Publicado: (2024)
por: Liu, Alexander H., et al.
Publicado: (2024)
DiVA: Fine-grained Factuality Verification with Agentic-Discriminative Verifier
por: Huang, Hui, et al.
Publicado: (2026)
por: Huang, Hui, et al.
Publicado: (2026)
Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured training
por: Yang, Yanlai, et al.
Publicado: (2024)
por: Yang, Yanlai, et al.
Publicado: (2024)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
por: Balasubramanian, Sriram, et al.
Publicado: (2025)
por: Balasubramanian, Sriram, et al.
Publicado: (2025)
Trust but Verify! A Survey on Verification Design for Test-time Scaling
por: Venktesh, V, et al.
Publicado: (2025)
por: Venktesh, V, et al.
Publicado: (2025)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
por: Liu, Shudong, et al.
Publicado: (2025)
por: Liu, Shudong, et al.
Publicado: (2025)
VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems
por: Qiao, Hezhe, et al.
Publicado: (2026)
por: Qiao, Hezhe, et al.
Publicado: (2026)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
por: Su, Xin, et al.
Publicado: (2026)
por: Su, Xin, et al.
Publicado: (2026)
DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs
por: Yin, Lake, et al.
Publicado: (2025)
por: Yin, Lake, et al.
Publicado: (2025)
Ejemplares similares
-
Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
por: Dai, Hui, et al.
Publicado: (2024) -
Context Tuning for In-Context Optimization
por: Lu, Jack, et al.
Publicado: (2025) -
CoLLEGe: Concept Embedding Generation for Large Language Models
por: Teehan, Ryan, et al.
Publicado: (2024) -
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
por: Dai, Hui, et al.
Publicado: (2026) -
ProCreate, Don't Reproduce! Propulsive Energy Diffusion for Creative Generation
por: Lu, Jack, et al.
Publicado: (2024)