Saved in:
| Main Author: | Young, Robin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2501.15446 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Empirical Sufficiency Lower Bounds for Language Modeling with Locally-Bootstrapped Semantic Structures
by: Prange, Jakob, et al.
Published: (2023)
by: Prange, Jakob, et al.
Published: (2023)
NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
by: Wang, Yuyao, et al.
Published: (2025)
by: Wang, Yuyao, et al.
Published: (2025)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
Evergreen: Efficient Claim Verification for Semantic Aggregates
by: Lee, Alexander W., et al.
Published: (2026)
by: Lee, Alexander W., et al.
Published: (2026)
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
by: Chen, Jiefeng, et al.
Published: (2025)
by: Chen, Jiefeng, et al.
Published: (2025)
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
by: Kumar, Adarsh, et al.
Published: (2025)
by: Kumar, Adarsh, et al.
Published: (2025)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
by: Yoon, Kanghoon, et al.
Published: (2025)
by: Yoon, Kanghoon, et al.
Published: (2025)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
BiDeV: Bilateral Defusing Verification for Complex Claim Fact-Checking
by: Liu, Yuxuan, et al.
Published: (2025)
by: Liu, Yuxuan, et al.
Published: (2025)
Self-Verification is All You Need To Pass The Japanese Bar Examination
by: Shin, Andrew
Published: (2026)
by: Shin, Andrew
Published: (2026)
Self-Trained Verification for Training- and Test-Time Self-Improvement
by: Wu, Chen Henry, et al.
Published: (2026)
by: Wu, Chen Henry, et al.
Published: (2026)
Are Machines Better at Complex Reasoning? Unveiling Human-Machine Inference Gaps in Entailment Verification
by: Sanyal, Soumya, et al.
Published: (2024)
by: Sanyal, Soumya, et al.
Published: (2024)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
by: Liu, Xiaoyuan, et al.
Published: (2025)
by: Liu, Xiaoyuan, et al.
Published: (2025)
Take It Easy: Label-Adaptive Self-Rationalization for Fact Verification and Explanation Generation
by: Yang, Jing, et al.
Published: (2024)
by: Yang, Jing, et al.
Published: (2024)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
by: Fan, Dongyang, et al.
Published: (2026)
by: Fan, Dongyang, et al.
Published: (2026)
LLM Self-Explanations Fail Semantic Invariance
by: Szeider, Stefan
Published: (2026)
by: Szeider, Stefan
Published: (2026)
Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
by: Jang, Chaeyun, et al.
Published: (2025)
by: Jang, Chaeyun, et al.
Published: (2025)
A Closer Look at the Self-Verification Abilities of Large Language Models in Logical Reasoning
by: Hong, Ruixin, et al.
Published: (2023)
by: Hong, Ruixin, et al.
Published: (2023)
Self-Supervised Learning Based Handwriting Verification
by: Chauhan, Mihir, et al.
Published: (2024)
by: Chauhan, Mihir, et al.
Published: (2024)
Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
by: Zhang, Anqi, et al.
Published: (2025)
by: Zhang, Anqi, et al.
Published: (2025)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
by: Zhang, Ziyin, et al.
Published: (2024)
by: Zhang, Ziyin, et al.
Published: (2024)
Language Models can Infer Action Semantics for Symbolic Planners from Environment Feedback
by: Zhu, Wang, et al.
Published: (2024)
by: Zhu, Wang, et al.
Published: (2024)
SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration
by: Wen, Zhuofan, et al.
Published: (2026)
by: Wen, Zhuofan, et al.
Published: (2026)
Circuit Complexity Bounds for Visual Autoregressive Model
by: Ke, Yekun, et al.
Published: (2025)
by: Ke, Yekun, et al.
Published: (2025)
Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP
by: Rozner, Josh, et al.
Published: (2021)
by: Rozner, Josh, et al.
Published: (2021)
Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
by: Creo, Aldan, et al.
Published: (2025)
by: Creo, Aldan, et al.
Published: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
by: Soligo, Anna, et al.
Published: (2026)
by: Soligo, Anna, et al.
Published: (2026)
Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
by: Jiang, Chunyang, et al.
Published: (2025)
by: Jiang, Chunyang, et al.
Published: (2025)
Self-Verification Dilemma: Experience-Driven Suppression of Overused Checking in LLM Reasoning
by: Long, Quanyu, et al.
Published: (2026)
by: Long, Quanyu, et al.
Published: (2026)
Semantic Matters: Multimodal Features for Affective Analysis
by: Hallmen, Tobias, et al.
Published: (2025)
by: Hallmen, Tobias, et al.
Published: (2025)
Optimizing Decomposition for Optimal Claim Verification
by: Lu, Yining, et al.
Published: (2025)
by: Lu, Yining, et al.
Published: (2025)
CAVE: Controllable Authorship Verification Explanations
by: Ramnath, Sahana, et al.
Published: (2024)
by: Ramnath, Sahana, et al.
Published: (2024)
Rethinking Loss Functions for Fact Verification
by: Mukobara, Yuta, et al.
Published: (2024)
by: Mukobara, Yuta, et al.
Published: (2024)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
Universal NP-Hardness of Clustering under General Utilities
by: Majumdar, Angshul
Published: (2026)
by: Majumdar, Angshul
Published: (2026)
Sorting by Strip Swaps is NP-Hard
by: Roy, Swapnoneel, et al.
Published: (2025)
by: Roy, Swapnoneel, et al.
Published: (2025)
Zero-Shot Verification-guided Chain of Thoughts
by: Chowdhury, Jishnu Ray, et al.
Published: (2025)
by: Chowdhury, Jishnu Ray, et al.
Published: (2025)
Faithful Autoformalization via Roundtrip Verification and Repair
by: Amrollahi, Daneshvar, et al.
Published: (2026)
by: Amrollahi, Daneshvar, et al.
Published: (2026)
General Purpose Verification for Chain of Thought Prompting
by: Vacareanu, Robert, et al.
Published: (2024)
by: Vacareanu, Robert, et al.
Published: (2024)
Similar Items
-
Empirical Sufficiency Lower Bounds for Language Modeling with Locally-Bootstrapped Semantic Structures
by: Prange, Jakob, et al.
Published: (2023) -
NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
by: Wang, Yuyao, et al.
Published: (2025) -
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026) -
Evergreen: Efficient Claim Verification for Semantic Aggregates
by: Lee, Alexander W., et al.
Published: (2026) -
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
by: Chen, Jiefeng, et al.
Published: (2025)