NP-Hard Lower Bound Complexity for Semantic Self-Verification
Fuente:
arXiv
Salvato in:
| Autore principale: | Young, Robin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Empirical Sufficiency Lower Bounds for Language Modeling with Locally-Bootstrapped Semantic Structures
di: Prange, Jakob, et al.
Pubblicazione: (2023)
di: Prange, Jakob, et al.
Pubblicazione: (2023)
NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
di: Wang, Yuyao, et al.
Pubblicazione: (2025)
di: Wang, Yuyao, et al.
Pubblicazione: (2025)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
di: Barone, Antonio Valerio Miceli, et al.
Pubblicazione: (2026)
di: Barone, Antonio Valerio Miceli, et al.
Pubblicazione: (2026)
Evergreen: Efficient Claim Verification for Semantic Aggregates
di: Lee, Alexander W., et al.
Pubblicazione: (2026)
di: Lee, Alexander W., et al.
Pubblicazione: (2026)
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
di: Chen, Jiefeng, et al.
Pubblicazione: (2025)
di: Chen, Jiefeng, et al.
Pubblicazione: (2025)
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
di: Kumar, Adarsh, et al.
Pubblicazione: (2025)
di: Kumar, Adarsh, et al.
Pubblicazione: (2025)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
BiDeV: Bilateral Defusing Verification for Complex Claim Fact-Checking
di: Liu, Yuxuan, et al.
Pubblicazione: (2025)
di: Liu, Yuxuan, et al.
Pubblicazione: (2025)
Self-Verification is All You Need To Pass The Japanese Bar Examination
di: Shin, Andrew
Pubblicazione: (2026)
di: Shin, Andrew
Pubblicazione: (2026)
Are Machines Better at Complex Reasoning? Unveiling Human-Machine Inference Gaps in Entailment Verification
di: Sanyal, Soumya, et al.
Pubblicazione: (2024)
di: Sanyal, Soumya, et al.
Pubblicazione: (2024)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
di: Liu, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Liu, Xiaoyuan, et al.
Pubblicazione: (2025)
Take It Easy: Label-Adaptive Self-Rationalization for Fact Verification and Explanation Generation
di: Yang, Jing, et al.
Pubblicazione: (2024)
di: Yang, Jing, et al.
Pubblicazione: (2024)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
di: Fan, Dongyang, et al.
Pubblicazione: (2026)
di: Fan, Dongyang, et al.
Pubblicazione: (2026)
Self-Trained Verification for Training- and Test-Time Self-Improvement
di: Wu, Chen Henry, et al.
Pubblicazione: (2026)
di: Wu, Chen Henry, et al.
Pubblicazione: (2026)
LLM Self-Explanations Fail Semantic Invariance
di: Szeider, Stefan
Pubblicazione: (2026)
di: Szeider, Stefan
Pubblicazione: (2026)
Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
di: Jang, Chaeyun, et al.
Pubblicazione: (2025)
di: Jang, Chaeyun, et al.
Pubblicazione: (2025)
A Closer Look at the Self-Verification Abilities of Large Language Models in Logical Reasoning
di: Hong, Ruixin, et al.
Pubblicazione: (2023)
di: Hong, Ruixin, et al.
Pubblicazione: (2023)
Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
di: Zhang, Anqi, et al.
Pubblicazione: (2025)
di: Zhang, Anqi, et al.
Pubblicazione: (2025)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
Self-Supervised Learning Based Handwriting Verification
di: Chauhan, Mihir, et al.
Pubblicazione: (2024)
di: Chauhan, Mihir, et al.
Pubblicazione: (2024)
Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP
di: Rozner, Josh, et al.
Pubblicazione: (2021)
di: Rozner, Josh, et al.
Pubblicazione: (2021)
Language Models can Infer Action Semantics for Symbolic Planners from Environment Feedback
di: Zhu, Wang, et al.
Pubblicazione: (2024)
di: Zhu, Wang, et al.
Pubblicazione: (2024)
Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
di: Creo, Aldan, et al.
Pubblicazione: (2025)
di: Creo, Aldan, et al.
Pubblicazione: (2025)
SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration
di: Wen, Zhuofan, et al.
Pubblicazione: (2026)
di: Wen, Zhuofan, et al.
Pubblicazione: (2026)
Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
di: Jiang, Chunyang, et al.
Pubblicazione: (2025)
di: Jiang, Chunyang, et al.
Pubblicazione: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
di: Soligo, Anna, et al.
Pubblicazione: (2026)
di: Soligo, Anna, et al.
Pubblicazione: (2026)
Circuit Complexity Bounds for Visual Autoregressive Model
di: Ke, Yekun, et al.
Pubblicazione: (2025)
di: Ke, Yekun, et al.
Pubblicazione: (2025)
Optimizing Decomposition for Optimal Claim Verification
di: Lu, Yining, et al.
Pubblicazione: (2025)
di: Lu, Yining, et al.
Pubblicazione: (2025)
CAVE: Controllable Authorship Verification Explanations
di: Ramnath, Sahana, et al.
Pubblicazione: (2024)
di: Ramnath, Sahana, et al.
Pubblicazione: (2024)
Rethinking Loss Functions for Fact Verification
di: Mukobara, Yuta, et al.
Pubblicazione: (2024)
di: Mukobara, Yuta, et al.
Pubblicazione: (2024)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
di: Maniparambil, Mayug, et al.
Pubblicazione: (2026)
di: Maniparambil, Mayug, et al.
Pubblicazione: (2026)
Self-Verification Dilemma: Experience-Driven Suppression of Overused Checking in LLM Reasoning
di: Long, Quanyu, et al.
Pubblicazione: (2026)
di: Long, Quanyu, et al.
Pubblicazione: (2026)
Universal NP-Hardness of Clustering under General Utilities
di: Majumdar, Angshul
Pubblicazione: (2026)
di: Majumdar, Angshul
Pubblicazione: (2026)
Zero-Shot Verification-guided Chain of Thoughts
di: Chowdhury, Jishnu Ray, et al.
Pubblicazione: (2025)
di: Chowdhury, Jishnu Ray, et al.
Pubblicazione: (2025)
Faithful Autoformalization via Roundtrip Verification and Repair
di: Amrollahi, Daneshvar, et al.
Pubblicazione: (2026)
di: Amrollahi, Daneshvar, et al.
Pubblicazione: (2026)
General Purpose Verification for Chain of Thought Prompting
di: Vacareanu, Robert, et al.
Pubblicazione: (2024)
di: Vacareanu, Robert, et al.
Pubblicazione: (2024)
RVISA: Reasoning and Verification for Implicit Sentiment Analysis
di: Lai, Wenna, et al.
Pubblicazione: (2024)
di: Lai, Wenna, et al.
Pubblicazione: (2024)
The Alignment Bottleneck in Decomposition-Based Claim Verification
di: Akhter, Mahmud Elahi, et al.
Pubblicazione: (2026)
di: Akhter, Mahmud Elahi, et al.
Pubblicazione: (2026)
Robust Claim Verification Through Fact Detection
di: Jafari, Nazanin, et al.
Pubblicazione: (2024)
di: Jafari, Nazanin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Empirical Sufficiency Lower Bounds for Language Modeling with Locally-Bootstrapped Semantic Structures
di: Prange, Jakob, et al.
Pubblicazione: (2023) -
NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
di: Wang, Yuyao, et al.
Pubblicazione: (2025) -
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
di: Barone, Antonio Valerio Miceli, et al.
Pubblicazione: (2026) -
Evergreen: Efficient Claim Verification for Semantic Aggregates
di: Lee, Alexander W., et al.
Pubblicazione: (2026) -
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
di: Chen, Jiefeng, et al.
Pubblicazione: (2025)