Propose, Solve, Verify: Self-Play Through Formal Verification
Fuente:
arXiv
Saved in:
| Main Authors: | Wilf, Alex, Aggarwal, Pranjal, Parno, Bryan, Fried, Daniel, Morency, Louis-Philippe, Liang, Paul Pu, Welleck, Sean |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement
by: Aggarwal, Pranjal, et al.
Published: (2024)
by: Aggarwal, Pranjal, et al.
Published: (2024)
From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking
by: Kim, Gyeongwon James, et al.
Published: (2025)
by: Kim, Gyeongwon James, et al.
Published: (2025)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
by: Agarwal, Anmol, et al.
Published: (2026)
by: Agarwal, Anmol, et al.
Published: (2026)
Agentic-R1: Distilled Dual-Strategy Reasoning
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023)
by: Mo, Shentong, et al.
Published: (2023)
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
by: Wang, Zhiruo, et al.
Published: (2024)
by: Wang, Zhiruo, et al.
Published: (2024)
Verifying Non-friendly Formal Verification Designs: Can We Start Earlier?
by: Olmos, Bryan, et al.
Published: (2024)
by: Olmos, Bryan, et al.
Published: (2024)
Reasoning with Latent Tokens in Diffusion Language Models
by: He, Andre, et al.
Published: (2026)
by: He, Andre, et al.
Published: (2026)
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
by: He, Andre, et al.
Published: (2025)
by: He, Andre, et al.
Published: (2025)
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
by: Mathur, Leena, et al.
Published: (2024)
by: Mathur, Leena, et al.
Published: (2024)
Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications
by: Liang, Paul Pu, et al.
Published: (2023)
by: Liang, Paul Pu, et al.
Published: (2023)
Explorable Theorems: Making Written Theorems Explorable by Grounding Them in Formal Representations
by: Kambhamettu, Hita, et al.
Published: (2026)
by: Kambhamettu, Hita, et al.
Published: (2026)
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
by: Wu, Yangzhen, et al.
Published: (2024)
by: Wu, Yangzhen, et al.
Published: (2024)
miniCodeProps: a Minimal Benchmark for Proving Code Properties
by: Lohn, Evan, et al.
Published: (2024)
by: Lohn, Evan, et al.
Published: (2024)
GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards
by: Ahn, Kyeongjin, et al.
Published: (2026)
by: Ahn, Kyeongjin, et al.
Published: (2026)
HEMM: Holistic Evaluation of Multimodal Foundation Models
by: Liang, Paul Pu, et al.
Published: (2024)
by: Liang, Paul Pu, et al.
Published: (2024)
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
by: Hu, Jiewen, et al.
Published: (2025)
by: Hu, Jiewen, et al.
Published: (2025)
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
by: Mathur, Leena, et al.
Published: (2025)
by: Mathur, Leena, et al.
Published: (2025)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
by: Liu, Xiaoyuan, et al.
Published: (2025)
by: Liu, Xiaoyuan, et al.
Published: (2025)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
ImProver 2: Iteratively Self-Improving LMs for Neurosymbolic Proof Optimization
by: Ahuja, Riyaz, et al.
Published: (2026)
by: Ahuja, Riyaz, et al.
Published: (2026)
Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving
by: Zhou, Kuo, et al.
Published: (2025)
by: Zhou, Kuo, et al.
Published: (2025)
miniCTX: Neural Theorem Proving with (Long-)Contexts
by: Hu, Jiewen, et al.
Published: (2024)
by: Hu, Jiewen, et al.
Published: (2024)
Optimizing Temperature for Language Models with Multi-Sample Inference
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
by: Singhi, Nishad, et al.
Published: (2025)
by: Singhi, Nishad, et al.
Published: (2025)
Automated Conjecture Resolution with Formal Verification
by: Ju, Haocheng, et al.
Published: (2026)
by: Ju, Haocheng, et al.
Published: (2026)
Lean-STaR: Learning to Interleave Thinking and Proving
by: Lin, Haohan, et al.
Published: (2024)
by: Lin, Haohan, et al.
Published: (2024)
Generating Pragmatic Examples to Train Neural Program Synthesizers
by: Vaduguru, Saujas, et al.
Published: (2023)
by: Vaduguru, Saujas, et al.
Published: (2023)
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
by: Zhou, Xuhui, et al.
Published: (2023)
by: Zhou, Xuhui, et al.
Published: (2023)
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
by: Fu, Daocheng, et al.
Published: (2025)
by: Fu, Daocheng, et al.
Published: (2025)
Fuse, Reason and Verify: Geometry Problem Solving with Parsed Clauses from Diagram
by: Zhang, Ming-Liang, et al.
Published: (2024)
by: Zhang, Ming-Liang, et al.
Published: (2024)
ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
by: Tanmay, Kumar, et al.
Published: (2025)
by: Tanmay, Kumar, et al.
Published: (2025)
ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning
by: Yang, Chengcao
Published: (2026)
by: Yang, Chengcao
Published: (2026)
Interactive Sketchpad: A Multimodal Tutoring System for Collaborative, Visual Problem-Solving
by: Chen, Steven-Shine, et al.
Published: (2025)
by: Chen, Steven-Shine, et al.
Published: (2025)
Similar Items
-
AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement
by: Aggarwal, Pranjal, et al.
Published: (2024) -
From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking
by: Kim, Gyeongwon James, et al.
Published: (2025) -
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
by: Aggarwal, Pranjal, et al.
Published: (2025) -
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026) -
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
by: Aggarwal, Pranjal, et al.
Published: (2025)