Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
Fuente:
arXiv
Saved in:
| Main Authors: | Agarwal, Anmol, Neamtu, Natalie, Aggarwal, Pranjal, Kim, Seungone, Limperg, Jannis, Flamant, Cedric, Shimizu, Kanna, Parno, Bryan, Welleck, Sean |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement
by: Aggarwal, Pranjal, et al.
Published: (2024)
by: Aggarwal, Pranjal, et al.
Published: (2024)
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
by: Flamant, Cedric, et al.
Published: (2026)
by: Flamant, Cedric, et al.
Published: (2026)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
Agentic-R1: Distilled Dual-Strategy Reasoning
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification
by: Shea, Ryan, et al.
Published: (2025)
by: Shea, Ryan, et al.
Published: (2025)
ExVerus: Verus Proof Repair via Counterexample Reasoning
by: Yang, Jun, et al.
Published: (2026)
by: Yang, Jun, et al.
Published: (2026)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
WaveCert: Translation Validation for Asynchronous Dataflow Programs via Dynamic Fractional Permissions
by: Lin, Zhengyao, et al.
Published: (2023)
by: Lin, Zhengyao, et al.
Published: (2023)
SpecLoop: An Agentic RTL-to-Specification Framework with Formal Verification Feedback Loop
by: Chang, Fu-Chieh, et al.
Published: (2026)
by: Chang, Fu-Chieh, et al.
Published: (2026)
Evaluating Language Models as Synthetic Data Generators
by: Kim, Seungone, et al.
Published: (2024)
by: Kim, Seungone, et al.
Published: (2024)
Consistent Autoformalization for Constructing Mathematical Libraries
by: Zhang, Lan, et al.
Published: (2024)
by: Zhang, Lan, et al.
Published: (2024)
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
by: Kim, Seungone, et al.
Published: (2024)
by: Kim, Seungone, et al.
Published: (2024)
CktFormalizer: Autoformalization of Natural Language into Circuit Representations
by: Xiong, Jing, et al.
Published: (2026)
by: Xiong, Jing, et al.
Published: (2026)
GEM: A Gym for Agentic LLMs
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
MMFormalizer: Multimodal Autoformalization in the Wild
by: Xiong, Jing, et al.
Published: (2026)
by: Xiong, Jing, et al.
Published: (2026)
Scaling Evaluation-time Compute with Reasoning Models as Evaluators
by: Kim, Seungone, et al.
Published: (2025)
by: Kim, Seungone, et al.
Published: (2025)
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
by: Lee, Young-Jun, et al.
Published: (2025)
by: Lee, Young-Jun, et al.
Published: (2025)
An Evaluation Benchmark for Autoformalization in Lean4
by: Gulati, Aryan, et al.
Published: (2024)
by: Gulati, Aryan, et al.
Published: (2024)
SpecEval: Evaluating Model Adherence to Behavior Specifications
by: Ahmed, Ahmed, et al.
Published: (2025)
by: Ahmed, Ahmed, et al.
Published: (2025)
MASA: LLM-Driven Multi-Agent Systems for Autoformalization
by: Zhang, Lan, et al.
Published: (2025)
by: Zhang, Lan, et al.
Published: (2025)
miniCTX: Neural Theorem Proving with (Long-)Contexts
by: Hu, Jiewen, et al.
Published: (2024)
by: Hu, Jiewen, et al.
Published: (2024)
Optimizing Temperature for Language Models with Multi-Sample Inference
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
Monotonic Reference-Free Refinement for Autoformalization
by: Zhang, Lan, et al.
Published: (2026)
by: Zhang, Lan, et al.
Published: (2026)
Can Language Models Evaluate Human Written Text? Case Study on Korean Student Writing for Education
by: Kim, Seungyoon, et al.
Published: (2024)
by: Kim, Seungyoon, et al.
Published: (2024)
AutoVerus: Automated Proof Generation for Rust Code
by: Yang, Chenyuan, et al.
Published: (2024)
by: Yang, Chenyuan, et al.
Published: (2024)
Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitions
by: Zhang, Lan, et al.
Published: (2025)
by: Zhang, Lan, et al.
Published: (2025)
Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing
by: Gulati, Anmol, et al.
Published: (2026)
by: Gulati, Anmol, et al.
Published: (2026)
FormalAlign: Automated Alignment Evaluation for Autoformalization
by: Lu, Jianqiao, et al.
Published: (2024)
by: Lu, Jianqiao, et al.
Published: (2024)
Propose, Solve, Verify: Self-Play Through Formal Verification
by: Wilf, Alex, et al.
Published: (2025)
by: Wilf, Alex, et al.
Published: (2025)
Autoformalize Mathematical Statements by Symbolic Equivalence and Semantic Consistency
by: Li, Zenan, et al.
Published: (2024)
by: Li, Zenan, et al.
Published: (2024)
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
A New Approach Towards Autoformalization
by: Patel, Nilay, et al.
Published: (2023)
by: Patel, Nilay, et al.
Published: (2023)
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
by: Lee, Seongyun, et al.
Published: (2025)
by: Lee, Seongyun, et al.
Published: (2025)
ReForm: Reflective Autoformalization with Prospective Bounded Sequence Optimization
by: Chen, Guoxin, et al.
Published: (2025)
by: Chen, Guoxin, et al.
Published: (2025)
HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
by: Chen, Hongzheng, et al.
Published: (2025)
by: Chen, Hongzheng, et al.
Published: (2025)
Towards Autoformalization of LLM-generated Outputs for Requirement Verification
by: Gupte, Mihir, et al.
Published: (2025)
by: Gupte, Mihir, et al.
Published: (2025)
Argument Reconstruction as Supervision for Critical Thinking in LLMs
by: Ryu, Hyun, et al.
Published: (2026)
by: Ryu, Hyun, et al.
Published: (2026)
Faithful Autoformalization via Roundtrip Verification and Repair
by: Amrollahi, Daneshvar, et al.
Published: (2026)
by: Amrollahi, Daneshvar, et al.
Published: (2026)
Similar Items
-
AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement
by: Aggarwal, Pranjal, et al.
Published: (2024) -
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
by: Flamant, Cedric, et al.
Published: (2026) -
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
by: Aggarwal, Pranjal, et al.
Published: (2025) -
Agentic-R1: Distilled Dual-Strategy Reasoning
by: Du, Weihua, et al.
Published: (2025) -
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
by: Aggarwal, Pranjal, et al.
Published: (2025)