TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhiruo, Fried, Daniel, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Compute-Matched Re-Evaluation of TroVE on MATH
by: Sesterhenn, Tobias, et al.
Published: (2025)
by: Sesterhenn, Tobias, et al.
Published: (2025)
Inducing Programmatic Skills for Agentic Tasks
by: Wang, Zora Zhiruo, et al.
Published: (2025)
by: Wang, Zora Zhiruo, et al.
Published: (2025)
What Are Tools Anyway? A Survey from the Language Model Perspective
by: Wang, Zhiruo, et al.
Published: (2024)
by: Wang, Zhiruo, et al.
Published: (2024)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
by: Wang, Zora Zhiruo, et al.
Published: (2025)
by: Wang, Zora Zhiruo, et al.
Published: (2025)
How Well Does Agent Development Reflect Real-World Work?
by: Wang, Zora Zhiruo, et al.
Published: (2026)
by: Wang, Zora Zhiruo, et al.
Published: (2026)
ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
by: Waghjale, Siddhant, et al.
Published: (2024)
by: Waghjale, Siddhant, et al.
Published: (2024)
TOM-SWE: User Mental Modeling For Software Engineering Agents
by: Zhou, Xuhui, et al.
Published: (2025)
by: Zhou, Xuhui, et al.
Published: (2025)
API-Assisted Code Generation for Question Answering on Varied Table Structures
by: Cao, Yihan, et al.
Published: (2023)
by: Cao, Yihan, et al.
Published: (2023)
Propose, Solve, Verify: Self-Play Through Formal Verification
by: Wilf, Alex, et al.
Published: (2025)
by: Wilf, Alex, et al.
Published: (2025)
Agent Workflow Memory
by: Wang, Zora Zhiruo, et al.
Published: (2024)
by: Wang, Zora Zhiruo, et al.
Published: (2024)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
by: Vijayvargiya, Sanidhya, et al.
Published: (2026)
by: Vijayvargiya, Sanidhya, et al.
Published: (2026)
Trust but Verify: Programmatic VLM Evaluation in the Wild
by: Prabhu, Viraj, et al.
Published: (2024)
by: Prabhu, Viraj, et al.
Published: (2024)
Effective Strategies for Asynchronous Software Engineering Agents
by: Geng, Jiayi, et al.
Published: (2026)
by: Geng, Jiayi, et al.
Published: (2026)
Calibrated Reasoning: An Explanatory Verifier for Dynamic and Efficient Problem-Solving
by: Garg, Anisha, et al.
Published: (2025)
by: Garg, Anisha, et al.
Published: (2025)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
by: Huq, Faria, et al.
Published: (2025)
by: Huq, Faria, et al.
Published: (2025)
SELF-GUIDE: Better Task-Specific Instruction Following via Self-Synthetic Finetuning
by: Zhao, Chenyang, et al.
Published: (2024)
by: Zhao, Chenyang, et al.
Published: (2024)
Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier
by: Zhong, Jianyuan, et al.
Published: (2025)
by: Zhong, Jianyuan, et al.
Published: (2025)
Reclaiming the Source of Programmatic Policies: Programmatic versus Latent Spaces
by: Carvalho, Tales H., et al.
Published: (2024)
by: Carvalho, Tales H., et al.
Published: (2024)
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
EvoPool: Evolutionary Programmatic Annotation for Label-Efficient Specialized Supervision
by: Xu, Tianyi, et al.
Published: (2026)
by: Xu, Tianyi, et al.
Published: (2026)
A Rubric-Supervised Critic from Sparse Real-World Outcomes
by: Wang, Xingyao, et al.
Published: (2026)
by: Wang, Xingyao, et al.
Published: (2026)
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
by: Ge, Chris, et al.
Published: (2026)
by: Ge, Chris, et al.
Published: (2026)
Lifting Traces to Logic: Programmatic Skill Induction with Neuro-Symbolic Learning for Long-Horizon Agentic Tasks
by: Shao, Jie-Jing, et al.
Published: (2026)
by: Shao, Jie-Jing, et al.
Published: (2026)
Training Versatile Coding Agents in Synthetic Environments
by: Zhu, Yiqi, et al.
Published: (2025)
by: Zhu, Yiqi, et al.
Published: (2025)
Programmatic Context Augmentation for LLM-based Symbolic Regression
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
by: Zhang, Qiran, et al.
Published: (2026)
by: Zhang, Qiran, et al.
Published: (2026)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
by: Helu, Zhi, et al.
Published: (2025)
by: Helu, Zhi, et al.
Published: (2025)
ACQ: A Unified Framework for Automated Programmatic Creativity in Online Advertising
by: Wang, Ruizhi, et al.
Published: (2024)
by: Wang, Ruizhi, et al.
Published: (2024)
SPELL: Synthesis of Programmatic Edits using LLMs
by: Ramos, Daniel, et al.
Published: (2026)
by: Ramos, Daniel, et al.
Published: (2026)
DeepThink3D: Enhancing Large Language Models with Programmatic Reasoning in Complex 3D Situated Reasoning Tasks
by: Song, Jiayi, et al.
Published: (2025)
by: Song, Jiayi, et al.
Published: (2025)
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
by: Zhou, Xuhui, et al.
Published: (2026)
by: Zhou, Xuhui, et al.
Published: (2026)
CADSmith: Multi-Agent CAD Generation with Programmatic Geometric Validation
by: Barkley, Jesse, et al.
Published: (2026)
by: Barkley, Jesse, et al.
Published: (2026)
From Solving to Verifying: A Unified Objective for Robust Reasoning in LLMs
by: Wang, Xiaoxuan, et al.
Published: (2025)
by: Wang, Xiaoxuan, et al.
Published: (2025)
Fuse, Reason and Verify: Geometry Problem Solving with Parsed Clauses from Diagram
by: Zhang, Ming-Liang, et al.
Published: (2024)
by: Zhang, Ming-Liang, et al.
Published: (2024)
LOGIGEN: Logic-Driven Generation of Verifiable Agentic Tasks
by: Zeng, Yucheng, et al.
Published: (2026)
by: Zeng, Yucheng, et al.
Published: (2026)
NeuroWeaver: An Autonomous Evolutionary Agent for Exploring the Programmatic Space of EEG Analysis Pipelines
by: Wang, Guoan, et al.
Published: (2026)
by: Wang, Guoan, et al.
Published: (2026)
Language Models For Generalised PDDL Planning: Synthesising Sound and Programmatic Policies
by: Chen, Dillon Z., et al.
Published: (2025)
by: Chen, Dillon Z., et al.
Published: (2025)
Similar Items
-
A Compute-Matched Re-Evaluation of TroVE on MATH
by: Sesterhenn, Tobias, et al.
Published: (2025) -
Inducing Programmatic Skills for Agentic Tasks
by: Wang, Zora Zhiruo, et al.
Published: (2025) -
What Are Tools Anyway? A Survey from the Language Model Perspective
by: Wang, Zhiruo, et al.
Published: (2024) -
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
by: Wang, Zora Zhiruo, et al.
Published: (2025) -
How Well Does Agent Development Reflect Real-World Work?
by: Wang, Zora Zhiruo, et al.
Published: (2026)