RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
Fuente:
arXiv
Saved in:
| Main Authors: | LeVine, Will, Evers, Brendan, Saltwick, Sam, Venkatesh, Abhay |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
It's LIT! Reliability-Optimized LLMs with Inspectable Tools
by: Zhang, Ruixin, et al.
Published: (2025)
by: Zhang, Ruixin, et al.
Published: (2025)
Refining Joint Text and Source Code Embeddings for Retrieval Task with Parameter-Efficient Fine-Tuning
by: Galliamov, Karim, et al.
Published: (2024)
by: Galliamov, Karim, et al.
Published: (2024)
CellScientist: Dual-Space Hierarchical Orchestration for Closed-Loop Refinement of Virtual Cell Models
by: Li, Mengran, et al.
Published: (2026)
by: Li, Mengran, et al.
Published: (2026)
Towards Refining Developer Questions using LLM-Based Named Entity Recognition for Developer Chatroom Conversations
by: Fathollahzadeh, Pouya, et al.
Published: (2025)
by: Fathollahzadeh, Pouya, et al.
Published: (2025)
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
by: Huang, Jiawei, et al.
Published: (2026)
by: Huang, Jiawei, et al.
Published: (2026)
iML: Executable, Problem-Grounded, and Broadly Exploratory Code-Driven AutoML
by: Le, Dat, et al.
Published: (2026)
by: Le, Dat, et al.
Published: (2026)
Pre-Training Representations of Binary Code Using Contrastive Learning
by: Zhang, Yifan, et al.
Published: (2022)
by: Zhang, Yifan, et al.
Published: (2022)
Impact of ML Optimization Tactics on Greener Pre-Trained ML Models
by: Álvarez, Alexandra González, et al.
Published: (2024)
by: Álvarez, Alexandra González, et al.
Published: (2024)
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
by: Jin, Yiyang, et al.
Published: (2025)
by: Jin, Yiyang, et al.
Published: (2025)
Recommending Pre-Trained Models for IoT Devices
by: Patil, Parth V., et al.
Published: (2024)
by: Patil, Parth V., et al.
Published: (2024)
Refining GPT-3 Embeddings with a Siamese Structure for Technical Post Duplicate Detection
by: Wu, Xingfang, et al.
Published: (2023)
by: Wu, Xingfang, et al.
Published: (2023)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
by: Kuntz, Thomas, et al.
Published: (2025)
by: Kuntz, Thomas, et al.
Published: (2025)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
by: Choi, Hyeonje, et al.
Published: (2026)
by: Choi, Hyeonje, et al.
Published: (2026)
The Dual-State Architecture for Reliable LLM Agents
by: Thompson, Matthew
Published: (2025)
by: Thompson, Matthew
Published: (2025)
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
by: Castellani, Tommaso, et al.
Published: (2025)
by: Castellani, Tommaso, et al.
Published: (2025)
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
by: LeVine, Will, et al.
Published: (2025)
by: LeVine, Will, et al.
Published: (2025)
Good Tools are Half the Work: Tool Usage in Deep Learning Projects
by: Panourgia, Evangelia, et al.
Published: (2023)
by: Panourgia, Evangelia, et al.
Published: (2023)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
by: Haque, Mirazul, et al.
Published: (2025)
by: Haque, Mirazul, et al.
Published: (2025)
SMARTCAL: An Approach to Self-Aware Tool-Use Evaluation and Calibration
by: Shen, Yuanhao, et al.
Published: (2024)
by: Shen, Yuanhao, et al.
Published: (2024)
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
by: Yu, Shasha, et al.
Published: (2026)
by: Yu, Shasha, et al.
Published: (2026)
Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap
by: Kim, Naryeong, et al.
Published: (2026)
by: Kim, Naryeong, et al.
Published: (2026)
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
by: Xi, Haoran, et al.
Published: (2025)
by: Xi, Haoran, et al.
Published: (2025)
Model Cascading for Code: A Cascaded Black-Box Multi-Model Framework for Cost-Efficient Code Completion with Self-Testing
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026)
by: Rank, Ben, et al.
Published: (2026)
Adaptive Reinforcement Learning for Dynamic Configuration Allocation in Pre-Production Testing
by: Zhu, Yu
Published: (2025)
by: Zhu, Yu
Published: (2025)
Signature in Code Backdoor Detection, how far are we?
by: Le, Quoc Hung, et al.
Published: (2025)
by: Le, Quoc Hung, et al.
Published: (2025)
Towards Continuous Assurance Case Creation for ADS with the Evidential Tool Bus
by: Sorokin, Lev, et al.
Published: (2024)
by: Sorokin, Lev, et al.
Published: (2024)
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
by: Majdinasab, Vahid, et al.
Published: (2024)
by: Majdinasab, Vahid, et al.
Published: (2024)
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
by: Yoon, Juyeon, et al.
Published: (2025)
by: Yoon, Juyeon, et al.
Published: (2025)
Tracing Stereotypes in Pre-trained Transformers: From Biased Neurons to Fairer Models
by: Voria, Gianmario, et al.
Published: (2026)
by: Voria, Gianmario, et al.
Published: (2026)
A Machine Learning-Based Error Mitigation Approach For Reliable Software Development On IBM'S Quantum Computers
by: Muqeet, Asmar, et al.
Published: (2024)
by: Muqeet, Asmar, et al.
Published: (2024)
Ensuring Reliability in Programming Knowledge Tracing: A Re-evaluation of Attention-augmented Models and Experimental Protocols
by: Kim, Jaewook, et al.
Published: (2026)
by: Kim, Jaewook, et al.
Published: (2026)
Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
Graph-Free Root Cause Analysis
by: Pham, Luan
Published: (2026)
by: Pham, Luan
Published: (2026)
LEANCODE: Understanding Models Better for Code Simplification of Pre-trained Large Language Models
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code
by: Weyssow, Martin, et al.
Published: (2023)
by: Weyssow, Martin, et al.
Published: (2023)
Towards MLOps: A DevOps Tools Recommender System for Machine Learning System
by: Shah, Pir Sami Ullah, et al.
Published: (2024)
by: Shah, Pir Sami Ullah, et al.
Published: (2024)
DesCartes Builder: A Tool to Develop Machine-Learning Based Digital Twins
by: de Conto, Eduardo, et al.
Published: (2025)
by: de Conto, Eduardo, et al.
Published: (2025)
You Only Train Once: A Flexible Training Framework for Code Vulnerability Detection Driven by Vul-Vector
by: Tian, Bowen, et al.
Published: (2025)
by: Tian, Bowen, et al.
Published: (2025)
Similar Items
-
It's LIT! Reliability-Optimized LLMs with Inspectable Tools
by: Zhang, Ruixin, et al.
Published: (2025) -
Refining Joint Text and Source Code Embeddings for Retrieval Task with Parameter-Efficient Fine-Tuning
by: Galliamov, Karim, et al.
Published: (2024) -
CellScientist: Dual-Space Hierarchical Orchestration for Closed-Loop Refinement of Virtual Cell Models
by: Li, Mengran, et al.
Published: (2026) -
Towards Refining Developer Questions using LLM-Based Named Entity Recognition for Developer Chatroom Conversations
by: Fathollahzadeh, Pouya, et al.
Published: (2025) -
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
by: Huang, Jiawei, et al.
Published: (2026)