AutoHarness: improving LLM agents by automatically synthesizing a code harness
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lou, Xinghua, Lázaro-Gredilla, Miguel, Dedieu, Antoine, Wendelken, Carter, Lehrach, Wolfgang, Murphy, Kevin P. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Improving Transformer World Models for Data-Efficient RL
par: Dedieu, Antoine, et autres
Publié: (2025)
par: Dedieu, Antoine, et autres
Publié: (2025)
DMC-VB: A Benchmark for Representation Learning for Control with Visual Distractors
par: Ortiz, Joseph, et autres
Publié: (2024)
par: Ortiz, Joseph, et autres
Publié: (2024)
Learning Cognitive Maps from Transformer Representations for Efficient Planning in Partially Observed Environments
par: Dedieu, Antoine, et autres
Publié: (2024)
par: Dedieu, Antoine, et autres
Publié: (2024)
Code World Models for General Game Playing
par: Lehrach, Wolfgang, et autres
Publié: (2025)
par: Lehrach, Wolfgang, et autres
Publié: (2025)
Diffusion Model Predictive Control
par: Zhou, Guangyao, et autres
Publié: (2024)
par: Zhou, Guangyao, et autres
Publié: (2024)
Model Predictive Simulation Using Structured Graphical Models and Transformers
par: Lou, Xinghua, et autres
Publié: (2024)
par: Lou, Xinghua, et autres
Publié: (2024)
What type of inference is planning?
par: Lázaro-Gredilla, Miguel, et autres
Publié: (2024)
par: Lázaro-Gredilla, Miguel, et autres
Publié: (2024)
Joint Learning of Hierarchical Neural Options and Abstract World Model
par: Piriyakulkij, Wasu Top, et autres
Publié: (2026)
par: Piriyakulkij, Wasu Top, et autres
Publié: (2026)
Question-Answering Based Summarization of Electronic Health Records using Retrieval Augmented Generation
par: Saba, Walid, et autres
Publié: (2024)
par: Saba, Walid, et autres
Publié: (2024)
LLM-AutoDiff: Auto-Differentiate Any LLM Workflow
par: Yin, Li, et autres
Publié: (2025)
par: Yin, Li, et autres
Publié: (2025)
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
par: Zhang, Xiechi, et autres
Publié: (2025)
par: Zhang, Xiechi, et autres
Publié: (2025)
FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents
par: Jia, Haoxuan, et autres
Publié: (2026)
par: Jia, Haoxuan, et autres
Publié: (2026)
Finding codes on infinite grids automatically
par: Salo, Ville, et autres
Publié: (2023)
par: Salo, Ville, et autres
Publié: (2023)
HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
par: Liu, Yujian, et autres
Publié: (2025)
par: Liu, Yujian, et autres
Publié: (2025)
AutoDCWorkflow: LLM-based Data Cleaning Workflow Auto-Generation and Benchmark
par: Li, Lan, et autres
Publié: (2024)
par: Li, Lan, et autres
Publié: (2024)
Harnessing Deep LLM Participation for Robust Entity Linking
par: Hou, Jiajun, et autres
Publié: (2025)
par: Hou, Jiajun, et autres
Publié: (2025)
AgriLLM: Harnessing Transformers for Farmer Queries
par: Didwania, Krish, et autres
Publié: (2024)
par: Didwania, Krish, et autres
Publié: (2024)
AutoCBT: An Autonomous Multi-agent Framework for Cognitive Behavioral Therapy in Psychological Counseling
par: Xu, Ancheng, et autres
Publié: (2025)
par: Xu, Ancheng, et autres
Publié: (2025)
Two CFG Nahuatl for automatic corpora expansion
par: Guzmán-Landa, Juan-José, et autres
Publié: (2025)
par: Guzmán-Landa, Juan-José, et autres
Publié: (2025)
The evaluation of a code-switched Sepedi-English automatic speech recognition system
par: Phaladi, Amanda, et autres
Publié: (2024)
par: Phaladi, Amanda, et autres
Publié: (2024)
StraGo: Harnessing Strategic Guidance for Prompt Optimization
par: Wu, Yurong, et autres
Publié: (2024)
par: Wu, Yurong, et autres
Publié: (2024)
Harnessing Consistency for Robust Test-Time LLM Ensemble
par: Zeng, Zhichen, et autres
Publié: (2025)
par: Zeng, Zhichen, et autres
Publié: (2025)
AutoChip: Automating HDL Generation Using LLM Feedback
par: Thakur, Shailja, et autres
Publié: (2023)
par: Thakur, Shailja, et autres
Publié: (2023)
AutoScreen-FW: An LLM-based Framework for Resume Screening
par: Xu, Zhelin, et autres
Publié: (2026)
par: Xu, Zhelin, et autres
Publié: (2026)
Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing
par: Srivatsa, KV Aditya, et autres
Publié: (2024)
par: Srivatsa, KV Aditya, et autres
Publié: (2024)
Harnessing Multiple Large Language Models: A Survey on LLM Ensemble
par: Chen, Zhijun, et autres
Publié: (2025)
par: Chen, Zhijun, et autres
Publié: (2025)
Review-LLM: Harnessing Large Language Models for Personalized Review Generation
par: Peng, Qiyao, et autres
Publié: (2024)
par: Peng, Qiyao, et autres
Publié: (2024)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
par: Cao, Hongliu, et autres
Publié: (2025)
par: Cao, Hongliu, et autres
Publié: (2025)
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
par: Yuan, Dong, et autres
Publié: (2024)
par: Yuan, Dong, et autres
Publié: (2024)
Deepfake tweets automatic detection
par: Frej, Adam, et autres
Publié: (2024)
par: Frej, Adam, et autres
Publié: (2024)
Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
par: Boggust, Angie, et autres
Publié: (2025)
par: Boggust, Angie, et autres
Publié: (2025)
HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection
par: Du, Xuefeng, et autres
Publié: (2024)
par: Du, Xuefeng, et autres
Publié: (2024)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
par: Bhattacharjee, Arijit, et autres
Publié: (2025)
par: Bhattacharjee, Arijit, et autres
Publié: (2025)
RExBench: Can coding agents autonomously implement AI research extensions?
par: Edwards, Nicholas, et autres
Publié: (2025)
par: Edwards, Nicholas, et autres
Publié: (2025)
Controlled LLM Decoding via Discrete Auto-regressive Biasing
par: Pynadath, Patrick, et autres
Publié: (2025)
par: Pynadath, Patrick, et autres
Publié: (2025)
AutoContext: Instance-Level Context Learning for LLM Agents
par: Cai, Kuntai, et autres
Publié: (2025)
par: Cai, Kuntai, et autres
Publié: (2025)
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
par: Zhou, Karen, et autres
Publié: (2026)
par: Zhou, Karen, et autres
Publié: (2026)
Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
par: Zhao, Ruochen, et autres
Publié: (2024)
par: Zhao, Ruochen, et autres
Publié: (2024)
PGA-SciRE: Harnessing LLM on Data Augmentation for Enhancing Scientific Relation Extraction
par: Zhou, Yang, et autres
Publié: (2024)
par: Zhou, Yang, et autres
Publié: (2024)
MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning
par: Wang, Hanqing, et autres
Publié: (2024)
par: Wang, Hanqing, et autres
Publié: (2024)
Documents similaires
-
Improving Transformer World Models for Data-Efficient RL
par: Dedieu, Antoine, et autres
Publié: (2025) -
DMC-VB: A Benchmark for Representation Learning for Control with Visual Distractors
par: Ortiz, Joseph, et autres
Publié: (2024) -
Learning Cognitive Maps from Transformer Representations for Efficient Planning in Partially Observed Environments
par: Dedieu, Antoine, et autres
Publié: (2024) -
Code World Models for General Game Playing
par: Lehrach, Wolfgang, et autres
Publié: (2025) -
Diffusion Model Predictive Control
par: Zhou, Guangyao, et autres
Publié: (2024)