Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Naryeong, Yoo, Shin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Capturing Semantic Flow of ML-based Systems
by: Yoo, Shin, et al.
Published: (2025)
by: Yoo, Shin, et al.
Published: (2025)
Lachesis: Predicting LLM Inference Accuracy using Structural Properties of Reasoning Paths
by: Kim, Naryeong, et al.
Published: (2024)
by: Kim, Naryeong, et al.
Published: (2024)
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
by: Yoon, Juyeon, et al.
Published: (2025)
by: Yoon, Juyeon, et al.
Published: (2025)
DANDI: Diffusion as Normative Distribution for Deep Neural Network Input
by: Kim, Somin, et al.
Published: (2025)
by: Kim, Somin, et al.
Published: (2025)
Identifying Bug Inducing Commits by Combining Fault Localisation and Code Change Histories
by: An, Gabin, et al.
Published: (2025)
by: An, Gabin, et al.
Published: (2025)
COSMosFL: Ensemble of Small Language Models for Fault Localisation
by: Cho, Hyunjoon, et al.
Published: (2025)
by: Cho, Hyunjoon, et al.
Published: (2025)
Predictive Prompt Analysis
by: Lee, Jae Yong, et al.
Published: (2025)
by: Lee, Jae Yong, et al.
Published: (2025)
ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning
by: Jiang, Juyong, et al.
Published: (2026)
by: Jiang, Juyong, et al.
Published: (2026)
BugGen: A Self-Correcting Multi-Agent LLM Pipeline for Realistic RTL Bug Synthesis
by: Jasper, Surya, et al.
Published: (2025)
by: Jasper, Surya, et al.
Published: (2025)
Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI's LLM with Open Source SLMs in Production
by: Irugalbandara, Chandra, et al.
Published: (2023)
by: Irugalbandara, Chandra, et al.
Published: (2023)
Model Cascading for Code: A Cascaded Black-Box Multi-Model Framework for Cost-Efficient Code Completion with Self-Testing
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
by: Jahangirova, Gunel, et al.
Published: (2024)
by: Jahangirova, Gunel, et al.
Published: (2024)
An Empirical Study of Fault Localisation Techniques for Deep Learning
by: Humbatova, Nargiz, et al.
Published: (2024)
by: Humbatova, Nargiz, et al.
Published: (2024)
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
by: Jin, Yiyang, et al.
Published: (2025)
by: Jin, Yiyang, et al.
Published: (2025)
Aligning the Objective of LLM-based Program Repair
by: Xu, Junjielong, et al.
Published: (2024)
by: Xu, Junjielong, et al.
Published: (2024)
Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents
by: Cai, Yuandao, et al.
Published: (2026)
by: Cai, Yuandao, et al.
Published: (2026)
FROAV: A Framework for RAG Observation and Agent Verification -- Lowering the Barrier to LLM Agent Research
by: Lin, Tzu-Hsuan, et al.
Published: (2026)
by: Lin, Tzu-Hsuan, et al.
Published: (2026)
PyTorch-based Geometric Learning with Non-CUDA Processing Units: Experiences from Intel Gaudi-v2 HPUs
by: Bu, Fanchen, et al.
Published: (2025)
by: Bu, Fanchen, et al.
Published: (2025)
EET: Experience-Driven Early Termination for Cost-Efficient Software Engineering Agents
by: Guo, Yaoqi, et al.
Published: (2026)
by: Guo, Yaoqi, et al.
Published: (2026)
Exploring LLM-based Agents for Root Cause Analysis
by: Roy, Devjeet, et al.
Published: (2024)
by: Roy, Devjeet, et al.
Published: (2024)
Cleaning Maintenance Logs with LLM Agents for Improved Predictive Maintenance
by: Dimidov, Valeriu, et al.
Published: (2025)
by: Dimidov, Valeriu, et al.
Published: (2025)
MooseAgent: A LLM Based Multi-agent Framework for Automating Moose Simulation
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
Breaking Memorization Barriers in LLM Code Fine-Tuning via Information Bottleneck for Improved Generalization
by: Wang, Changsheng, et al.
Published: (2025)
by: Wang, Changsheng, et al.
Published: (2025)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
by: LeVine, Will, et al.
Published: (2026)
by: LeVine, Will, et al.
Published: (2026)
EVALOOOP: A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming
by: Fang, Sen, et al.
Published: (2025)
by: Fang, Sen, et al.
Published: (2025)
Ontology- and LLM-based Data Harmonization for Federated Learning in Healthcare
by: Kokash, Natallia, et al.
Published: (2025)
by: Kokash, Natallia, et al.
Published: (2025)
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems
by: Xing, Jun, et al.
Published: (2025)
by: Xing, Jun, et al.
Published: (2025)
Should Code Models Learn Pedagogically? A Preliminary Evaluation of Curriculum Learning for Real-World Software Engineering Tasks
by: Khant, Kyi Shin, et al.
Published: (2025)
by: Khant, Kyi Shin, et al.
Published: (2025)
Beyond pip install: Evaluating LLM Agents for the Automated Installation of Python Projects
by: Milliken, Louis, et al.
Published: (2024)
by: Milliken, Louis, et al.
Published: (2024)
Ensuring Reliability in Programming Knowledge Tracing: A Re-evaluation of Attention-augmented Models and Experimental Protocols
by: Kim, Jaewook, et al.
Published: (2026)
by: Kim, Jaewook, et al.
Published: (2026)
Try with Simpler -- An Evaluation of Improved Principal Component Analysis in Log-based Anomaly Detection
by: Yang, Lin, et al.
Published: (2023)
by: Yang, Lin, et al.
Published: (2023)
CigaR: Cost-efficient Program Repair with LLMs
by: Hidvégi, Dávid, et al.
Published: (2024)
by: Hidvégi, Dávid, et al.
Published: (2024)
Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?
by: Voria, Gianmario, et al.
Published: (2024)
by: Voria, Gianmario, et al.
Published: (2024)
LLM-Guided Runtime Parameter Optimization for Energy-Efficient Model Inference
by: Crumpacker, Katelyn, et al.
Published: (2026)
by: Crumpacker, Katelyn, et al.
Published: (2026)
Pimp My LLM: Leveraging Variability Modeling to Tune Inference Hyperparameters
by: Zine, Nada, et al.
Published: (2026)
by: Zine, Nada, et al.
Published: (2026)
A Survey on LLM-based Code Generation for Low-Resource and Domain-Specific Programming Languages
by: Joel, Sathvik, et al.
Published: (2024)
by: Joel, Sathvik, et al.
Published: (2024)
LLM Critics Help Catch LLM Bugs
by: McAleese, Nat, et al.
Published: (2024)
by: McAleese, Nat, et al.
Published: (2024)
Strategic Heterogeneous Multi-Agent Architecture for Cost-Effective Code Vulnerability Detection
by: Wang, Zhaohui Geoffrey
Published: (2026)
by: Wang, Zhaohui Geoffrey
Published: (2026)
Deep Learning Model Reuse in the HuggingFace Community: Challenges, Benefit and Trends
by: Taraghi, Mina, et al.
Published: (2024)
by: Taraghi, Mina, et al.
Published: (2024)
What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
by: Errica, Federico, et al.
Published: (2024)
by: Errica, Federico, et al.
Published: (2024)
Similar Items
-
Capturing Semantic Flow of ML-based Systems
by: Yoo, Shin, et al.
Published: (2025) -
Lachesis: Predicting LLM Inference Accuracy using Structural Properties of Reasoning Paths
by: Kim, Naryeong, et al.
Published: (2024) -
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
by: Yoon, Juyeon, et al.
Published: (2025) -
DANDI: Diffusion as Normative Distribution for Deep Neural Network Input
by: Kim, Somin, et al.
Published: (2025) -
Identifying Bug Inducing Commits by Combining Fault Localisation and Code Change Histories
by: An, Gabin, et al.
Published: (2025)