Consistency Amplifies: How Behavioral Variance Shapes Agent Accuracy
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Mehta, Aman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Architecture Without Architects: How AI Coding Agents Shape Software Architecture
von: Konrad, Phongsakon Mark, et al.
Veröffentlicht: (2026)
von: Konrad, Phongsakon Mark, et al.
Veröffentlicht: (2026)
How does information access affect LLM monitors' ability to detect sabotage?
von: Arike, Rauno, et al.
Veröffentlicht: (2026)
von: Arike, Rauno, et al.
Veröffentlicht: (2026)
CASET: Complexity Analysis using Simple Execution Traces for CS* submissions
von: Mehta, Aaryen, et al.
Veröffentlicht: (2024)
von: Mehta, Aaryen, et al.
Veröffentlicht: (2024)
PARCER as an Operational Contract to Reduce Variance, Cost, and Risk in LLM Systems
von: Filho, Elzo Brito dos Santos
Veröffentlicht: (2026)
von: Filho, Elzo Brito dos Santos
Veröffentlicht: (2026)
How Do Agents Perform Code Optimization? An Empirical Study
von: Peng, Huiyun, et al.
Veröffentlicht: (2025)
von: Peng, Huiyun, et al.
Veröffentlicht: (2025)
Learning Correct Behavior from Examples: Validating Sequential Execution in Autonomous Agents
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2026)
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2026)
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
von: Li, Zeping, et al.
Veröffentlicht: (2026)
von: Li, Zeping, et al.
Veröffentlicht: (2026)
When AI Teammates Meet Code Review: Collaboration Signals Shaping the Integration of Agent-Authored Pull Requests
von: Nachuma, Costain, et al.
Veröffentlicht: (2026)
von: Nachuma, Costain, et al.
Veröffentlicht: (2026)
Same Signal, Different Semantics: A Cross-Framework Behavioral Analysis of Software Engineering Agents
von: Ma, Wei, et al.
Veröffentlicht: (2026)
von: Ma, Wei, et al.
Veröffentlicht: (2026)
The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment
von: Yu, Shasha, et al.
Veröffentlicht: (2026)
von: Yu, Shasha, et al.
Veröffentlicht: (2026)
Accuracy Can Lie: On the Impact of Surrogate Model in Configuration Tuning
von: Chen, Pengzhou, et al.
Veröffentlicht: (2025)
von: Chen, Pengzhou, et al.
Veröffentlicht: (2025)
How AI Coding Agents Communicate: A Study of Pull Request Description Characteristics and Human Review Responses
von: Watanabe, Kan, et al.
Veröffentlicht: (2026)
von: Watanabe, Kan, et al.
Veröffentlicht: (2026)
Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs
von: Hida, Gilberto Sussumu, et al.
Veröffentlicht: (2026)
von: Hida, Gilberto Sussumu, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Characterizing Code Comprehension Capabilities in (Large) Language Models
von: Mächtle, Felix, et al.
Veröffentlicht: (2026)
von: Mächtle, Felix, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
von: Weng, Shihao, et al.
Veröffentlicht: (2026)
von: Weng, Shihao, et al.
Veröffentlicht: (2026)
Bridging Forecast Accuracy and Inventory KPIs: A Simulation-Based Software Framework
von: Fukuhara, So, et al.
Veröffentlicht: (2026)
von: Fukuhara, So, et al.
Veröffentlicht: (2026)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
Beyond Accuracy: An Empirical Study on Unit Testing in Open-source Deep Learning Projects
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation
von: Xianpeng, et al.
Veröffentlicht: (2026)
von: Xianpeng, et al.
Veröffentlicht: (2026)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
von: Li, Ziyu, et al.
Veröffentlicht: (2024)
von: Li, Ziyu, et al.
Veröffentlicht: (2024)
Can Agents Fix Agent Issues?
von: Rahardja, Alfin Wijaya, et al.
Veröffentlicht: (2025)
von: Rahardja, Alfin Wijaya, et al.
Veröffentlicht: (2025)
Agent Behavioral Contracts: Formal Specification and Runtime Enforcement for Reliable Autonomous AI Agents
von: Bhardwaj, Varun Pratap
Veröffentlicht: (2026)
von: Bhardwaj, Varun Pratap
Veröffentlicht: (2026)
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
von: Sharma, Aman, et al.
Veröffentlicht: (2026)
von: Sharma, Aman, et al.
Veröffentlicht: (2026)
Rethinking Scientific Modeling: Toward Physically Consistent and Simulation-Executable Programmatic Generation
von: Jiang, Yongqing, et al.
Veröffentlicht: (2026)
von: Jiang, Yongqing, et al.
Veröffentlicht: (2026)
A Semi-Formal Verification Methodology for Efficient Configuration Coverage of Highly Configurable Digital Designs
von: Kumar, Aman, et al.
Veröffentlicht: (2024)
von: Kumar, Aman, et al.
Veröffentlicht: (2024)
Unify and Triumph: Polyglot, Diverse, and Self-Consistent Generation of Unit Tests with LLMs
von: Khelladi, Djamel Eddine, et al.
Veröffentlicht: (2025)
von: Khelladi, Djamel Eddine, et al.
Veröffentlicht: (2025)
Accurate and Consistent Graph Model Generation from Text with Large Language Models
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
AgentGuard: Runtime Verification of AI Agents
von: Koohestani, Roham
Veröffentlicht: (2025)
von: Koohestani, Roham
Veröffentlicht: (2025)
Enhancing Code Consistency in AI Research with Large Language Models and Retrieval-Augmented Generation
von: Keshri, Rajat, et al.
Veröffentlicht: (2025)
von: Keshri, Rajat, et al.
Veröffentlicht: (2025)
AgentStepper: Interactive Debugging of Software Development Agents
von: Hutter, Robert, et al.
Veröffentlicht: (2026)
von: Hutter, Robert, et al.
Veröffentlicht: (2026)
HAFixAgent: History-Aware Program Repair Agent
von: Shi, Yu, et al.
Veröffentlicht: (2025)
von: Shi, Yu, et al.
Veröffentlicht: (2025)
Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as Tools
von: Agarwal, Prerna, et al.
Veröffentlicht: (2025)
von: Agarwal, Prerna, et al.
Veröffentlicht: (2025)
An Empirical Study of Agent Developer Practices in AI Agent Frameworks
von: Wang, Yanlin, et al.
Veröffentlicht: (2025)
von: Wang, Yanlin, et al.
Veröffentlicht: (2025)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
von: Sahoo, Priyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Priyam, et al.
Veröffentlicht: (2026)
AgentTrace: A Structured Logging Framework for Agent System Observability
von: AlSayyad, Adam, et al.
Veröffentlicht: (2026)
von: AlSayyad, Adam, et al.
Veröffentlicht: (2026)
EmbeWebAgent: Embedding Web Agents into Any Customized UI
von: Ma, Chenyang, et al.
Veröffentlicht: (2026)
von: Ma, Chenyang, et al.
Veröffentlicht: (2026)
SpecAgent: A Speculative Retrieval and Forecasting Agent for Code Completion
von: Ma, George, et al.
Veröffentlicht: (2025)
von: Ma, George, et al.
Veröffentlicht: (2025)
AgentSLA : Towards a Service Level Agreement for AI Agents
von: Jouneaux, Gwendal, et al.
Veröffentlicht: (2025)
von: Jouneaux, Gwendal, et al.
Veröffentlicht: (2025)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025)
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Architecture Without Architects: How AI Coding Agents Shape Software Architecture
von: Konrad, Phongsakon Mark, et al.
Veröffentlicht: (2026) -
How does information access affect LLM monitors' ability to detect sabotage?
von: Arike, Rauno, et al.
Veröffentlicht: (2026) -
CASET: Complexity Analysis using Simple Execution Traces for CS* submissions
von: Mehta, Aaryen, et al.
Veröffentlicht: (2024) -
PARCER as an Operational Contract to Reduce Variance, Cost, and Risk in LLM Systems
von: Filho, Elzo Brito dos Santos
Veröffentlicht: (2026) -
How Do Agents Perform Code Optimization? An Empirical Study
von: Peng, Huiyun, et al.
Veröffentlicht: (2025)