The Dual-State Architecture for Reliable LLM Agents
Fuente:
arXiv
Guardado en:
| Autor principal: | Thompson, Matthew |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
por: Rank, Ben, et al.
Publicado: (2026)
por: Rank, Ben, et al.
Publicado: (2026)
Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production
por: Fehlis, Yao, et al.
Publicado: (2026)
por: Fehlis, Yao, et al.
Publicado: (2026)
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
por: Yu, Shasha, et al.
Publicado: (2026)
por: Yu, Shasha, et al.
Publicado: (2026)
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
por: Manglik, Akshay, et al.
Publicado: (2026)
por: Manglik, Akshay, et al.
Publicado: (2026)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
por: Han, Xiaoke, et al.
Publicado: (2025)
por: Han, Xiaoke, et al.
Publicado: (2025)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
por: Xiao, Yijia, et al.
Publicado: (2025)
por: Xiao, Yijia, et al.
Publicado: (2025)
EGSS: Entropy-guided Stepwise Scaling for Reliable Software Engineering
por: Mao, Chenhui, et al.
Publicado: (2026)
por: Mao, Chenhui, et al.
Publicado: (2026)
Maestro: Joint Graph & Config Optimization for Reliable AI Agents
por: Wang, Wenxiao, et al.
Publicado: (2025)
por: Wang, Wenxiao, et al.
Publicado: (2025)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
por: Zhou, Zenghui, et al.
Publicado: (2026)
por: Zhou, Zenghui, et al.
Publicado: (2026)
Cleaning Maintenance Logs with LLM Agents for Improved Predictive Maintenance
por: Dimidov, Valeriu, et al.
Publicado: (2025)
por: Dimidov, Valeriu, et al.
Publicado: (2025)
Can Coding Agents Be General Agents?
por: Ivanov, Maksim, et al.
Publicado: (2026)
por: Ivanov, Maksim, et al.
Publicado: (2026)
A Reference Architecture of Reinforcement Learning Frameworks
por: Liu, Xiaoran, et al.
Publicado: (2026)
por: Liu, Xiaoran, et al.
Publicado: (2026)
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
por: Zhou, Yongxi, et al.
Publicado: (2026)
por: Zhou, Yongxi, et al.
Publicado: (2026)
Advancing Software Security and Reliability in Cloud Platforms through AI-based Anomaly Detection
por: Saleh, Sabbir M., et al.
Publicado: (2024)
por: Saleh, Sabbir M., et al.
Publicado: (2024)
ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization
por: Lian, Junbo Jacob, et al.
Publicado: (2026)
por: Lian, Junbo Jacob, et al.
Publicado: (2026)
DRAFT-ing Architectural Design Decisions using LLMs
por: Dhar, Rudra, et al.
Publicado: (2025)
por: Dhar, Rudra, et al.
Publicado: (2025)
Can LLMs Generate Architectural Design Decisions? -An Exploratory Empirical study
por: Dhar, Rudra, et al.
Publicado: (2024)
por: Dhar, Rudra, et al.
Publicado: (2024)
QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture
por: Prakash, Shvetank, et al.
Publicado: (2025)
por: Prakash, Shvetank, et al.
Publicado: (2025)
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
por: Jiralerspong, Thomas, et al.
Publicado: (2026)
por: Jiralerspong, Thomas, et al.
Publicado: (2026)
AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
por: Wang, Zhaohui Geoffrey
Publicado: (2026)
por: Wang, Zhaohui Geoffrey
Publicado: (2026)
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
por: Bisztray, Tamas, et al.
Publicado: (2025)
por: Bisztray, Tamas, et al.
Publicado: (2025)
Predicting Open Source Software Sustainability with Deep Temporal Neural Hierarchical Architectures and Explainable AI
por: Karim, S M Rakib Ul, et al.
Publicado: (2026)
por: Karim, S M Rakib Ul, et al.
Publicado: (2026)
Understanding LLM-Driven Test Oracle Generation
por: Bodicoat, Adam, et al.
Publicado: (2026)
por: Bodicoat, Adam, et al.
Publicado: (2026)
A Survey on Code Generation with LLM-based Agents
por: Dong, Yihong, et al.
Publicado: (2025)
por: Dong, Yihong, et al.
Publicado: (2025)
Agentless: Demystifying LLM-based Software Engineering Agents
por: Xia, Chunqiu Steven, et al.
Publicado: (2024)
por: Xia, Chunqiu Steven, et al.
Publicado: (2024)
The BrowserGym Ecosystem for Web Agent Research
por: De Chezelles, Thibault Le Sellier, et al.
Publicado: (2024)
por: De Chezelles, Thibault Le Sellier, et al.
Publicado: (2024)
TriagerX: Dual Transformers for Bug Triaging Tasks with Content and Interaction Based Rankings
por: Mamun, Md Afif Al, et al.
Publicado: (2025)
por: Mamun, Md Afif Al, et al.
Publicado: (2025)
Mutation-Guided LLM-based Test Generation at Meta
por: Foster, Christopher, et al.
Publicado: (2025)
por: Foster, Christopher, et al.
Publicado: (2025)
SWE-Bench-CL: Continual Learning for Coding Agents
por: Joshi, Thomas, et al.
Publicado: (2025)
por: Joshi, Thomas, et al.
Publicado: (2025)
Automated Cloud Infrastructure-as-Code Reconciliation with AI Agents
por: Yang, Zhenning, et al.
Publicado: (2025)
por: Yang, Zhenning, et al.
Publicado: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
por: Xu, WeiZhe, et al.
Publicado: (2026)
por: Xu, WeiZhe, et al.
Publicado: (2026)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
por: Jiang, Shan, et al.
Publicado: (2026)
por: Jiang, Shan, et al.
Publicado: (2026)
AgentOpt v0.1 Technical Report: Client-Side Optimization for LLM-Based Agent
por: Hua, Wenyue, et al.
Publicado: (2026)
por: Hua, Wenyue, et al.
Publicado: (2026)
Engineering LLM Powered Multi-agent Framework for Autonomous CloudOps
por: Parthasarathy, Kannan, et al.
Publicado: (2025)
por: Parthasarathy, Kannan, et al.
Publicado: (2025)
Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family Classification
por: Wu, Zehao, et al.
Publicado: (2025)
por: Wu, Zehao, et al.
Publicado: (2025)
A Regression Framework for Understanding Prompt Component Impact on LLM Performance
por: Lauziere, Andrew, et al.
Publicado: (2026)
por: Lauziere, Andrew, et al.
Publicado: (2026)
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
por: Castellani, Tommaso, et al.
Publicado: (2025)
por: Castellani, Tommaso, et al.
Publicado: (2025)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
por: Shah, Mehil B, et al.
Publicado: (2025)
por: Shah, Mehil B, et al.
Publicado: (2025)
SMARLA: A Safety Monitoring Approach for Deep Reinforcement Learning Agents
por: Zolfagharian, Amirhossein, et al.
Publicado: (2023)
por: Zolfagharian, Amirhossein, et al.
Publicado: (2023)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
por: Ding, Yifeng, et al.
Publicado: (2026)
por: Ding, Yifeng, et al.
Publicado: (2026)
Ejemplares similares
-
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
por: Rank, Ben, et al.
Publicado: (2026) -
Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production
por: Fehlis, Yao, et al.
Publicado: (2026) -
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
por: Yu, Shasha, et al.
Publicado: (2026) -
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
por: Manglik, Akshay, et al.
Publicado: (2026) -
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
por: Han, Xiaoke, et al.
Publicado: (2025)