AgentFly: Extensible and Scalable Reinforcement Learning for LM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Renxi, Genadi, Rifo Ahmad, Bouardi, Bilal El, Wang, Yongxin, Koto, Fajri, Liu, Zhengzhong, Baldwin, Timothy, Li, Haonan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
Neural Theorem Proving for Verification Conditions: A Real-World Benchmark
by: Xu, Qiyuan, et al.
Published: (2026)
by: Xu, Qiyuan, et al.
Published: (2026)
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
by: Wang, Renxi, et al.
Published: (2024)
by: Wang, Renxi, et al.
Published: (2024)
ContextBench: A Benchmark for Context Retrieval in Coding Agents
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
by: Ahmed, Syed Yusuf, et al.
Published: (2026)
by: Ahmed, Syed Yusuf, et al.
Published: (2026)
A Framework for Assessing AI Agent Decisions and Outcomes in AutoML Pipelines
by: Du, Gaoyuan, et al.
Published: (2026)
by: Du, Gaoyuan, et al.
Published: (2026)
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
by: Wu, Jiaqing, et al.
Published: (2026)
by: Wu, Jiaqing, et al.
Published: (2026)
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
by: Li, Kenan, et al.
Published: (2026)
by: Li, Kenan, et al.
Published: (2026)
AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
by: Bhardwaj, Varun Pratap
Published: (2026)
by: Bhardwaj, Varun Pratap
Published: (2026)
Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents
by: Khatchadourian, Raffi
Published: (2026)
by: Khatchadourian, Raffi
Published: (2026)
Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
by: Muzsai, Lajos, et al.
Published: (2025)
by: Muzsai, Lajos, et al.
Published: (2025)
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
by: Haseeb, Muhammad
Published: (2025)
by: Haseeb, Muhammad
Published: (2025)
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
SimuScene: Training and Benchmarking Code Generation to Simulate Physical Scenarios
by: Wang, Yanan, et al.
Published: (2026)
by: Wang, Yanan, et al.
Published: (2026)
elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings
by: Osborne, Philip, et al.
Published: (2025)
by: Osborne, Philip, et al.
Published: (2025)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
by: Li, Meiziniu, et al.
Published: (2024)
by: Li, Meiziniu, et al.
Published: (2024)
Demystifying the Silence of Correctness Bugs in PyTorch Compiler
by: Li, Meiziniu, et al.
Published: (2026)
by: Li, Meiziniu, et al.
Published: (2026)
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
by: Li, Meiziniu, et al.
Published: (2022)
by: Li, Meiziniu, et al.
Published: (2022)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
by: Rehan, Tzafrir
Published: (2026)
by: Rehan, Tzafrir
Published: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend
by: Meng, Haoming
Published: (2026)
by: Meng, Haoming
Published: (2026)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
by: Iscan, Mehmet
Published: (2026)
by: Iscan, Mehmet
Published: (2026)
Deep Generative Symbolic Regression
by: Holt, Samuel, et al.
Published: (2023)
by: Holt, Samuel, et al.
Published: (2023)
Cross-lingual Transfer in Programming Languages: An Extensive Empirical Study
by: Baltaji, Razan, et al.
Published: (2023)
by: Baltaji, Razan, et al.
Published: (2023)
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
by: Gautam, Dhruv, et al.
Published: (2025)
by: Gautam, Dhruv, et al.
Published: (2025)
CodeTracer: Towards Traceable Agent States
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
ToolGen: Unified Tool Retrieval and Calling via Generation
by: Wang, Renxi, et al.
Published: (2024)
by: Wang, Renxi, et al.
Published: (2024)
The LLMbda Calculus: AI Agents, Conversations, and Information Flow
by: Garby, Zac, et al.
Published: (2026)
by: Garby, Zac, et al.
Published: (2026)
Multi-Agent Code Verification via Information Theory
by: Rajan, Shreshth
Published: (2025)
by: Rajan, Shreshth
Published: (2025)
NeuroLog: Reasoning You Can Audit -- Neuro-Symbolic Vulnerability Discovery via LLM Facts, Datalog, and SMT
by: Rawat, Sanjay
Published: (2026)
by: Rawat, Sanjay
Published: (2026)
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
by: Bercovich, Ivan
Published: (2026)
by: Bercovich, Ivan
Published: (2026)
Generating Multidimensional Clusters With Support Lines
by: Fachada, Nuno, et al.
Published: (2023)
by: Fachada, Nuno, et al.
Published: (2023)
Predicting Future Actions of Reinforcement Learning Agents
by: Chung, Stephen, et al.
Published: (2024)
by: Chung, Stephen, et al.
Published: (2024)
On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations
by: Hundal, Rajdeep Singh, et al.
Published: (2025)
by: Hundal, Rajdeep Singh, et al.
Published: (2025)
Scalable Bayesian Clustering for Integrative Analysis of Multi-View Data
by: Cabral, Rafael, et al.
Published: (2024)
by: Cabral, Rafael, et al.
Published: (2024)
RepoAudit: An Autonomous LLM-Agent for Repository-Level Code Auditing
by: Guo, Jinyao, et al.
Published: (2025)
by: Guo, Jinyao, et al.
Published: (2025)
Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery
by: Agarwal, Abhinav
Published: (2026)
by: Agarwal, Abhinav
Published: (2026)
LLMLogAnalyzer: A Clustering-Based Log Analysis Chatbot using Large Language Models
by: Cai, Peng, et al.
Published: (2025)
by: Cai, Peng, et al.
Published: (2025)
CodeEvolve: LLM-Driven Evolutionary Optimization with Runtime-Enriched Target Selection for Multi-Language Code Enhancement
by: Borra, Ajay Krishna, et al.
Published: (2026)
by: Borra, Ajay Krishna, et al.
Published: (2026)
Similar Items
-
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025) -
Neural Theorem Proving for Verification Conditions: A Real-World Benchmark
by: Xu, Qiyuan, et al.
Published: (2026) -
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
by: Wang, Renxi, et al.
Published: (2024) -
ContextBench: A Benchmark for Context Retrieval in Coding Agents
by: Li, Han, et al.
Published: (2026) -
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
by: Ahmed, Syed Yusuf, et al.
Published: (2026)