FROAV: A Framework for RAG Observation and Agent Verification -- Lowering the Barrier to LLM Agent Research
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Tzu-Hsuan, Kao, Chih-Hsuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Workflow-Level Design Principles for Trustworthy GenAI in Automotive System Engineering
di: Cheng, Chih-Hong, et al.
Pubblicazione: (2026)
di: Cheng, Chih-Hong, et al.
Pubblicazione: (2026)
MooseAgent: A LLM Based Multi-agent Framework for Automating Moose Simulation
di: Zhang, Tao, et al.
Pubblicazione: (2025)
di: Zhang, Tao, et al.
Pubblicazione: (2025)
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
di: Jin, Yiyang, et al.
Pubblicazione: (2025)
di: Jin, Yiyang, et al.
Pubblicazione: (2025)
Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents
di: Cai, Yuandao, et al.
Pubblicazione: (2026)
di: Cai, Yuandao, et al.
Pubblicazione: (2026)
BugGen: A Self-Correcting Multi-Agent LLM Pipeline for Realistic RTL Bug Synthesis
di: Jasper, Surya, et al.
Pubblicazione: (2025)
di: Jasper, Surya, et al.
Pubblicazione: (2025)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
AgentForge: A Flexible Low-Code Platform for Reinforcement Learning Agent Design
di: Junior, Francisco Erivaldo Fernandes, et al.
Pubblicazione: (2024)
di: Junior, Francisco Erivaldo Fernandes, et al.
Pubblicazione: (2024)
Towards Continuous Assurance Case Creation for ADS with the Evidential Tool Bus
di: Sorokin, Lev, et al.
Pubblicazione: (2024)
di: Sorokin, Lev, et al.
Pubblicazione: (2024)
ReCodeAgent: A Multi-Agent Workflow for Language-agnostic Translation and Validation of Large-scale Repositories
di: Ibrahimzada, Ali Reza, et al.
Pubblicazione: (2026)
di: Ibrahimzada, Ali Reza, et al.
Pubblicazione: (2026)
The Dual-State Architecture for Reliable LLM Agents
di: Thompson, Matthew
Pubblicazione: (2025)
di: Thompson, Matthew
Pubblicazione: (2025)
SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
di: Wang, Shuaiqi, et al.
Pubblicazione: (2026)
di: Wang, Shuaiqi, et al.
Pubblicazione: (2026)
Exploring LLM-based Agents for Root Cause Analysis
di: Roy, Devjeet, et al.
Pubblicazione: (2024)
di: Roy, Devjeet, et al.
Pubblicazione: (2024)
Bootstrapping Coding Agents: The Specification Is the Program
di: Monperrus, Martin
Pubblicazione: (2026)
di: Monperrus, Martin
Pubblicazione: (2026)
Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap
di: Kim, Naryeong, et al.
Pubblicazione: (2026)
di: Kim, Naryeong, et al.
Pubblicazione: (2026)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
di: Rank, Ben, et al.
Pubblicazione: (2026)
di: Rank, Ben, et al.
Pubblicazione: (2026)
Breaking Memorization Barriers in LLM Code Fine-Tuning via Information Bottleneck for Improved Generalization
di: Wang, Changsheng, et al.
Pubblicazione: (2025)
di: Wang, Changsheng, et al.
Pubblicazione: (2025)
The BrowserGym Ecosystem for Web Agent Research
di: De Chezelles, Thibault Le Sellier, et al.
Pubblicazione: (2024)
di: De Chezelles, Thibault Le Sellier, et al.
Pubblicazione: (2024)
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
di: Yu, Shasha, et al.
Pubblicazione: (2026)
di: Yu, Shasha, et al.
Pubblicazione: (2026)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
di: Han, Xiaoke, et al.
Pubblicazione: (2025)
di: Han, Xiaoke, et al.
Pubblicazione: (2025)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
di: Kuntz, Thomas, et al.
Pubblicazione: (2025)
di: Kuntz, Thomas, et al.
Pubblicazione: (2025)
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
di: Castellani, Tommaso, et al.
Pubblicazione: (2025)
di: Castellani, Tommaso, et al.
Pubblicazione: (2025)
Agint: Agentic Graph Compilation for Software Engineering Agents
di: Chivukula, Abhi, et al.
Pubblicazione: (2025)
di: Chivukula, Abhi, et al.
Pubblicazione: (2025)
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
di: Manglik, Akshay, et al.
Pubblicazione: (2026)
di: Manglik, Akshay, et al.
Pubblicazione: (2026)
Can Coding Agents Be General Agents?
di: Ivanov, Maksim, et al.
Pubblicazione: (2026)
di: Ivanov, Maksim, et al.
Pubblicazione: (2026)
CA2: Code-Aware Agent for Automated Game Testing
di: Adaikkappan, Valliappan Chidambaram, et al.
Pubblicazione: (2026)
di: Adaikkappan, Valliappan Chidambaram, et al.
Pubblicazione: (2026)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
di: Raghavendra, Mohit, et al.
Pubblicazione: (2026)
di: Raghavendra, Mohit, et al.
Pubblicazione: (2026)
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
di: Xi, Haoran, et al.
Pubblicazione: (2025)
di: Xi, Haoran, et al.
Pubblicazione: (2025)
SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow
di: Bula, Timothy, et al.
Pubblicazione: (2025)
di: Bula, Timothy, et al.
Pubblicazione: (2025)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2025)
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2025)
SEVerA: Verified Synthesis of Self-Evolving Agents
di: Banerjee, Debangshu, et al.
Pubblicazione: (2026)
di: Banerjee, Debangshu, et al.
Pubblicazione: (2026)
Methodological Framework for Quantifying Semantic Test Coverage in RAG Systems
di: Broestl, Noah, et al.
Pubblicazione: (2025)
di: Broestl, Noah, et al.
Pubblicazione: (2025)
RAMBO: Enhancing RAG-based Repository-Level Method Body Completion
di: Bui, Tuan-Dung, et al.
Pubblicazione: (2024)
di: Bui, Tuan-Dung, et al.
Pubblicazione: (2024)
Automating Formal Verification with Reinforcement Learning and Recursive Inference
di: Tan, Max
Pubblicazione: (2026)
di: Tan, Max
Pubblicazione: (2026)
Cerberus: Multi-Agent Reasoning and Coverage-Guided Exploration for Static Detection of Runtime Errors
di: Dhulipala, Hridya, et al.
Pubblicazione: (2025)
di: Dhulipala, Hridya, et al.
Pubblicazione: (2025)
Concept-Guided LLM Agents for Human-AI Safety Codesign
di: Geissler, Florian, et al.
Pubblicazione: (2024)
di: Geissler, Florian, et al.
Pubblicazione: (2024)
Cleaning Maintenance Logs with LLM Agents for Improved Predictive Maintenance
di: Dimidov, Valeriu, et al.
Pubblicazione: (2025)
di: Dimidov, Valeriu, et al.
Pubblicazione: (2025)
A Survey on Code Generation with LLM-based Agents
di: Dong, Yihong, et al.
Pubblicazione: (2025)
di: Dong, Yihong, et al.
Pubblicazione: (2025)
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
di: Wu, Jie JW, et al.
Pubblicazione: (2025)
di: Wu, Jie JW, et al.
Pubblicazione: (2025)
VNN: Verification-Friendly Neural Networks with Hard Robustness Guarantees
di: Baninajjar, Anahita, et al.
Pubblicazione: (2023)
di: Baninajjar, Anahita, et al.
Pubblicazione: (2023)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
di: LeVine, Will, et al.
Pubblicazione: (2026)
di: LeVine, Will, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Workflow-Level Design Principles for Trustworthy GenAI in Automotive System Engineering
di: Cheng, Chih-Hong, et al.
Pubblicazione: (2026) -
MooseAgent: A LLM Based Multi-agent Framework for Automating Moose Simulation
di: Zhang, Tao, et al.
Pubblicazione: (2025) -
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
di: Jin, Yiyang, et al.
Pubblicazione: (2025) -
Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents
di: Cai, Yuandao, et al.
Pubblicazione: (2026) -
BugGen: A Self-Correcting Multi-Agent LLM Pipeline for Realistic RTL Bug Synthesis
di: Jasper, Surya, et al.
Pubblicazione: (2025)