Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
Fuente:
arXiv
Guardado en:
| Autores principales: | Kohl, Jens, Gloger, Luisa, Costa, Rui, Kruse, Otto, Luitz, Manuel P., Katz, David, Barbeito, Gonzalo, Schweier, Markus, French, Ryan, Schroeder, Jonas, Riedl, Thomas, Perri, Raphael, Mostafa, Youssef |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Automated structural testing of LLM-based agents: methods, framework, and case studies
por: Kohl, Jens, et al.
Publicado: (2026)
por: Kohl, Jens, et al.
Publicado: (2026)
Emergent Coordination in Multi-Agent Language Models
por: Riedl, Christoph
Publicado: (2025)
por: Riedl, Christoph
Publicado: (2025)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
por: Tang, Wenjie, et al.
Publicado: (2026)
por: Tang, Wenjie, et al.
Publicado: (2026)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
por: Wang, Yuchen, et al.
Publicado: (2026)
por: Wang, Yuchen, et al.
Publicado: (2026)
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
por: Costa, Igor
Publicado: (2026)
por: Costa, Igor
Publicado: (2026)
Semantic Risk-Aware Heuristic Planning for Robotic Navigation in Dynamic Environments: An LLM-Inspired Approach
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
FORMICA: Decision-Focused Learning for Communication-Free Multi-Robot Task Allocation
por: Lopez, Antonio, et al.
Publicado: (2026)
por: Lopez, Antonio, et al.
Publicado: (2026)
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
por: Ray, Aninda
Publicado: (2026)
por: Ray, Aninda
Publicado: (2026)
PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and Convergent Feedback
por: Annapureddy, Sasank
Publicado: (2026)
por: Annapureddy, Sasank
Publicado: (2026)
FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory
por: Gu, Yingjie, et al.
Publicado: (2026)
por: Gu, Yingjie, et al.
Publicado: (2026)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
por: Hu, Pan
Publicado: (2025)
por: Hu, Pan
Publicado: (2025)
Exploring Robust Multi-Agent Workflows for Environmental Data Management
por: Guan, Boyuan, et al.
Publicado: (2026)
por: Guan, Boyuan, et al.
Publicado: (2026)
MOMA-AC: A preference-driven actor-critic framework for continuous multi-objective multi-agent reinforcement learning
por: Callaghan, Adam, et al.
Publicado: (2025)
por: Callaghan, Adam, et al.
Publicado: (2025)
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
por: Costa, Rimom
Publicado: (2025)
por: Costa, Rimom
Publicado: (2025)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
por: Wang, Xiaohua, et al.
Publicado: (2026)
por: Wang, Xiaohua, et al.
Publicado: (2026)
Latent Cache Flow: Model-to-Model Communication Without Text
por: Rossi, Maximillian, et al.
Publicado: (2026)
por: Rossi, Maximillian, et al.
Publicado: (2026)
A Super-Learner with Large Language Models for Medical Emergency Advising
por: Aityan, Sergey K., et al.
Publicado: (2025)
por: Aityan, Sergey K., et al.
Publicado: (2025)
PRISM-Consult: A Panel-of-Experts Architecture for Clinician-Aligned Diagnosis
por: Levine, Lionel, et al.
Publicado: (2025)
por: Levine, Lionel, et al.
Publicado: (2025)
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
por: Hong, Yoosung
Publicado: (2026)
por: Hong, Yoosung
Publicado: (2026)
Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory
por: Jiang, Rongjie, et al.
Publicado: (2026)
por: Jiang, Rongjie, et al.
Publicado: (2026)
Applying Cognitive Design Patterns to General LLM Agents
por: Wray, Robert E., et al.
Publicado: (2025)
por: Wray, Robert E., et al.
Publicado: (2025)
Privacy Preserving Multi Agent Path Finding
por: Lehman, Rotem Lev, et al.
Publicado: (2026)
por: Lehman, Rotem Lev, et al.
Publicado: (2026)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
por: Wang, Xiaohua, et al.
Publicado: (2026)
por: Wang, Xiaohua, et al.
Publicado: (2026)
Collaborative On-Sensor Array Cameras
por: Sun, Jipeng, et al.
Publicado: (2025)
por: Sun, Jipeng, et al.
Publicado: (2025)
Learning To Help: Training Models to Assist Legacy Devices
por: Wu, Yu, et al.
Publicado: (2024)
por: Wu, Yu, et al.
Publicado: (2024)
ScrapMem: A Bio-inspired Framework for On-device Personalized Agent Memory via Optical Forgetting
por: Chang, Jiale, et al.
Publicado: (2026)
por: Chang, Jiale, et al.
Publicado: (2026)
SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing
por: Meng, Zi, et al.
Publicado: (2026)
por: Meng, Zi, et al.
Publicado: (2026)
Right-to-Act: A Pre-Execution Non-Compensatory Decision Protocol for AI Systems
por: Lavi, Gadi
Publicado: (2026)
por: Lavi, Gadi
Publicado: (2026)
Post Hoc Extraction of Pareto Fronts for Continuous Control
por: Thakar, Raghav, et al.
Publicado: (2026)
por: Thakar, Raghav, et al.
Publicado: (2026)
Event-Triggered Adaptive Consensus for Multi-Robot Task Allocation
por: Aznar, Fidel, et al.
Publicado: (2026)
por: Aznar, Fidel, et al.
Publicado: (2026)
ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning
por: Chen, Wen-Tse, et al.
Publicado: (2024)
por: Chen, Wen-Tse, et al.
Publicado: (2024)
Robust and Diverse Multi-Agent Learning via Rational Policy Gradient
por: Lauffer, Niklas, et al.
Publicado: (2025)
por: Lauffer, Niklas, et al.
Publicado: (2025)
Building Large-Scale Drone Defenses from Small-Team Strategies
por: Douglas, Grant, et al.
Publicado: (2026)
por: Douglas, Grant, et al.
Publicado: (2026)
Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems
por: Wang, Zixu, et al.
Publicado: (2026)
por: Wang, Zixu, et al.
Publicado: (2026)
Toward Constraint Compliant Goal Formulation and Planning
por: Jones, Steven J., et al.
Publicado: (2024)
por: Jones, Steven J., et al.
Publicado: (2024)
MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents
por: Sidik, Bronislav, et al.
Publicado: (2026)
por: Sidik, Bronislav, et al.
Publicado: (2026)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
por: Li, Bowen, et al.
Publicado: (2026)
por: Li, Bowen, et al.
Publicado: (2026)
Analysing Factorizations of Action-Value Networks for Cooperative Multi-Agent Reinforcement Learning
por: Castellini, Jacopo, et al.
Publicado: (2019)
por: Castellini, Jacopo, et al.
Publicado: (2019)
ChromaFlow: A Negative Ablation Study of Orchestration Overhead in Tool-Augmented Agent Evaluation
por: Mittal, Tarun
Publicado: (2026)
por: Mittal, Tarun
Publicado: (2026)
StatePlane: A Cognitive State Plane for Long-Horizon AI Systems Under Bounded Context
por: Annapureddy, Sasank, et al.
Publicado: (2026)
por: Annapureddy, Sasank, et al.
Publicado: (2026)
Ejemplares similares
-
Automated structural testing of LLM-based agents: methods, framework, and case studies
por: Kohl, Jens, et al.
Publicado: (2026) -
Emergent Coordination in Multi-Agent Language Models
por: Riedl, Christoph
Publicado: (2025) -
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
por: Tang, Wenjie, et al.
Publicado: (2026) -
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
por: Wang, Yuchen, et al.
Publicado: (2026) -
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
por: Costa, Igor
Publicado: (2026)