A Framework for Assessing AI Agent Decisions and Outcomes in AutoML Pipelines
Fuente:
arXiv
Guardado en:
| Autores principales: | Du, Gaoyuan, Ahlawat, Amit, Liu, Xiaoyang, Wu, Jing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CodeTracer: Towards Traceable Agent States
por: Li, Han, et al.
Publicado: (2026)
por: Li, Han, et al.
Publicado: (2026)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
por: Tang, Wenjie, et al.
Publicado: (2026)
por: Tang, Wenjie, et al.
Publicado: (2026)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
por: Wang, Yuchen, et al.
Publicado: (2026)
por: Wang, Yuchen, et al.
Publicado: (2026)
Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI
por: Pasupuleti, Vinil, et al.
Publicado: (2026)
por: Pasupuleti, Vinil, et al.
Publicado: (2026)
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
por: Costa, Rimom
Publicado: (2025)
por: Costa, Rimom
Publicado: (2025)
TML-Bench: Benchmark for Data Science Agents on Tabular ML Tasks
por: Pinchuk, Mykola
Publicado: (2026)
por: Pinchuk, Mykola
Publicado: (2026)
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
por: Maryanskyy, Artem
Publicado: (2026)
por: Maryanskyy, Artem
Publicado: (2026)
Robust and Diverse Multi-Agent Learning via Rational Policy Gradient
por: Lauffer, Niklas, et al.
Publicado: (2025)
por: Lauffer, Niklas, et al.
Publicado: (2025)
Analysing Factorizations of Action-Value Networks for Cooperative Multi-Agent Reinforcement Learning
por: Castellini, Jacopo, et al.
Publicado: (2019)
por: Castellini, Jacopo, et al.
Publicado: (2019)
ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning
por: Chen, Wen-Tse, et al.
Publicado: (2024)
por: Chen, Wen-Tse, et al.
Publicado: (2024)
ChromaFlow: A Negative Ablation Study of Orchestration Overhead in Tool-Augmented Agent Evaluation
por: Mittal, Tarun
Publicado: (2026)
por: Mittal, Tarun
Publicado: (2026)
StatePlane: A Cognitive State Plane for Long-Horizon AI Systems Under Bounded Context
por: Annapureddy, Sasank, et al.
Publicado: (2026)
por: Annapureddy, Sasank, et al.
Publicado: (2026)
When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape
por: Mitchell, Richard Joseph
Publicado: (2026)
por: Mitchell, Richard Joseph
Publicado: (2026)
Knowledge Equivalence in Digital Twins of Intelligent Systems
por: Zhang, Nan, et al.
Publicado: (2022)
por: Zhang, Nan, et al.
Publicado: (2022)
AI-Assisted Engineering Should Track the Epistemic Status and Temporal Validity of Architectural Decisions
por: Gilda, Sankalp, et al.
Publicado: (2026)
por: Gilda, Sankalp, et al.
Publicado: (2026)
Generative AI and the Transformation of Software Development Practices
por: Acharya, Vivek
Publicado: (2025)
por: Acharya, Vivek
Publicado: (2025)
Bounded Autonomy for Enterprise AI: Typed Action Contracts and Consumer-Side Execution
por: Sohail, Sarmad, et al.
Publicado: (2026)
por: Sohail, Sarmad, et al.
Publicado: (2026)
Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory
por: Jiang, Rongjie, et al.
Publicado: (2026)
por: Jiang, Rongjie, et al.
Publicado: (2026)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
por: Rehan, Tzafrir
Publicado: (2026)
por: Rehan, Tzafrir
Publicado: (2026)
AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents
por: Tupe, Vaibhav, et al.
Publicado: (2025)
por: Tupe, Vaibhav, et al.
Publicado: (2025)
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
por: Hong, Yoosung
Publicado: (2026)
por: Hong, Yoosung
Publicado: (2026)
Learning To Help: Training Models to Assist Legacy Devices
por: Wu, Yu, et al.
Publicado: (2024)
por: Wu, Yu, et al.
Publicado: (2024)
BONSAI: A Mixed-Initiative Workspace for Human-AI Co-Development of Visual Analytics Applications
por: Spinner, Thilo, et al.
Publicado: (2026)
por: Spinner, Thilo, et al.
Publicado: (2026)
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
por: Rao, Swanand
Publicado: (2026)
por: Rao, Swanand
Publicado: (2026)
PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
por: Schipper, Olivier, et al.
Publicado: (2025)
por: Schipper, Olivier, et al.
Publicado: (2025)
N-Agent Ad Hoc Teamwork
por: Wang, Caroline, et al.
Publicado: (2024)
por: Wang, Caroline, et al.
Publicado: (2024)
Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents
por: Khatchadourian, Raffi
Publicado: (2026)
por: Khatchadourian, Raffi
Publicado: (2026)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
por: Wu, Shuai, et al.
Publicado: (2026)
por: Wu, Shuai, et al.
Publicado: (2026)
Audience Amplified: Virtual Audiences in Asynchronously Performed AR Theater
por: Kim, You-Jin, et al.
Publicado: (2025)
por: Kim, You-Jin, et al.
Publicado: (2025)
Extending NGU to Multi-Agent RL: A Preliminary Study
por: Hernandez, Juan, et al.
Publicado: (2025)
por: Hernandez, Juan, et al.
Publicado: (2025)
Latent Cache Flow: Model-to-Model Communication Without Text
por: Rossi, Maximillian, et al.
Publicado: (2026)
por: Rossi, Maximillian, et al.
Publicado: (2026)
CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend
por: Meng, Haoming
Publicado: (2026)
por: Meng, Haoming
Publicado: (2026)
FlowSteer: Towards Agents Designing Agentic Workflows via Reinforced Progressive Canvas Editing
por: Zhang, Mingda, et al.
Publicado: (2026)
por: Zhang, Mingda, et al.
Publicado: (2026)
QTypeMix: Enhancing Multi-Agent Cooperative Strategies through Heterogeneous and Homogeneous Value Decomposition
por: Fu, Songchen, et al.
Publicado: (2024)
por: Fu, Songchen, et al.
Publicado: (2024)
Nurture-First Agent Development: Building Domain-Expert AI Agents Through Conversational Knowledge Crystallization
por: Zhang, Linghao
Publicado: (2026)
por: Zhang, Linghao
Publicado: (2026)
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
por: Ray, Aninda
Publicado: (2026)
por: Ray, Aninda
Publicado: (2026)
When Can Human-AI Teams Outperform Individuals? Tight Bounds with Impossibility Guarantees
por: Guo, Dongxin, et al.
Publicado: (2026)
por: Guo, Dongxin, et al.
Publicado: (2026)
The Stochastic Gap: A Markovian Framework for Pre-Deployment Reliability and Oversight-Cost Auditing in Agentic Artificial Intelligence
por: Pal, Biplab, et al.
Publicado: (2026)
por: Pal, Biplab, et al.
Publicado: (2026)
Umwelt Engineering: Designing the Cognitive Worlds of Linguistic Agents
por: Jehu-Appiah, Rodney
Publicado: (2026)
por: Jehu-Appiah, Rodney
Publicado: (2026)
SIA: Self Improving AI with Harness & Weight Updates
por: Hebbar, Prannay, et al.
Publicado: (2026)
por: Hebbar, Prannay, et al.
Publicado: (2026)
Ejemplares similares
-
CodeTracer: Towards Traceable Agent States
por: Li, Han, et al.
Publicado: (2026) -
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
por: Tang, Wenjie, et al.
Publicado: (2026) -
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
por: Wang, Yuchen, et al.
Publicado: (2026) -
Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI
por: Pasupuleti, Vinil, et al.
Publicado: (2026) -
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
por: Costa, Rimom
Publicado: (2025)