Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Jinhu, Li, Yifan, Zhao, Minghao, Zhang, Wentao, Zhang, Zijian, Li, Yaoman, King, Irwin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
by: Li, Bowen, et al.
Published: (2026)
by: Li, Bowen, et al.
Published: (2026)
Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems
by: Wang, Zixu, et al.
Published: (2026)
by: Wang, Zixu, et al.
Published: (2026)
MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing
by: Chen, Han, et al.
Published: (2026)
by: Chen, Han, et al.
Published: (2026)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
by: Wang, Yuchen, et al.
Published: (2026)
by: Wang, Yuchen, et al.
Published: (2026)
MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
by: Costa, Rimom
Published: (2025)
by: Costa, Rimom
Published: (2025)
Tool-RoCo: An Agent-as-Tool Self-organization Large Language Model Benchmark in Multi-robot Cooperation
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems
by: Lan, Guangchen
Published: (2026)
by: Lan, Guangchen
Published: (2026)
CPEMH: An Agentic Framework for Prompt-Driven Behavior Evaluation and Assurance in Foundation-Model Systems for Mental Health Screening
by: Lorenzoni, Giuliano, et al.
Published: (2026)
by: Lorenzoni, Giuliano, et al.
Published: (2026)
Applying Cognitive Design Patterns to General LLM Agents
by: Wray, Robert E., et al.
Published: (2025)
by: Wray, Robert E., et al.
Published: (2025)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
ScrapMem: A Bio-inspired Framework for On-device Personalized Agent Memory via Optical Forgetting
by: Chang, Jiale, et al.
Published: (2026)
by: Chang, Jiale, et al.
Published: (2026)
Post Hoc Extraction of Pareto Fronts for Continuous Control
by: Thakar, Raghav, et al.
Published: (2026)
by: Thakar, Raghav, et al.
Published: (2026)
AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents
by: Tupe, Vaibhav, et al.
Published: (2025)
by: Tupe, Vaibhav, et al.
Published: (2025)
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
by: Qi, Jinhu, et al.
Published: (2026)
by: Qi, Jinhu, et al.
Published: (2026)
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
by: Ray, Aninda
Published: (2026)
by: Ray, Aninda
Published: (2026)
PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and Convergent Feedback
by: Annapureddy, Sasank
Published: (2026)
by: Annapureddy, Sasank
Published: (2026)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Latent Cache Flow: Model-to-Model Communication Without Text
by: Rossi, Maximillian, et al.
Published: (2026)
by: Rossi, Maximillian, et al.
Published: (2026)
Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
by: Gokdemir, Ozan, et al.
Published: (2025)
by: Gokdemir, Ozan, et al.
Published: (2025)
RefiningGPT: Specialized language Models for Automated Refinery Unit-level Process Diagram Synthesis
by: Liu, Dongxiao, et al.
Published: (2026)
by: Liu, Dongxiao, et al.
Published: (2026)
Agent WARPP: Workflow Adherence via Runtime Parallel Personalization
by: Mazzolenis, Maria Emilia, et al.
Published: (2025)
by: Mazzolenis, Maria Emilia, et al.
Published: (2025)
Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
by: Saini, Mayank, et al.
Published: (2025)
by: Saini, Mayank, et al.
Published: (2025)
Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
by: Usman, Rana Muhammad
Published: (2026)
by: Usman, Rana Muhammad
Published: (2026)
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
by: Kohl, Jens, et al.
Published: (2024)
by: Kohl, Jens, et al.
Published: (2024)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
by: Wu, Shuai, et al.
Published: (2026)
by: Wu, Shuai, et al.
Published: (2026)
Agentic Automation of BT-RADS Scoring: End-to-End Multi-Agent System for Standardized Brain Tumor Follow-up Assessment
by: Jabal, Mohamed Sobhi, et al.
Published: (2026)
by: Jabal, Mohamed Sobhi, et al.
Published: (2026)
MMiC: Mitigating Modality Incompleteness in Clustered Federated Learning
by: Yang, Lishan, et al.
Published: (2025)
by: Yang, Lishan, et al.
Published: (2025)
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
by: Huang, Yunpeng, et al.
Published: (2023)
by: Huang, Yunpeng, et al.
Published: (2023)
LLM Scalability Risk for Agentic-AI and Model Supply Chain Security
by: Ahi, Kiarash, et al.
Published: (2026)
by: Ahi, Kiarash, et al.
Published: (2026)
A Scalable Communication Protocol for Networks of Large Language Models
by: Marro, Samuele, et al.
Published: (2024)
by: Marro, Samuele, et al.
Published: (2024)
elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings
by: Osborne, Philip, et al.
Published: (2025)
by: Osborne, Philip, et al.
Published: (2025)
FediLoRA: Practical Federated Fine-Tuning of Foundation Models Under Missing-Modality Constraints
by: Yang, Lishan, et al.
Published: (2025)
by: Yang, Lishan, et al.
Published: (2025)
Bounded Autonomy for Enterprise AI: Typed Action Contracts and Consumer-Side Execution
by: Sohail, Sarmad, et al.
Published: (2026)
by: Sohail, Sarmad, et al.
Published: (2026)
Agentic AI Translate: An Agentic Translator Prototype for Translation as Communication Design
by: Yamada, Masaru
Published: (2026)
by: Yamada, Masaru
Published: (2026)
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
by: Sellami, Khaled, et al.
Published: (2025)
by: Sellami, Khaled, et al.
Published: (2025)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
by: Ge, Yuxu
Published: (2026)
by: Ge, Yuxu
Published: (2026)
From Multi-Agent Systems and the Semantic Web to Agentic AI: A Unified Narrative of the Web of Agents
by: Petrova, Tatiana, et al.
Published: (2025)
by: Petrova, Tatiana, et al.
Published: (2025)
Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI
by: Pasupuleti, Vinil, et al.
Published: (2026)
by: Pasupuleti, Vinil, et al.
Published: (2026)
CRAwDAD: Causal Reasoning Augmentation with Dual-Agent Debate
by: Vamosi, Finn G., et al.
Published: (2025)
by: Vamosi, Finn G., et al.
Published: (2025)
Similar Items
-
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
by: Li, Bowen, et al.
Published: (2026) -
Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems
by: Wang, Zixu, et al.
Published: (2026) -
MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing
by: Chen, Han, et al.
Published: (2026) -
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
by: Wang, Yuchen, et al.
Published: (2026) -
MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents
by: Sidik, Bronislav, et al.
Published: (2026)