Towards More Standardized AI Evaluation: From Models to Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Filali, Ali El, Bedar, Inès |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
par: AlShikh, Waseem, et autres
Publié: (2025)
par: AlShikh, Waseem, et autres
Publié: (2025)
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
par: Tan, Haoran, et autres
Publié: (2025)
par: Tan, Haoran, et autres
Publié: (2025)
SalamahBench: Toward Standardized Safety Evaluation for Arabic Language Models
par: Abdelnasser, Omar, et autres
Publié: (2026)
par: Abdelnasser, Omar, et autres
Publié: (2026)
Towards a More Inclusive AI: Progress and Perspectives in Large Language Model Training for the Sámi Language
par: Paul, Ronny, et autres
Publié: (2024)
par: Paul, Ronny, et autres
Publié: (2024)
Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications
par: Shu, Raphael, et autres
Publié: (2024)
par: Shu, Raphael, et autres
Publié: (2024)
SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
par: Li, Zekun, et autres
Publié: (2025)
par: Li, Zekun, et autres
Publié: (2025)
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
par: Röttger, Paul, et autres
Publié: (2024)
par: Röttger, Paul, et autres
Publié: (2024)
OLMES: A Standard for Language Model Evaluations
par: Gu, Yuling, et autres
Publié: (2024)
par: Gu, Yuling, et autres
Publié: (2024)
MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models
par: Liu, Zhiwei, et autres
Publié: (2025)
par: Liu, Zhiwei, et autres
Publié: (2025)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
par: Kapoor, Sayash, et autres
Publié: (2025)
par: Kapoor, Sayash, et autres
Publié: (2025)
Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
par: Qian, Cheng, et autres
Publié: (2024)
par: Qian, Cheng, et autres
Publié: (2024)
Holistic Evaluation and Failure Diagnosis of AI Agents
par: Madvil, Netta, et autres
Publié: (2026)
par: Madvil, Netta, et autres
Publié: (2026)
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents
par: Komoravolu, Sameer, et autres
Publié: (2025)
par: Komoravolu, Sameer, et autres
Publié: (2025)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
par: Al-Tawaha, Ahmad, et autres
Publié: (2026)
par: Al-Tawaha, Ahmad, et autres
Publié: (2026)
Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
par: Münker, Simon
Publié: (2025)
par: Münker, Simon
Publié: (2025)
From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLM
par: Luo, Chanyong, et autres
Publié: (2026)
par: Luo, Chanyong, et autres
Publié: (2026)
Towards stable AI systems for Evaluating Arabic Pronunciations
par: Zaatiti, Hadi, et autres
Publié: (2025)
par: Zaatiti, Hadi, et autres
Publié: (2025)
AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
par: Shu, Yiheng, et autres
Publié: (2026)
par: Shu, Yiheng, et autres
Publié: (2026)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
par: Kartik, NVJK, et autres
Publié: (2025)
par: Kartik, NVJK, et autres
Publié: (2025)
Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model
par: Ding, Bowen, et autres
Publié: (2025)
par: Ding, Bowen, et autres
Publié: (2025)
Evaluating Multimodal Generative AI with Korean Educational Standards
par: Park, Sanghee, et autres
Publié: (2025)
par: Park, Sanghee, et autres
Publié: (2025)
Standardizing Longitudinal Radiology Report Evaluation via Large Language Model Annotation
par: Wang, Xinyi, et autres
Publié: (2026)
par: Wang, Xinyi, et autres
Publié: (2026)
From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes
par: Garzón, Rubén, et autres
Publié: (2026)
par: Garzón, Rubén, et autres
Publié: (2026)
SMATCH++: Standardized and Extended Evaluation of Semantic Graphs
par: Opitz, Juri
Publié: (2023)
par: Opitz, Juri
Publié: (2023)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
par: Mekky, Ali, et autres
Publié: (2025)
par: Mekky, Ali, et autres
Publié: (2025)
FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
par: Chen, Haorui, et autres
Publié: (2025)
par: Chen, Haorui, et autres
Publié: (2025)
Towards More Effective Table-to-Text Generation: Assessing In-Context Learning and Self-Evaluation with Open-Source Models
par: Iravani, Sahar, et autres
Publié: (2024)
par: Iravani, Sahar, et autres
Publié: (2024)
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
par: Zhu, Jiachen, et autres
Publié: (2025)
par: Zhu, Jiachen, et autres
Publié: (2025)
RoToR: Towards More Reliable Responses for Order-Invariant Inputs
par: Yoon, Soyoung, et autres
Publié: (2025)
par: Yoon, Soyoung, et autres
Publié: (2025)
Characteristic AI Agents via Large Language Models
par: Wang, Xi, et autres
Publié: (2024)
par: Wang, Xi, et autres
Publié: (2024)
From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes
par: Zhou, Karen, et autres
Publié: (2025)
par: Zhou, Karen, et autres
Publié: (2025)
AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
par: Carro, María Victoria, et autres
Publié: (2025)
par: Carro, María Victoria, et autres
Publié: (2025)
More Agents Improve Math Problem Solving but Adversarial Robustness Gap Persists
par: Alavi, Khashayar, et autres
Publié: (2025)
par: Alavi, Khashayar, et autres
Publié: (2025)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
par: Testini, Irene, et autres
Publié: (2025)
par: Testini, Irene, et autres
Publié: (2025)
Role-Playing Evaluation for Large Language Models
par: Boudouri, Yassine El, et autres
Publié: (2025)
par: Boudouri, Yassine El, et autres
Publié: (2025)
More Agents Is All You Need
par: Li, Junyou, et autres
Publié: (2024)
par: Li, Junyou, et autres
Publié: (2024)
StaICC: Standardized Evaluation for Classification Task in In-context Learning
par: Cho, Hakaze, et autres
Publié: (2025)
par: Cho, Hakaze, et autres
Publié: (2025)
Rosetta Stone at KSAA-RD Shared Task: A Hop From Language Modeling To Word--Definition Alignment
par: ElBakry, Ahmed, et autres
Publié: (2023)
par: ElBakry, Ahmed, et autres
Publié: (2023)
Towards Autonomous Agents: Adaptive-planning, Reasoning, and Acting in Language Models
par: Dutta, Abhishek, et autres
Publié: (2024)
par: Dutta, Abhishek, et autres
Publié: (2024)
ACIArena: Toward Unified Evaluation for Agent Cascading Injection
par: An, Hengyu, et autres
Publié: (2026)
par: An, Hengyu, et autres
Publié: (2026)
Documents similaires
-
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
par: AlShikh, Waseem, et autres
Publié: (2025) -
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
par: Tan, Haoran, et autres
Publié: (2025) -
SalamahBench: Toward Standardized Safety Evaluation for Arabic Language Models
par: Abdelnasser, Omar, et autres
Publié: (2026) -
Towards a More Inclusive AI: Progress and Perspectives in Large Language Model Training for the Sámi Language
par: Paul, Ronny, et autres
Publié: (2024) -
Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications
par: Shu, Raphael, et autres
Publié: (2024)