Towards More Standardized AI Evaluation: From Models to Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Filali, Ali El, Bedar, Inès |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
di: AlShikh, Waseem, et al.
Pubblicazione: (2025)
di: AlShikh, Waseem, et al.
Pubblicazione: (2025)
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
di: Tan, Haoran, et al.
Pubblicazione: (2025)
di: Tan, Haoran, et al.
Pubblicazione: (2025)
SalamahBench: Toward Standardized Safety Evaluation for Arabic Language Models
di: Abdelnasser, Omar, et al.
Pubblicazione: (2026)
di: Abdelnasser, Omar, et al.
Pubblicazione: (2026)
Towards a More Inclusive AI: Progress and Perspectives in Large Language Model Training for the Sámi Language
di: Paul, Ronny, et al.
Pubblicazione: (2024)
di: Paul, Ronny, et al.
Pubblicazione: (2024)
Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications
di: Shu, Raphael, et al.
Pubblicazione: (2024)
di: Shu, Raphael, et al.
Pubblicazione: (2024)
SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
di: Li, Zekun, et al.
Pubblicazione: (2025)
di: Li, Zekun, et al.
Pubblicazione: (2025)
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
di: Röttger, Paul, et al.
Pubblicazione: (2024)
di: Röttger, Paul, et al.
Pubblicazione: (2024)
OLMES: A Standard for Language Model Evaluations
di: Gu, Yuling, et al.
Pubblicazione: (2024)
di: Gu, Yuling, et al.
Pubblicazione: (2024)
MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models
di: Liu, Zhiwei, et al.
Pubblicazione: (2025)
di: Liu, Zhiwei, et al.
Pubblicazione: (2025)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
di: Kapoor, Sayash, et al.
Pubblicazione: (2025)
di: Kapoor, Sayash, et al.
Pubblicazione: (2025)
Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
di: Qian, Cheng, et al.
Pubblicazione: (2024)
di: Qian, Cheng, et al.
Pubblicazione: (2024)
Holistic Evaluation and Failure Diagnosis of AI Agents
di: Madvil, Netta, et al.
Pubblicazione: (2026)
di: Madvil, Netta, et al.
Pubblicazione: (2026)
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents
di: Komoravolu, Sameer, et al.
Pubblicazione: (2025)
di: Komoravolu, Sameer, et al.
Pubblicazione: (2025)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
di: Al-Tawaha, Ahmad, et al.
Pubblicazione: (2026)
di: Al-Tawaha, Ahmad, et al.
Pubblicazione: (2026)
Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
di: Münker, Simon
Pubblicazione: (2025)
di: Münker, Simon
Pubblicazione: (2025)
From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLM
di: Luo, Chanyong, et al.
Pubblicazione: (2026)
di: Luo, Chanyong, et al.
Pubblicazione: (2026)
Towards stable AI systems for Evaluating Arabic Pronunciations
di: Zaatiti, Hadi, et al.
Pubblicazione: (2025)
di: Zaatiti, Hadi, et al.
Pubblicazione: (2025)
AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
di: Shu, Yiheng, et al.
Pubblicazione: (2026)
di: Shu, Yiheng, et al.
Pubblicazione: (2026)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
di: Kartik, NVJK, et al.
Pubblicazione: (2025)
di: Kartik, NVJK, et al.
Pubblicazione: (2025)
Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model
di: Ding, Bowen, et al.
Pubblicazione: (2025)
di: Ding, Bowen, et al.
Pubblicazione: (2025)
Evaluating Multimodal Generative AI with Korean Educational Standards
di: Park, Sanghee, et al.
Pubblicazione: (2025)
di: Park, Sanghee, et al.
Pubblicazione: (2025)
Standardizing Longitudinal Radiology Report Evaluation via Large Language Model Annotation
di: Wang, Xinyi, et al.
Pubblicazione: (2026)
di: Wang, Xinyi, et al.
Pubblicazione: (2026)
From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes
di: Garzón, Rubén, et al.
Pubblicazione: (2026)
di: Garzón, Rubén, et al.
Pubblicazione: (2026)
SMATCH++: Standardized and Extended Evaluation of Semantic Graphs
di: Opitz, Juri
Pubblicazione: (2023)
di: Opitz, Juri
Pubblicazione: (2023)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
di: Mekky, Ali, et al.
Pubblicazione: (2025)
di: Mekky, Ali, et al.
Pubblicazione: (2025)
FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
di: Chen, Haorui, et al.
Pubblicazione: (2025)
di: Chen, Haorui, et al.
Pubblicazione: (2025)
Towards More Effective Table-to-Text Generation: Assessing In-Context Learning and Self-Evaluation with Open-Source Models
di: Iravani, Sahar, et al.
Pubblicazione: (2024)
di: Iravani, Sahar, et al.
Pubblicazione: (2024)
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
di: Zhu, Jiachen, et al.
Pubblicazione: (2025)
di: Zhu, Jiachen, et al.
Pubblicazione: (2025)
RoToR: Towards More Reliable Responses for Order-Invariant Inputs
di: Yoon, Soyoung, et al.
Pubblicazione: (2025)
di: Yoon, Soyoung, et al.
Pubblicazione: (2025)
Characteristic AI Agents via Large Language Models
di: Wang, Xi, et al.
Pubblicazione: (2024)
di: Wang, Xi, et al.
Pubblicazione: (2024)
From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes
di: Zhou, Karen, et al.
Pubblicazione: (2025)
di: Zhou, Karen, et al.
Pubblicazione: (2025)
AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
di: Carro, María Victoria, et al.
Pubblicazione: (2025)
di: Carro, María Victoria, et al.
Pubblicazione: (2025)
More Agents Improve Math Problem Solving but Adversarial Robustness Gap Persists
di: Alavi, Khashayar, et al.
Pubblicazione: (2025)
di: Alavi, Khashayar, et al.
Pubblicazione: (2025)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
di: Testini, Irene, et al.
Pubblicazione: (2025)
di: Testini, Irene, et al.
Pubblicazione: (2025)
Role-Playing Evaluation for Large Language Models
di: Boudouri, Yassine El, et al.
Pubblicazione: (2025)
di: Boudouri, Yassine El, et al.
Pubblicazione: (2025)
More Agents Is All You Need
di: Li, Junyou, et al.
Pubblicazione: (2024)
di: Li, Junyou, et al.
Pubblicazione: (2024)
StaICC: Standardized Evaluation for Classification Task in In-context Learning
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Rosetta Stone at KSAA-RD Shared Task: A Hop From Language Modeling To Word--Definition Alignment
di: ElBakry, Ahmed, et al.
Pubblicazione: (2023)
di: ElBakry, Ahmed, et al.
Pubblicazione: (2023)
Towards Autonomous Agents: Adaptive-planning, Reasoning, and Acting in Language Models
di: Dutta, Abhishek, et al.
Pubblicazione: (2024)
di: Dutta, Abhishek, et al.
Pubblicazione: (2024)
ACIArena: Toward Unified Evaluation for Agent Cascading Injection
di: An, Hengyu, et al.
Pubblicazione: (2026)
di: An, Hengyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
di: AlShikh, Waseem, et al.
Pubblicazione: (2025) -
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
di: Tan, Haoran, et al.
Pubblicazione: (2025) -
SalamahBench: Toward Standardized Safety Evaluation for Arabic Language Models
di: Abdelnasser, Omar, et al.
Pubblicazione: (2026) -
Towards a More Inclusive AI: Progress and Perspectives in Large Language Model Training for the Sámi Language
di: Paul, Ronny, et al.
Pubblicazione: (2024) -
Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications
di: Shu, Raphael, et al.
Pubblicazione: (2024)