GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Jisoo, Chang, Raeyoung, Kwon, Dongwook, Singh, Harmanpreet, Verma, Nikhil |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CascadeDebate: Multi-Agent Deliberation for Cost-Aware LLM Cascades
di: Chang, Raeyoung, et al.
Pubblicazione: (2026)
di: Chang, Raeyoung, et al.
Pubblicazione: (2026)
SMaRT: Select, Mix, and ReinvenT -- A Strategy Fusion Framework for LLM-Driven Reasoning and Planning
di: Verma, Nikhil, et al.
Pubblicazione: (2025)
di: Verma, Nikhil, et al.
Pubblicazione: (2025)
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2026)
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2026)
Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
di: Choi, Dongwook, et al.
Pubblicazione: (2025)
di: Choi, Dongwook, et al.
Pubblicazione: (2025)
Counterfactual Graph for Multi-Agent LLM Calibration
di: Huang, Jiatan, et al.
Pubblicazione: (2026)
di: Huang, Jiatan, et al.
Pubblicazione: (2026)
The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context
di: Verma, Nikhil, et al.
Pubblicazione: (2025)
di: Verma, Nikhil, et al.
Pubblicazione: (2025)
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2026)
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2026)
Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation
di: Dua, Radhika, et al.
Pubblicazione: (2025)
di: Dua, Radhika, et al.
Pubblicazione: (2025)
Overthinking Loops in Agents: A Structural Risk via MCP Tools
di: Lee, Yohan, et al.
Pubblicazione: (2026)
di: Lee, Yohan, et al.
Pubblicazione: (2026)
Agent-as-a-Graph: Knowledge Graph-Based Tool and Agent Retrieval for LLM Multi-Agent Systems
di: Nizar, Faheem, et al.
Pubblicazione: (2025)
di: Nizar, Faheem, et al.
Pubblicazione: (2025)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
di: Kwon, Deuksin, et al.
Pubblicazione: (2025)
di: Kwon, Deuksin, et al.
Pubblicazione: (2025)
GPT-DETOX: An In-Context Learning-Based Paraphraser for Text Detoxification
di: Pesaranghader, Ali, et al.
Pubblicazione: (2024)
di: Pesaranghader, Ali, et al.
Pubblicazione: (2024)
Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization
di: Kwon, Taeyoon, et al.
Pubblicazione: (2025)
di: Kwon, Taeyoon, et al.
Pubblicazione: (2025)
Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes
di: Nath, Abhijnan, et al.
Pubblicazione: (2025)
di: Nath, Abhijnan, et al.
Pubblicazione: (2025)
EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents
di: Choi, Dongwook, et al.
Pubblicazione: (2026)
di: Choi, Dongwook, et al.
Pubblicazione: (2026)
Knowledge Tagging with Large Language Model based Multi-Agent System
di: Li, Hang, et al.
Pubblicazione: (2024)
di: Li, Hang, et al.
Pubblicazione: (2024)
Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis
di: Mok, Jisoo, et al.
Pubblicazione: (2025)
di: Mok, Jisoo, et al.
Pubblicazione: (2025)
Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
di: Lee, Dongwook, et al.
Pubblicazione: (2026)
di: Lee, Dongwook, et al.
Pubblicazione: (2026)
Reinforcement World Model Learning for LLM-based Agents
di: Yu, Xiao, et al.
Pubblicazione: (2026)
di: Yu, Xiao, et al.
Pubblicazione: (2026)
PersonalHomeBench: Evaluating Agents in Personalized Smart Homes
di: Bharadwaj, Manasa, et al.
Pubblicazione: (2026)
di: Bharadwaj, Manasa, et al.
Pubblicazione: (2026)
KULCQ: An Unsupervised Keyword-based Utterance Level Clustering Quality Metric
di: Guruprasad, Pranav, et al.
Pubblicazione: (2024)
di: Guruprasad, Pranav, et al.
Pubblicazione: (2024)
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
di: Choi, Junhyuk, et al.
Pubblicazione: (2026)
di: Choi, Junhyuk, et al.
Pubblicazione: (2026)
LLM-Based Multi-Agent Systems are Scalable Graph Generative Models
di: Ji, Jiarui, et al.
Pubblicazione: (2024)
di: Ji, Jiarui, et al.
Pubblicazione: (2024)
Adaptive Graph Pruning for Multi-Agent Communication
di: Li, Boyi, et al.
Pubblicazione: (2025)
di: Li, Boyi, et al.
Pubblicazione: (2025)
Evaluating Automatic Metrics with Incremental Machine Translation Systems
di: Wu, Guojun, et al.
Pubblicazione: (2024)
di: Wu, Guojun, et al.
Pubblicazione: (2024)
Quality-Aware Translation Tagging in Multilingual RAG system
di: Moon, Hoyeon, et al.
Pubblicazione: (2025)
di: Moon, Hoyeon, et al.
Pubblicazione: (2025)
Region4Web: Rethinking Observation Space Granularity for Web Agents
di: Kwon, Donguk, et al.
Pubblicazione: (2026)
di: Kwon, Donguk, et al.
Pubblicazione: (2026)
CRAFT: Grounded Multi-Agent Coordination Under Partial Information
di: Nath, Abhijnan, et al.
Pubblicazione: (2026)
di: Nath, Abhijnan, et al.
Pubblicazione: (2026)
LLM-based Frameworks for API Argument Filling in Task-Oriented Conversational Systems
di: Mok, Jisoo, et al.
Pubblicazione: (2024)
di: Mok, Jisoo, et al.
Pubblicazione: (2024)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
di: Kartik, NVJK, et al.
Pubblicazione: (2025)
di: Kartik, NVJK, et al.
Pubblicazione: (2025)
PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses
di: Hong, Minki, et al.
Pubblicazione: (2026)
di: Hong, Minki, et al.
Pubblicazione: (2026)
How to Choose How to Choose Your Chatbot: A Massively Multi-System MultiReference Data Set for Dialog Metric Evaluation
di: Khayrallah, Huda, et al.
Pubblicazione: (2023)
di: Khayrallah, Huda, et al.
Pubblicazione: (2023)
Creativity in LLM-based Multi-Agent Systems: A Survey
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
TeXBLEU: Automatic Metric for Evaluate LaTeX Format
di: Jung, Kyudan, et al.
Pubblicazione: (2024)
di: Jung, Kyudan, et al.
Pubblicazione: (2024)
Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research
di: Zhang, Qianqian, et al.
Pubblicazione: (2025)
di: Zhang, Qianqian, et al.
Pubblicazione: (2025)
Enhancing Knowledge Graph Construction: Evaluating with Emphasis on Hallucination, Omission, and Graph Similarity Metrics
di: Ghanem, Hussam, et al.
Pubblicazione: (2025)
di: Ghanem, Hussam, et al.
Pubblicazione: (2025)
MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables
di: Seo, Kwangwook, et al.
Pubblicazione: (2025)
di: Seo, Kwangwook, et al.
Pubblicazione: (2025)
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
di: Ma, Zi-Ao, et al.
Pubblicazione: (2024)
di: Ma, Zi-Ao, et al.
Pubblicazione: (2024)
BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems
di: Wang, Wei, et al.
Pubblicazione: (2024)
di: Wang, Wei, et al.
Pubblicazione: (2024)
Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
di: Hua, Yilun, et al.
Pubblicazione: (2026)
di: Hua, Yilun, et al.
Pubblicazione: (2026)
Documenti analoghi
-
CascadeDebate: Multi-Agent Deliberation for Cost-Aware LLM Cascades
di: Chang, Raeyoung, et al.
Pubblicazione: (2026) -
SMaRT: Select, Mix, and ReinvenT -- A Strategy Fusion Framework for LLM-Driven Reasoning and Planning
di: Verma, Nikhil, et al.
Pubblicazione: (2025) -
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
di: Martin-Boyle, Anna, et al.
Pubblicazione: (2026) -
Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
di: Choi, Dongwook, et al.
Pubblicazione: (2025) -
Counterfactual Graph for Multi-Agent LLM Calibration
di: Huang, Jiatan, et al.
Pubblicazione: (2026)