MASEval: Extending Multi-Agent Evaluation from Models to Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Emde, Cornelius, Rubinstein, Alexander, Goel, Anmol, Heakl, Ahmed, Yun, Sangdoo, Oh, Seong Joon, Gubri, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models
von: Goel, Anmol, et al.
Veröffentlicht: (2026)
von: Goel, Anmol, et al.
Veröffentlicht: (2026)
Dr.LLM: Dynamic Layer Routing in LLMs
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
von: Puerto, Haritz, et al.
Veröffentlicht: (2024)
von: Puerto, Haritz, et al.
Veröffentlicht: (2024)
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
von: Green, Tommaso, et al.
Veröffentlicht: (2025)
von: Green, Tommaso, et al.
Veröffentlicht: (2025)
Calibrating Large Language Models Using Their Generations Only
von: Ulmer, Dennis, et al.
Veröffentlicht: (2024)
von: Ulmer, Dennis, et al.
Veröffentlicht: (2024)
C-SEO Bench: Does Conversational SEO Work?
von: Puerto, Haritz, et al.
Veröffentlicht: (2025)
von: Puerto, Haritz, et al.
Veröffentlicht: (2025)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
von: Gubri, Martin, et al.
Veröffentlicht: (2024)
von: Gubri, Martin, et al.
Veröffentlicht: (2024)
DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
von: Rubinstein, Alexander, et al.
Veröffentlicht: (2025)
von: Rubinstein, Alexander, et al.
Veröffentlicht: (2025)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
von: Do, Heejin, et al.
Veröffentlicht: (2025)
von: Do, Heejin, et al.
Veröffentlicht: (2025)
MEME: Multi-entity & Evolving Memory Evaluation
von: Jung, Seokwon, et al.
Veröffentlicht: (2026)
von: Jung, Seokwon, et al.
Veröffentlicht: (2026)
Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation
von: Mohamed, Asim, et al.
Veröffentlicht: (2025)
von: Mohamed, Asim, et al.
Veröffentlicht: (2025)
Do Deep Neural Network Solutions Form a Star Domain?
von: Sonthalia, Ankit, et al.
Veröffentlicht: (2024)
von: Sonthalia, Ankit, et al.
Veröffentlicht: (2024)
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
Scalable Ensemble Diversification for OOD Generalization and Detection
von: Rubinstein, Alexander, et al.
Veröffentlicht: (2024)
von: Rubinstein, Alexander, et al.
Veröffentlicht: (2024)
Are We Done with Object-Centric Learning?
von: Rubinstein, Alexander, et al.
Veröffentlicht: (2025)
von: Rubinstein, Alexander, et al.
Veröffentlicht: (2025)
AraSpider: Democratizing Arabic-to-SQL
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics
von: Soh, Yun Joon, et al.
Veröffentlicht: (2024)
von: Soh, Yun Joon, et al.
Veröffentlicht: (2024)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
von: Aneja, Krishak, et al.
Veröffentlicht: (2026)
von: Aneja, Krishak, et al.
Veröffentlicht: (2026)
Planner and Executor: Collaboration between Discrete Diffusion And Autoregressive Models in Reasoning
von: Berrayana, Lina, et al.
Veröffentlicht: (2025)
von: Berrayana, Lina, et al.
Veröffentlicht: (2025)
LLM generation novelty through the lens of semantic similarity
von: Davydov, Philipp, et al.
Veröffentlicht: (2025)
von: Davydov, Philipp, et al.
Veröffentlicht: (2025)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
von: Kodali, Prashant, et al.
Veröffentlicht: (2024)
von: Kodali, Prashant, et al.
Veröffentlicht: (2024)
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
von: Scimeca, Luca, et al.
Veröffentlicht: (2023)
von: Scimeca, Luca, et al.
Veröffentlicht: (2023)
Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
von: Shim, Jung-Woo, et al.
Veröffentlicht: (2025)
von: Shim, Jung-Woo, et al.
Veröffentlicht: (2025)
Domain-Partitioned Hybrid RAG for Legal Reasoning: Toward Modular and Explainable Legal AI for India
von: Goel, Rakshita, et al.
Veröffentlicht: (2025)
von: Goel, Rakshita, et al.
Veröffentlicht: (2025)
ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
SysTemp: A Multi-Agent System for Template-Based Generation of SysML v2
von: Bouamra, Yasmine, et al.
Veröffentlicht: (2025)
von: Bouamra, Yasmine, et al.
Veröffentlicht: (2025)
MATA: Multi-Agent Framework for Reliable and Flexible Table Question Answering
von: Hyeon, Sieun, et al.
Veröffentlicht: (2026)
von: Hyeon, Sieun, et al.
Veröffentlicht: (2026)
MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
von: Levi, Elad, et al.
Veröffentlicht: (2025)
von: Levi, Elad, et al.
Veröffentlicht: (2025)
Synthetic Feature Augmentation Improves Generalization Performance of Language Models
von: Choudhary, Ashok, et al.
Veröffentlicht: (2025)
von: Choudhary, Ashok, et al.
Veröffentlicht: (2025)
Multi-Agent Simulator Drives Language Models for Legal Intensive Interaction
von: Yue, Shengbin, et al.
Veröffentlicht: (2025)
von: Yue, Shengbin, et al.
Veröffentlicht: (2025)
Knowledge Tagging with Large Language Model based Multi-Agent System
von: Li, Hang, et al.
Veröffentlicht: (2024)
von: Li, Hang, et al.
Veröffentlicht: (2024)
ASIC-Agent: An Autonomous Multi-Agent System for ASIC Design with Benchmark Evaluation
von: Allam, Ahmed, et al.
Veröffentlicht: (2025)
von: Allam, Ahmed, et al.
Veröffentlicht: (2025)
Shh, don't say that! Domain Certification in LLMs
von: Emde, Cornelius, et al.
Veröffentlicht: (2025)
von: Emde, Cornelius, et al.
Veröffentlicht: (2025)
Auditing Language Model Unlearning via Information Decomposition
von: Goel, Anmol, et al.
Veröffentlicht: (2026)
von: Goel, Anmol, et al.
Veröffentlicht: (2026)
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
von: Kazi, Taaha, et al.
Veröffentlicht: (2024)
von: Kazi, Taaha, et al.
Veröffentlicht: (2024)
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
von: Huang, James Y., et al.
Veröffentlicht: (2025)
von: Huang, James Y., et al.
Veröffentlicht: (2025)
MIRIX: Multi-Agent Memory System for LLM-Based Agents
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
First Hallucination Tokens Are Different from Conditional Ones
von: Snel, Jakob, et al.
Veröffentlicht: (2025)
von: Snel, Jakob, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models
von: Goel, Anmol, et al.
Veröffentlicht: (2026) -
Dr.LLM: Dynamic Layer Routing in LLMs
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025) -
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
von: Puerto, Haritz, et al.
Veröffentlicht: (2024) -
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
von: Green, Tommaso, et al.
Veröffentlicht: (2025) -
Calibrating Large Language Models Using Their Generations Only
von: Ulmer, Dennis, et al.
Veröffentlicht: (2024)