Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
Fuente:
arXiv
Salvato in:
| Autori principali: | Mushtaq, Abdullah, Naeem, Muhammad Rafay, Ghaznavi, Ibrahim, Abd-alrazaq, Alaa, Tabassum, Aliya, Qadir, Junaid |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
Harnessing Multi-Agent LLMs for Complex Engineering Problem-Solving: A Framework for Senior Design Projects
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
Toward Inclusive Educational AI: Auditing Frontier LLMs through a Multiplexity Lens
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
di: Elmahjub, Ezieddin, et al.
Pubblicazione: (2026)
di: Elmahjub, Ezieddin, et al.
Pubblicazione: (2026)
Inducing Personality in LLM-Based Honeypot Agents: Measuring the Effect on Human-Like Agenda Generation
di: Newsham, Lewis, et al.
Pubblicazione: (2025)
di: Newsham, Lewis, et al.
Pubblicazione: (2025)
Can LLM Agents Really Debate? A Controlled Study of Multi-Agent Debate in Logical Reasoning
di: Wu, Haolun, et al.
Pubblicazione: (2025)
di: Wu, Haolun, et al.
Pubblicazione: (2025)
Evaluating Collective Behaviour of Hundreds of LLM Agents
di: Willis, Richard, et al.
Pubblicazione: (2026)
di: Willis, Richard, et al.
Pubblicazione: (2026)
When AI Agents Disagree Like Humans: Reasoning Trace Analysis for Human-AI Collaborative Moderation
di: Wawer, Michał, et al.
Pubblicazione: (2026)
di: Wawer, Michał, et al.
Pubblicazione: (2026)
Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning
di: Atif, Muhammad Ahmed, et al.
Pubblicazione: (2026)
di: Atif, Muhammad Ahmed, et al.
Pubblicazione: (2026)
Code Like Humans: A Multi-Agent Solution for Medical Coding
di: Motzfeldt, Andreas, et al.
Pubblicazione: (2025)
di: Motzfeldt, Andreas, et al.
Pubblicazione: (2025)
Evaluating Multi-Agent LLM Architectures for Rare Disease Diagnosis
di: Almasoud, Ahmed
Pubblicazione: (2026)
di: Almasoud, Ahmed
Pubblicazione: (2026)
Can AI Chatbots Provide Coaching in Engineering? Beyond Information Processing Toward Mastery
di: Qadir, Junaid, et al.
Pubblicazione: (2026)
di: Qadir, Junaid, et al.
Pubblicazione: (2026)
All Models Are Wrong, But Can They Be Useful? Lessons from COVID-19 Agent-Based Models: A Systematic Review
di: Von Hoene, Emma, et al.
Pubblicazione: (2025)
di: Von Hoene, Emma, et al.
Pubblicazione: (2025)
Can We Govern the Agent-to-Agent Economy?
di: Chaffer, Tomer Jordi
Pubblicazione: (2025)
di: Chaffer, Tomer Jordi
Pubblicazione: (2025)
SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents
di: Hamblin, Grant, et al.
Pubblicazione: (2026)
di: Hamblin, Grant, et al.
Pubblicazione: (2026)
When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM Judges
di: Liang, Sichu, et al.
Pubblicazione: (2026)
di: Liang, Sichu, et al.
Pubblicazione: (2026)
Talk, Judge, Cooperate: Gossip-Driven Indirect Reciprocity in Self-Interested LLM Agents
di: Zhu, Shuhui, et al.
Pubblicazione: (2026)
di: Zhu, Shuhui, et al.
Pubblicazione: (2026)
Can AI Agents Agree?
di: Berdoz, Frédéric, et al.
Pubblicazione: (2026)
di: Berdoz, Frédéric, et al.
Pubblicazione: (2026)
Bias-Adjusted LLM Agents for Human-Like Decision-Making via Behavioral Economics
di: Kitadai, Ayato, et al.
Pubblicazione: (2025)
di: Kitadai, Ayato, et al.
Pubblicazione: (2025)
GoAgent: Group-of-Agents Communication Topology Generation for LLM-based Multi-Agent Systems
di: Chen, Hongjiang, et al.
Pubblicazione: (2026)
di: Chen, Hongjiang, et al.
Pubblicazione: (2026)
ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents
di: Bousetouane, Fouad
Pubblicazione: (2026)
di: Bousetouane, Fouad
Pubblicazione: (2026)
BioAgents: Democratizing Bioinformatics Analysis with Multi-Agent Systems
di: Mehandru, Nikita, et al.
Pubblicazione: (2025)
di: Mehandru, Nikita, et al.
Pubblicazione: (2025)
MobiAgent: A Systematic Framework for Customizable Mobile Agents
di: Zhang, Cheng, et al.
Pubblicazione: (2025)
di: Zhang, Cheng, et al.
Pubblicazione: (2025)
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
di: P, Vedanta S, et al.
Pubblicazione: (2026)
di: P, Vedanta S, et al.
Pubblicazione: (2026)
Interactional Fairness in LLM Multi-Agent Systems: An Evaluation Framework
di: Binkyte, Ruta
Pubblicazione: (2025)
di: Binkyte, Ruta
Pubblicazione: (2025)
LLM-Enabled Multi-Agent Systems: Empirical Evaluation and Insights into Emerging Design Patterns & Paradigms
di: Renney, Harri, et al.
Pubblicazione: (2026)
di: Renney, Harri, et al.
Pubblicazione: (2026)
AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent
di: Luo, Yinyi, et al.
Pubblicazione: (2026)
di: Luo, Yinyi, et al.
Pubblicazione: (2026)
LLM-ABM for Transportation: Assessing the Potential of LLM Agents in System Analysis
di: Liu, Tianming, et al.
Pubblicazione: (2025)
di: Liu, Tianming, et al.
Pubblicazione: (2025)
Towards Cognitive Synergy in LLM-Based Multi-Agent Systems: Integrating Theory of Mind and Critical Evaluation
di: Kostka, Adam, et al.
Pubblicazione: (2025)
di: Kostka, Adam, et al.
Pubblicazione: (2025)
MLC-Agent: Cognitive Model based on Memory-Learning Collaboration in LLM Empowered Agent Simulation Environment
di: Zhang, Ming, et al.
Pubblicazione: (2025)
di: Zhang, Ming, et al.
Pubblicazione: (2025)
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
di: Sinha, Aarush, et al.
Pubblicazione: (2026)
di: Sinha, Aarush, et al.
Pubblicazione: (2026)
The Social Laboratory: A Psychometric Framework for Multi-Agent LLM Evaluation
di: Reza, Zarreen
Pubblicazione: (2025)
di: Reza, Zarreen
Pubblicazione: (2025)
ALAS: Transactional and Dynamic Multi-Agent LLM Planning
di: Geng, Longling, et al.
Pubblicazione: (2025)
di: Geng, Longling, et al.
Pubblicazione: (2025)
A Plan Reuse Mechanism for LLM-Driven Agent
di: Li, Guopeng, et al.
Pubblicazione: (2025)
di: Li, Guopeng, et al.
Pubblicazione: (2025)
LLM Harmony: Multi-Agent Communication for Problem Solving
di: Rasal, Sumedh
Pubblicazione: (2024)
di: Rasal, Sumedh
Pubblicazione: (2024)
ScamAgents: How AI Agents Can Simulate Human-Level Scam Calls
di: Badhe, Sanket
Pubblicazione: (2025)
di: Badhe, Sanket
Pubblicazione: (2025)
AgentSchool: An LLM-Powered Multi-Agent Simulation for Education
di: Ye, Yulei, et al.
Pubblicazione: (2026)
di: Ye, Yulei, et al.
Pubblicazione: (2026)
MegaAgent: A Large-Scale Autonomous LLM-based Multi-Agent System Without Predefined SOPs
di: Wang, Qian, et al.
Pubblicazione: (2024)
di: Wang, Qian, et al.
Pubblicazione: (2024)
Runtime Composition in Dynamic System of Systems: A Systematic Review of Challenges, Solutions, Tools, and Evaluation Methods
di: Ashfaq, Muhammad, et al.
Pubblicazione: (2025)
di: Ashfaq, Muhammad, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025) -
Harnessing Multi-Agent LLMs for Complex Engineering Problem-Solving: A Framework for Senior Design Projects
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025) -
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025) -
Toward Inclusive Educational AI: Auditing Frontier LLMs through a Multiplexity Lens
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025) -
IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
di: Elmahjub, Ezieddin, et al.
Pubblicazione: (2026)