Diagnostics of cognitive failures in multi-agent expert systems using dynamic evaluation protocols and subsequent mutation of the processing context
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sorstkins, Andrejs, Bailey, Josh, Baron, Dr Alistair |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Assessing RAG and HyDE on 1B vs. 4B-Parameter Gemma LLMs for Personal Assistants Integretion
par: Sorstkins, Andrejs
Publié: (2025)
par: Sorstkins, Andrejs
Publié: (2025)
Learning to Undo: Rollback-Augmented Reinforcement Learning with Reversibility Signals
par: Sorstkins, Andrejs, et autres
Publié: (2025)
par: Sorstkins, Andrejs, et autres
Publié: (2025)
A process algebraic framework for multi-agent dynamic epistemic systems
par: Aldini, Alessandro
Publié: (2024)
par: Aldini, Alessandro
Publié: (2024)
Adaptive routing protocols for determining optimal paths in AI multi-agent systems: a priority- and learning-enhanced approach
par: Panayotov, Theodor, et autres
Publié: (2025)
par: Panayotov, Theodor, et autres
Publié: (2025)
The impact of multi-agent debate protocols on debate quality: a controlled case study
par: Marandi, Ramtin Zargari
Publié: (2026)
par: Marandi, Ramtin Zargari
Publié: (2026)
Fuzzy expert system for the process of collecting and purifying acidic water: a digital twin approach
par: Maratuly, Temirbolat, et autres
Publié: (2026)
par: Maratuly, Temirbolat, et autres
Publié: (2026)
Metric assessment protocol in the context of answer fluctuation on MCQ tasks
par: Goliakova, Ekaterina, et autres
Publié: (2025)
par: Goliakova, Ekaterina, et autres
Publié: (2025)
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
par: Wijk, Hjalmar, et autres
Publié: (2024)
par: Wijk, Hjalmar, et autres
Publié: (2024)
Context-Picker: Dynamic context selection using multi-stage reinforcement learning
par: Zhu, Siyuan, et autres
Publié: (2025)
par: Zhu, Siyuan, et autres
Publié: (2025)
System 0/1/2/3: Quad-process theory for multi-timescale embodied collective cognitive systems
par: Taniguchi, Tadahiro, et autres
Publié: (2025)
par: Taniguchi, Tadahiro, et autres
Publié: (2025)
Legal interpretation and AI: from expert systems to argumentation and LLMs
par: Janeček, Václav, et autres
Publié: (2026)
par: Janeček, Václav, et autres
Publié: (2026)
Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
par: Driouich, Ilias, et autres
Publié: (2025)
par: Driouich, Ilias, et autres
Publié: (2025)
Discovering mathematical concepts through a multi-agent system
par: Aggarwal, Daattavya, et autres
Publié: (2026)
par: Aggarwal, Daattavya, et autres
Publié: (2026)
Multi-agent cooperation through in-context co-player inference
par: Weis, Marissa A., et autres
Publié: (2026)
par: Weis, Marissa A., et autres
Publié: (2026)
Graders should cheat: privileged information enables expert-level automated evaluations
par: Zhou, Jin Peng, et autres
Publié: (2025)
par: Zhou, Jin Peng, et autres
Publié: (2025)
Memory poisoning and secure multi-agent systems
par: Torra, Vicenç, et autres
Publié: (2026)
par: Torra, Vicenç, et autres
Publié: (2026)
Enhanced Transformer architecture for in-context learning of dynamical systems
par: Rufolo, Matteo, et autres
Publié: (2024)
par: Rufolo, Matteo, et autres
Publié: (2024)
FREIDA: A Framework for developing quantitative agent based models based on qualitative expert knowledge
par: Oetker, Frederike, et autres
Publié: (2023)
par: Oetker, Frederike, et autres
Publié: (2023)
Setting up for failure: automatic discovery of the neural mechanisms of cognitive errors
par: Radmard, Puria, et autres
Publié: (2025)
par: Radmard, Puria, et autres
Publié: (2025)
Configurable multi-agent framework for scalable and realistic testing of llm-based agents
par: Wang, Sai, et autres
Publié: (2025)
par: Wang, Sai, et autres
Publié: (2025)
Log analysis is necessary for credible evaluation of AI agents
par: Kirgis, Peter, et autres
Publié: (2026)
par: Kirgis, Peter, et autres
Publié: (2026)
Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation
par: Wysocka, Magdalena, et autres
Publié: (2023)
par: Wysocka, Magdalena, et autres
Publié: (2023)
Robin: A multi-agent system for automating scientific discovery
par: Ghareeb, Ali Essam, et autres
Publié: (2025)
par: Ghareeb, Ali Essam, et autres
Publié: (2025)
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
par: Ruan, Shouwei, et autres
Publié: (2025)
par: Ruan, Shouwei, et autres
Publié: (2025)
Adaptive parameter sharing for multi-agent reinforcement learning
par: Li, Dapeng, et autres
Publié: (2023)
par: Li, Dapeng, et autres
Publié: (2023)
Realistic pedestrian-driver interaction modelling using multi-agent RL with human perceptual-motor constraints
par: Wang, Yueyang, et autres
Publié: (2025)
par: Wang, Yueyang, et autres
Publié: (2025)
LLMs learn governing principles of dynamical systems, revealing an in-context neural scaling law
par: Liu, Toni J. B., et autres
Publié: (2024)
par: Liu, Toni J. B., et autres
Publié: (2024)
An AI system to help scientists write expert-level empirical software
par: Aygün, Eser, et autres
Publié: (2025)
par: Aygün, Eser, et autres
Publié: (2025)
Stream-based perception for cognitive agents in mobile ecosystems
par: Dötterl, Jeremias, et autres
Publié: (2024)
par: Dötterl, Jeremias, et autres
Publié: (2024)
Agentic clinical reasoning over longitudinal myeloma records: a retrospective evaluation against expert consensus
par: Moll, Johannes, et autres
Publié: (2026)
par: Moll, Johannes, et autres
Publié: (2026)
Using multi-agent architecture to mitigate the risk of LLM hallucinations
par: Amer, Abd Elrahman, et autres
Publié: (2025)
par: Amer, Abd Elrahman, et autres
Publié: (2025)
Plancraft: an evaluation dataset for planning with LLM agents
par: Dagan, Gautier, et autres
Publié: (2024)
par: Dagan, Gautier, et autres
Publié: (2024)
A systematic review on expert systems for improving energy efficiency in the manufacturing industry
par: Ioshchikhes, Borys, et autres
Publié: (2024)
par: Ioshchikhes, Borys, et autres
Publié: (2024)
Acceleration method for generating perception failure scenarios based on editing Markov process
par: Cai, Canjie
Publié: (2024)
par: Cai, Canjie
Publié: (2024)
A Large Language Model-based multi-agent manufacturing system for intelligent shopfloor
par: Zhao, Zhen, et autres
Publié: (2024)
par: Zhao, Zhen, et autres
Publié: (2024)
WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search
par: Hu, Yuelin, et autres
Publié: (2026)
par: Hu, Yuelin, et autres
Publié: (2026)
MAFA: A multi-agent framework for annotation
par: Hegazy, Mahmood, et autres
Publié: (2025)
par: Hegazy, Mahmood, et autres
Publié: (2025)
Beyond the high score: Prosocial ability profiles of multi-agent populations
par: Tesic, Marko, et autres
Publié: (2025)
par: Tesic, Marko, et autres
Publié: (2025)
Dynamic fairness-aware recommendation through multi-agent social choice
par: Aird, Amanda, et autres
Publié: (2023)
par: Aird, Amanda, et autres
Publié: (2023)
Reshaping MOFs text mining with a dynamic multi-agents framework of large language model
par: Lin, Zuhong, et autres
Publié: (2025)
par: Lin, Zuhong, et autres
Publié: (2025)
Documents similaires
-
Assessing RAG and HyDE on 1B vs. 4B-Parameter Gemma LLMs for Personal Assistants Integretion
par: Sorstkins, Andrejs
Publié: (2025) -
Learning to Undo: Rollback-Augmented Reinforcement Learning with Reversibility Signals
par: Sorstkins, Andrejs, et autres
Publié: (2025) -
A process algebraic framework for multi-agent dynamic epistemic systems
par: Aldini, Alessandro
Publié: (2024) -
Adaptive routing protocols for determining optimal paths in AI multi-agent systems: a priority- and learning-enhanced approach
par: Panayotov, Theodor, et autres
Publié: (2025) -
The impact of multi-agent debate protocols on debate quality: a controlled case study
par: Marandi, Ramtin Zargari
Publié: (2026)