Saved in:
Bibliographic Details
Main Authors: Ozer, Onat, Wu, Grace, Wang, Yuchen, Dosti, Daniel, Zhang, Honghao, De La Rue, Vivi
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.20845
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914218010214400
author Ozer, Onat
Wu, Grace
Wang, Yuchen
Dosti, Daniel
Zhang, Honghao
De La Rue, Vivi
author_facet Ozer, Onat
Wu, Grace
Wang, Yuchen
Dosti, Daniel
Zhang, Honghao
De La Rue, Vivi
contents LLMs have shown the capacity to improve their performance on reasoning tasks through reflecting on their mistakes, and acting with these reflections in mind. However, continual reflections of the same LLM onto itself exhibit degeneration of thought, where the LLM continues to repeat the same errors again and again even with the knowledge that its wrong. To address this problem, we instead introduce multi-agent with multi-persona debators as the method to generate reflections. Through out extensive experimentation, we've found that the leads to better diversity of in the reflections generated by the llm agent. We demonstrate an accuracy of 47% EM HotPot QA (question answering) and 82.7% on HumanEval (programming), both performances surpassing reflection with a single llm.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20845
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAR:Multi-Agent Reflexion Improves Reasoning Abilities in LLMs
Ozer, Onat
Wu, Grace
Wang, Yuchen
Dosti, Daniel
Zhang, Honghao
De La Rue, Vivi
Artificial Intelligence
Multiagent Systems
LLMs have shown the capacity to improve their performance on reasoning tasks through reflecting on their mistakes, and acting with these reflections in mind. However, continual reflections of the same LLM onto itself exhibit degeneration of thought, where the LLM continues to repeat the same errors again and again even with the knowledge that its wrong. To address this problem, we instead introduce multi-agent with multi-persona debators as the method to generate reflections. Through out extensive experimentation, we've found that the leads to better diversity of in the reflections generated by the llm agent. We demonstrate an accuracy of 47% EM HotPot QA (question answering) and 82.7% on HumanEval (programming), both performances surpassing reflection with a single llm.
title MAR:Multi-Agent Reflexion Improves Reasoning Abilities in LLMs
topic Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2512.20845