Self-Improvement of Language Models by Post-Training on Multi-Agent Debate
Fuente:
arXiv
Saved in:
| Main Authors: | Samanta, Ankur, Magesh, Akshayaa, Wu, Runzhe, Jain, Ayush, Yu, Youliang, Jiang, Daniel, Vidolov, Boris, Sajda, Paul, Efroni, Yonathan, Hassani, Kaveh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structure Enables Effective Self-Localization of Errors in LLMs
by: Samanta, Ankur, et al.
Published: (2026)
by: Samanta, Ankur, et al.
Published: (2026)
Credit Assignment with Resets in Language Model Reasoning
by: Samanta, Ankur, et al.
Published: (2026)
by: Samanta, Ankur, et al.
Published: (2026)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
by: Wu, Runzhe, et al.
Published: (2025)
by: Wu, Runzhe, et al.
Published: (2025)
Heterogeneous Multi-Player Multi-Armed Bandits Robust To Adversarial Attacks
by: Magesh, Akshayaa, et al.
Published: (2025)
by: Magesh, Akshayaa, et al.
Published: (2025)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)
by: Roth, Amit, et al.
Published: (2026)
Principled Detection of Hallucinations in Large Language Models via Multiple Testing
by: Li, Jiawei, et al.
Published: (2025)
by: Li, Jiawei, et al.
Published: (2025)
Simple Optimizers for Convex Aligned Multi-Objective Optimization
by: Kretzu, Ben, et al.
Published: (2025)
by: Kretzu, Ben, et al.
Published: (2025)
Robust Multi-Hypothesis Testing with Moment Constrained Uncertainty Sets
by: Magesh, Akshayaa, et al.
Published: (2022)
by: Magesh, Akshayaa, et al.
Published: (2022)
Gaze patterns predict preference and confidence in pairwise AI image evaluation
by: Papadopoulos, Nikolas, et al.
Published: (2026)
by: Papadopoulos, Nikolas, et al.
Published: (2026)
Gradient Free Deep Reinforcement Learning With TabPFN
by: Schiff, David, et al.
Published: (2025)
by: Schiff, David, et al.
Published: (2025)
The Bias of Harmful Label Associations in Vision-Language Models
by: Hazirbas, Caner, et al.
Published: (2024)
by: Hazirbas, Caner, et al.
Published: (2024)
Exploiting Structure in Offline Multi-Agent RL: The Benefits of Low Interaction Rank
by: Zhan, Wenhao, et al.
Published: (2024)
by: Zhan, Wenhao, et al.
Published: (2024)
Aligned Multi Objective Optimization
by: Efroni, Yonathan, et al.
Published: (2025)
by: Efroni, Yonathan, et al.
Published: (2025)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs
by: Wu, Lili, et al.
Published: (2024)
by: Wu, Lili, et al.
Published: (2024)
Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
by: Yi, John Seon Keun, et al.
Published: (2026)
by: Yi, John Seon Keun, et al.
Published: (2026)
Enabling Multi-Robot Collaboration from Single-Human Guidance
by: Ji, Zhengran, et al.
Published: (2024)
by: Ji, Zhengran, et al.
Published: (2024)
Predicting Public Transportation Crowd Using Weather API and Social Media
by: Akshayaa Shree.M, Maheshwari.S
Published: (2026)
by: Akshayaa Shree.M, Maheshwari.S
Published: (2026)
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
Time After Time: Deep-Q Effect Estimation for Interventions on When and What to do
by: Wald, Yoav, et al.
Published: (2025)
by: Wald, Yoav, et al.
Published: (2025)
Ossarth : An Open-Source Customizable LLMOS
by: Magesh, Siddharth
Published: (2026)
by: Magesh, Siddharth
Published: (2026)
Perception of an AI Teammate in an Embodied Control Task Affects Team Performance, Reflected in Human Teammates' Behaviors and Physiological Responses
by: Qin, Yinuo, et al.
Published: (2025)
by: Qin, Yinuo, et al.
Published: (2025)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
by: Salimi, Moein, et al.
Published: (2026)
by: Salimi, Moein, et al.
Published: (2026)
Capitalisms of the “Global South” (c. 10th to 19th Centuries) - Old and New Contributions and Debates
by: Kaveh Yazdani
Published: (2023)
by: Kaveh Yazdani
Published: (2023)
Computationally Efficient RL under Linear Bellman Completeness for Deterministic Dynamics
by: Wu, Runzhe, et al.
Published: (2024)
by: Wu, Runzhe, et al.
Published: (2024)
Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate
by: Liu, Chenxi, et al.
Published: (2026)
by: Liu, Chenxi, et al.
Published: (2026)
Exact analysis of potential flow past bodies of irregular shapes
by: Jain, Ankur
Published: (2025)
by: Jain, Ankur
Published: (2025)
Detecting Causality with Symplectic Quandles
by: Jain, Ayush
Published: (2023)
by: Jain, Ayush
Published: (2023)
Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
by: Singla, Pratham, et al.
Published: (2025)
by: Singla, Pratham, et al.
Published: (2025)
Train on Validation (ToV): Fast data selection with applications to fine-tuning
by: Jain, Ayush, et al.
Published: (2025)
by: Jain, Ayush, et al.
Published: (2025)
EEG-estimated functional connectivity, and not behavior, differentiates Parkinson's patients from health controls during the Simon conflict task
by: Sun, Xiaoxiao, et al.
Published: (2024)
by: Sun, Xiaoxiao, et al.
Published: (2024)
Improvement in the Performance of Composite Cathode via Solid Electrolyte for High Electrochemical Performance of Li–S Battery
by: Tarun Patodia, et al.
Published: (2024)
by: Tarun Patodia, et al.
Published: (2024)
CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate
by: Sun, Yiliu, et al.
Published: (2025)
by: Sun, Yiliu, et al.
Published: (2025)
Contextual Experience Replay for Self-Improvement of Language Agents
by: Liu, Yitao, et al.
Published: (2025)
by: Liu, Yitao, et al.
Published: (2025)
GroupDebate: Enhancing the Efficiency of Multi-Agent Debate Using Group Discussion
by: Liu, Tongxuan, et al.
Published: (2024)
by: Liu, Tongxuan, et al.
Published: (2024)
SIDiffAgent: Self-Improving Diffusion Agent
by: Garg, Shivank, et al.
Published: (2026)
by: Garg, Shivank, et al.
Published: (2026)
Good things come in three: Generating SO Post Titles with Pre-Trained Models, Self Improvement and Post Ranking
by: Le, Duc Anh, et al.
Published: (2024)
by: Le, Duc Anh, et al.
Published: (2024)
Self-Trained Verification for Training- and Test-Time Self-Improvement
by: Wu, Chen Henry, et al.
Published: (2026)
by: Wu, Chen Henry, et al.
Published: (2026)
AI based Content Creation and Product Recommendation Applications in E-commerce: An Ethical overview
by: Jain, Aditi Madhusudan, et al.
Published: (2025)
by: Jain, Aditi Madhusudan, et al.
Published: (2025)
Rethinking Legal Compliance Automation: Opportunities with Large Language Models
by: Hassani, Shabnam, et al.
Published: (2024)
by: Hassani, Shabnam, et al.
Published: (2024)
Similar Items
-
Structure Enables Effective Self-Localization of Errors in LLMs
by: Samanta, Ankur, et al.
Published: (2026) -
Credit Assignment with Resets in Language Model Reasoning
by: Samanta, Ankur, et al.
Published: (2026) -
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
by: Wu, Runzhe, et al.
Published: (2025) -
Heterogeneous Multi-Player Multi-Armed Bandits Robust To Adversarial Attacks
by: Magesh, Akshayaa, et al.
Published: (2025) -
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)