Ensemble Debates with Local Large Language Models for AI Alignment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Sarabamoun, Ephraiem
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914159607676928
author Sarabamoun, Ephraiem
author_facet Sarabamoun, Ephraiem
contents As large language models (LLMs) take on greater roles in high-stakes decisions, alignment with human values is essential. Reliance on proprietary APIs limits reproducibility and broad participation. We study whether local open-source ensemble debates can improve alignmentoriented reasoning. Across 150 debates spanning 15 scenarios and five ensemble configurations, ensembles outperform single-model baselines on a 7-point rubric (overall: 3.48 vs. 3.13), with the largest gains in reasoning depth (+19.4%) and argument quality (+34.1%). Improvements are strongest for truthfulness (+1.25 points) and human enhancement (+0.80). We provide code, prompts, and a debate data set, providing an accessible and reproducible foundation for ensemble-based alignment evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00091
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ensemble Debates with Local Large Language Models for AI Alignment
Sarabamoun, Ephraiem
Artificial Intelligence
Computation and Language
As large language models (LLMs) take on greater roles in high-stakes decisions, alignment with human values is essential. Reliance on proprietary APIs limits reproducibility and broad participation. We study whether local open-source ensemble debates can improve alignmentoriented reasoning. Across 150 debates spanning 15 scenarios and five ensemble configurations, ensembles outperform single-model baselines on a 7-point rubric (overall: 3.48 vs. 3.13), with the largest gains in reasoning depth (+19.4%) and argument quality (+34.1%). Improvements are strongest for truthfulness (+1.25 points) and human enhancement (+0.80). We provide code, prompts, and a debate data set, providing an accessible and reproducible foundation for ensemble-based alignment evaluation.
title Ensemble Debates with Local Large Language Models for AI Alignment
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.00091