ESSA: Evolutionary Strategies for Scalable Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Korotyshova, Daria, Shaposhnikov, Boris, Malakhov, Alexey, Khokhulin, Alexey, Surnachev, Nikita, Ovcharenko, Kirill, Bredis, George, Gorbatovski, Alexey, Sinii, Viacheslav, Gavrilov, Daniil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917161734242304
author Korotyshova, Daria
Shaposhnikov, Boris
Malakhov, Alexey
Khokhulin, Alexey
Surnachev, Nikita
Ovcharenko, Kirill
Bredis, George
Gorbatovski, Alexey
Sinii, Viacheslav
Gavrilov, Daniil
author_facet Korotyshova, Daria
Shaposhnikov, Boris
Malakhov, Alexey
Khokhulin, Alexey
Surnachev, Nikita
Ovcharenko, Kirill
Bredis, George
Gorbatovski, Alexey
Sinii, Viacheslav
Gavrilov, Daniil
contents Alignment of Large Language Models (LLMs) typically relies on Reinforcement Learning from Human Feedback (RLHF) with gradient-based optimizers such as Proximal Policy Optimization (PPO) or Group Relative Policy Optimization (GRPO). While effective, these methods require complex distributed training, large memory budgets, and careful hyperparameter tuning, all of which become increasingly difficult at billion-parameter scale. We present ESSA, Evolutionary Strategies for Scalable Alignment, a gradient-free framework that aligns LLMs using only forward inference and black-box optimization. ESSA focuses optimization on Low-Rank Adapters (LoRA) and further compresses their parameter space by optimizing only the singular values from an singular value decomposition (SVD) of each adapter matrix. This dimensionality reduction makes evolutionary search practical even for very large models and allows efficient operation in quantized INT4 and INT8 inference mode. Across these benchmarks ESSA improves the test accuracy of Qwen2.5-Math-7B by 12.6% on GSM8K and 14.8% on PRM800K, and raises the accuracy of LLaMA3.1-8B on IFEval by 22.5%, all compared with GRPO. In large-scale settings ESSA shows stronger scaling than gradient-based methods: on Qwen2.5-32B for PRM800K it reaches near-optimal accuracy twice as fast on 16 GPUs and six times as fast on 128 GPUs compared with GRPO. These results position evolutionary strategies as a compelling, hardware-friendly alternative to gradient-based LLM alignment, combining competitive quality with substantially reduced wall-clock time and engineering overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04453
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ESSA: Evolutionary Strategies for Scalable Alignment
Korotyshova, Daria
Shaposhnikov, Boris
Malakhov, Alexey
Khokhulin, Alexey
Surnachev, Nikita
Ovcharenko, Kirill
Bredis, George
Gorbatovski, Alexey
Sinii, Viacheslav
Gavrilov, Daniil
Machine Learning
Alignment of Large Language Models (LLMs) typically relies on Reinforcement Learning from Human Feedback (RLHF) with gradient-based optimizers such as Proximal Policy Optimization (PPO) or Group Relative Policy Optimization (GRPO). While effective, these methods require complex distributed training, large memory budgets, and careful hyperparameter tuning, all of which become increasingly difficult at billion-parameter scale. We present ESSA, Evolutionary Strategies for Scalable Alignment, a gradient-free framework that aligns LLMs using only forward inference and black-box optimization. ESSA focuses optimization on Low-Rank Adapters (LoRA) and further compresses their parameter space by optimizing only the singular values from an singular value decomposition (SVD) of each adapter matrix. This dimensionality reduction makes evolutionary search practical even for very large models and allows efficient operation in quantized INT4 and INT8 inference mode. Across these benchmarks ESSA improves the test accuracy of Qwen2.5-Math-7B by 12.6% on GSM8K and 14.8% on PRM800K, and raises the accuracy of LLaMA3.1-8B on IFEval by 22.5%, all compared with GRPO. In large-scale settings ESSA shows stronger scaling than gradient-based methods: on Qwen2.5-32B for PRM800K it reaches near-optimal accuracy twice as fast on 16 GPUs and six times as fast on 128 GPUs compared with GRPO. These results position evolutionary strategies as a compelling, hardware-friendly alternative to gradient-based LLM alignment, combining competitive quality with substantially reduced wall-clock time and engineering overhead.
title ESSA: Evolutionary Strategies for Scalable Alignment
topic Machine Learning
url https://arxiv.org/abs/2507.04453