ESSA: Evolutionary Strategies for Scalable Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Korotyshova, Daria, Shaposhnikov, Boris, Malakhov, Alexey, Khokhulin, Alexey, Surnachev, Nikita, Ovcharenko, Kirill, Bredis, George, Gorbatovski, Alexey, Sinii, Viacheslav, Gavrilov, Daniil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Differences Between Direct Alignment Algorithms are a Blur
by: Gorbatovski, Alexey, et al.
Published: (2025)
by: Gorbatovski, Alexey, et al.
Published: (2025)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Trust-Region Behavior Blending for On-Policy Distillation
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Learn Your Reference Model for Real Good Alignment
by: Gorbatovski, Alexey, et al.
Published: (2024)
by: Gorbatovski, Alexey, et al.
Published: (2024)
Steering LLM Reasoning Through Bias-Only Adaptation
by: Sinii, Viacheslav, et al.
Published: (2025)
by: Sinii, Viacheslav, et al.
Published: (2025)
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
by: Sinii, Viacheslav, et al.
Published: (2025)
by: Sinii, Viacheslav, et al.
Published: (2025)
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
by: Bredis, George, et al.
Published: (2025)
by: Bredis, George, et al.
Published: (2025)
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
by: Aksenov, Yaroslav, et al.
Published: (2024)
by: Aksenov, Yaroslav, et al.
Published: (2024)
Next Embedding Prediction Makes World Models Stronger
by: Bredis, George, et al.
Published: (2026)
by: Bredis, George, et al.
Published: (2026)
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025)
by: Gerasimov, Gleb, et al.
Published: (2025)
Generative Flow Networks as Entropy-Regularized RL
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Mechanistic Permutability: Match Features Across Layers
by: Balagansky, Nikita, et al.
Published: (2024)
by: Balagansky, Nikita, et al.
Published: (2024)
Uncertainty Estimation of Transformers' Predictions via Topological Analysis of the Attention Matrices
by: Kostenok, Elizaveta, et al.
Published: (2023)
by: Kostenok, Elizaveta, et al.
Published: (2023)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
by: Laptev, Daniil, et al.
Published: (2025)
by: Laptev, Daniil, et al.
Published: (2025)
Evaluating Robustness in Latent Diffusion Models via Embedding Level Augmentation
by: Martirosyan, Boris, et al.
Published: (2025)
by: Martirosyan, Boris, et al.
Published: (2025)
Improving GFlowNets with Monte Carlo Tree Search
by: Morozov, Nikita, et al.
Published: (2024)
by: Morozov, Nikita, et al.
Published: (2024)
Adaptive Destruction Processes for Diffusion Samplers
by: Gritsaev, Timofei, et al.
Published: (2025)
by: Gritsaev, Timofei, et al.
Published: (2025)
XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX
by: Nikulin, Alexander, et al.
Published: (2023)
by: Nikulin, Alexander, et al.
Published: (2023)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025)
by: Kurochkin, Vadim, et al.
Published: (2025)
On Efficient Scaling of GNNs via IO-Aware Layers Implementations
by: Fomina, Daria, et al.
Published: (2026)
by: Fomina, Daria, et al.
Published: (2026)
Reinforcement learning for question answering in programming domain using public community scoring as a human feedback
by: Gorbatovski, Alexey, et al.
Published: (2024)
by: Gorbatovski, Alexey, et al.
Published: (2024)
Accelerating Transformers in Online RL
by: Zelezetsky, Daniil, et al.
Published: (2025)
by: Zelezetsky, Daniil, et al.
Published: (2025)
Teach Old SAEs New Domain Tricks with Boosting
by: Koriagin, Nikita, et al.
Published: (2025)
by: Koriagin, Nikita, et al.
Published: (2025)
Guided Star-Shaped Masked Diffusion
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
by: Balagansky, Nikita, et al.
Published: (2025)
by: Balagansky, Nikita, et al.
Published: (2025)
On (not) learning the Möbius function
by: Pozdnyakov, Alexey
Published: (2026)
by: Pozdnyakov, Alexey
Published: (2026)
Ensemble-based graph representation of fMRI data for cognitive brain state classification
by: Vlasenko, Daniil, et al.
Published: (2025)
by: Vlasenko, Daniil, et al.
Published: (2025)
Diffusion Language Models Generation Can Be Halted Early
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
Method for noise-induced regularization in quantum neural networks
by: Kuzmin, Viacheslav, et al.
Published: (2024)
by: Kuzmin, Viacheslav, et al.
Published: (2024)
Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
by: Alexey, Protopopov
Published: (2026)
by: Alexey, Protopopov
Published: (2026)
Emergence of In-Context Reinforcement Learning from Noise Distillation
by: Zisman, Ilya, et al.
Published: (2023)
by: Zisman, Ilya, et al.
Published: (2023)
The Euler characteristic of a triangulated manifold in terms of even-dimensional faces
by: Gavrilov, Alexey V.
Published: (2025)
by: Gavrilov, Alexey V.
Published: (2025)
Tight Bounds for Schrödinger Potential Estimation in Unpaired Data Translation
by: Puchkin, Nikita, et al.
Published: (2025)
by: Puchkin, Nikita, et al.
Published: (2025)
N-Gram Induction Heads for In-Context RL: Improving Stability and Reducing Data Needs
by: Zisman, Ilya, et al.
Published: (2024)
by: Zisman, Ilya, et al.
Published: (2024)
Uniting contrastive and generative learning for event sequences models
by: Yugay, Aleksandr, et al.
Published: (2024)
by: Yugay, Aleksandr, et al.
Published: (2024)
GoalLadder: Incremental Goal Discovery with Vision-Language Models
by: Zakharov, Alexey, et al.
Published: (2025)
by: Zakharov, Alexey, et al.
Published: (2025)
Probability-Generating Function Kernels for Spherical Data
by: Papamarkou, Theodore, et al.
Published: (2021)
by: Papamarkou, Theodore, et al.
Published: (2021)
Schrödinger bridge problem via empirical risk minimization
by: Belomestny, Denis, et al.
Published: (2026)
by: Belomestny, Denis, et al.
Published: (2026)
Similar Items
-
The Differences Between Direct Alignment Algorithms are a Blur
by: Gorbatovski, Alexey, et al.
Published: (2025) -
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026) -
Trust-Region Behavior Blending for On-Policy Distillation
by: Plyusov, Daniil, et al.
Published: (2026) -
Learn Your Reference Model for Real Good Alignment
by: Gorbatovski, Alexey, et al.
Published: (2024) -
Steering LLM Reasoning Through Bias-Only Adaptation
by: Sinii, Viacheslav, et al.
Published: (2025)