Trust-Region Behavior Blending for On-Policy Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Plyusov, Daniil, Gorbatovski, Alexey, Malakhov, Alexey, Balagansky, Nikita, Shaposhnikov, Boris, Korotyshova, Daria, Gavrilov, Daniil |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Steering LLM Reasoning Through Bias-Only Adaptation
by: Sinii, Viacheslav, et al.
Published: (2025)
by: Sinii, Viacheslav, et al.
Published: (2025)
The Differences Between Direct Alignment Algorithms are a Blur
by: Gorbatovski, Alexey, et al.
Published: (2025)
by: Gorbatovski, Alexey, et al.
Published: (2025)
Learn Your Reference Model for Real Good Alignment
by: Gorbatovski, Alexey, et al.
Published: (2024)
by: Gorbatovski, Alexey, et al.
Published: (2024)
ESSA: Evolutionary Strategies for Scalable Alignment
by: Korotyshova, Daria, et al.
Published: (2025)
by: Korotyshova, Daria, et al.
Published: (2025)
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
by: Aksenov, Yaroslav, et al.
Published: (2024)
by: Aksenov, Yaroslav, et al.
Published: (2024)
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
by: Sinii, Viacheslav, et al.
Published: (2025)
by: Sinii, Viacheslav, et al.
Published: (2025)
Next Embedding Prediction Makes World Models Stronger
by: Bredis, George, et al.
Published: (2026)
by: Bredis, George, et al.
Published: (2026)
Teach Old SAEs New Domain Tricks with Boosting
by: Koriagin, Nikita, et al.
Published: (2025)
by: Koriagin, Nikita, et al.
Published: (2025)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
by: Balagansky, Nikita, et al.
Published: (2025)
by: Balagansky, Nikita, et al.
Published: (2025)
Mechanistic Permutability: Match Features Across Layers
by: Balagansky, Nikita, et al.
Published: (2024)
by: Balagansky, Nikita, et al.
Published: (2024)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
by: Laptev, Daniil, et al.
Published: (2025)
by: Laptev, Daniil, et al.
Published: (2025)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025)
by: Kurochkin, Vadim, et al.
Published: (2025)
Improving GFlowNets with Monte Carlo Tree Search
by: Morozov, Nikita, et al.
Published: (2024)
by: Morozov, Nikita, et al.
Published: (2024)
Diffusion Language Models Generation Can Be Halted Early
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
On Efficient Scaling of GNNs via IO-Aware Layers Implementations
by: Fomina, Daria, et al.
Published: (2026)
by: Fomina, Daria, et al.
Published: (2026)
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025)
by: Gerasimov, Gleb, et al.
Published: (2025)
A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
by: Bredis, George, et al.
Published: (2025)
by: Bredis, George, et al.
Published: (2025)
Re:Frame -- Retrieving Experience From Associative Memory
by: Zelezetsky, Daniil, et al.
Published: (2025)
by: Zelezetsky, Daniil, et al.
Published: (2025)
Learning Shortest Paths with Generative Flow Networks
by: Morozov, Nikita, et al.
Published: (2026)
by: Morozov, Nikita, et al.
Published: (2026)
Performance Insights-based AI-driven Football Transfer Fee Prediction
by: Sulimov, Daniil
Published: (2024)
by: Sulimov, Daniil
Published: (2024)
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
by: Cherepanov, Egor, et al.
Published: (2026)
by: Cherepanov, Egor, et al.
Published: (2026)
Object-Centric Learning with Slot Mixture Module
by: Kirilenko, Daniil, et al.
Published: (2023)
by: Kirilenko, Daniil, et al.
Published: (2023)
The Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization
by: Kravatskiy, Alexey, et al.
Published: (2025)
by: Kravatskiy, Alexey, et al.
Published: (2025)
Universal time-series forecasting with mixture predictors
by: Ryabko, Daniil
Published: (2020)
by: Ryabko, Daniil
Published: (2020)
Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models
by: Chekalina, Viktoriia, et al.
Published: (2025)
by: Chekalina, Viktoriia, et al.
Published: (2025)
Guided Star-Shaped Masked Diffusion
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
Trust-Region Adaptive Policy Optimization
by: Su, Mingyu, et al.
Published: (2025)
by: Su, Mingyu, et al.
Published: (2025)
Extreme Region Policy Distillation
by: Chen, Changyu, et al.
Published: (2026)
by: Chen, Changyu, et al.
Published: (2026)
UPath: Universal Planner Across Topological Heterogeneity For Grid-Based Pathfinding
by: Ananikian, Aleksandr, et al.
Published: (2026)
by: Ananikian, Aleksandr, et al.
Published: (2026)
Bayesian Inverse Problems Meet Flow Matching: Efficient and Flexible Inference via Transformers
by: Sherki, Daniil, et al.
Published: (2025)
by: Sherki, Daniil, et al.
Published: (2025)
Generative Flow Networks as Entropy-Regularized RL
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Trust Regions Sell, But Who's Buying? Overlap Geometry as an Alternative Trust Region for Policy Optimization
by: Trivedi, Gaurish, et al.
Published: (2026)
by: Trivedi, Gaurish, et al.
Published: (2026)
Reinforcement learning for question answering in programming domain using public community scoring as a human feedback
by: Gorbatovski, Alexey, et al.
Published: (2024)
by: Gorbatovski, Alexey, et al.
Published: (2024)
Matrix Low-Rank Trust Region Policy Optimization
by: Rozada, Sergio, et al.
Published: (2024)
by: Rozada, Sergio, et al.
Published: (2024)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
by: Terekhov, Mikhail, et al.
Published: (2025)
by: Terekhov, Mikhail, et al.
Published: (2025)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Machine learning meets mass spectrometry: a focused perspective
by: Boiko, Daniil A., et al.
Published: (2024)
by: Boiko, Daniil A., et al.
Published: (2024)
Similar Items
-
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026) -
Steering LLM Reasoning Through Bias-Only Adaptation
by: Sinii, Viacheslav, et al.
Published: (2025) -
The Differences Between Direct Alignment Algorithms are a Blur
by: Gorbatovski, Alexey, et al.
Published: (2025) -
Learn Your Reference Model for Real Good Alignment
by: Gorbatovski, Alexey, et al.
Published: (2024) -
ESSA: Evolutionary Strategies for Scalable Alignment
by: Korotyshova, Daria, et al.
Published: (2025)