Reinforcement Learning via Self-Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Hübotter, Jonas, Lübeck, Frederike, Behric, Lejs, Baumann, Anton, Bagatella, Marco, Marta, Daniel, Hakimi, Ido, Shenfeld, Idan, Buening, Thomas Kleine, Guestrin, Carlos, Krause, Andreas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Aligning Language Models from User Interactions
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2026)
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2026)
Majority Voting for Code Generation
di: Launer, Tim, et al.
Pubblicazione: (2026)
di: Launer, Tim, et al.
Pubblicazione: (2026)
Self-Distillation Enables Continual Learning
di: Shenfeld, Idan, et al.
Pubblicazione: (2026)
di: Shenfeld, Idan, et al.
Pubblicazione: (2026)
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
di: Hübotter, Jonas, et al.
Pubblicazione: (2024)
di: Hübotter, Jonas, et al.
Pubblicazione: (2024)
Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning
di: Otth, Matthias, et al.
Pubblicazione: (2025)
di: Otth, Matthias, et al.
Pubblicazione: (2025)
Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
di: Bertolissi, Ryo, et al.
Pubblicazione: (2025)
di: Bertolissi, Ryo, et al.
Pubblicazione: (2025)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
di: Diaz-Bone, Leander, et al.
Pubblicazione: (2025)
di: Diaz-Bone, Leander, et al.
Pubblicazione: (2025)
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
di: Hübotter, Jonas, et al.
Pubblicazione: (2025)
di: Hübotter, Jonas, et al.
Pubblicazione: (2025)
Active Fine-Tuning of Multi-Task Policies
di: Bagatella, Marco, et al.
Pubblicazione: (2024)
di: Bagatella, Marco, et al.
Pubblicazione: (2024)
Test-time Offline Reinforcement Learning on Goal-related Experience
di: Bagatella, Marco, et al.
Pubblicazione: (2025)
di: Bagatella, Marco, et al.
Pubblicazione: (2025)
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
di: Pásztor, Barna, et al.
Pubblicazione: (2025)
di: Pásztor, Barna, et al.
Pubblicazione: (2025)
Probabilistic Artificial Intelligence
di: Krause, Andreas, et al.
Pubblicazione: (2025)
di: Krause, Andreas, et al.
Pubblicazione: (2025)
Strategyproof Reinforcement Learning from Human Feedback
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2025)
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2025)
Environment Design for Inverse Reinforcement Learning
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2022)
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2022)
Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding
di: Shao, Daqian, et al.
Pubblicazione: (2025)
di: Shao, Daqian, et al.
Pubblicazione: (2025)
ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
di: Behric, Lejs Deen, et al.
Pubblicazione: (2025)
di: Behric, Lejs Deen, et al.
Pubblicazione: (2025)
Directed Exploration in Reinforcement Learning from Linear Temporal Logic
di: Bagatella, Marco, et al.
Pubblicazione: (2024)
di: Bagatella, Marco, et al.
Pubblicazione: (2024)
Soft Forward-Backward Representations for Zero-shot Reinforcement Learning with General Utilities
di: Bagatella, Marco, et al.
Pubblicazione: (2026)
di: Bagatella, Marco, et al.
Pubblicazione: (2026)
A Minimax Approach to Ad Hoc Teamwork
di: Villin, Victor, et al.
Pubblicazione: (2025)
di: Villin, Victor, et al.
Pubblicazione: (2025)
LITE: Efficiently Estimating Gaussian Probability of Maximality
di: Menet, Nicolas, et al.
Pubblicazione: (2025)
di: Menet, Nicolas, et al.
Pubblicazione: (2025)
RL's Razor: Why Online Reinforcement Learning Forgets Less
di: Shenfeld, Idan, et al.
Pubblicazione: (2025)
di: Shenfeld, Idan, et al.
Pubblicazione: (2025)
Adaptable Cardiovascular Disease Risk Prediction from Heterogeneous Data using Large Language Models
di: Lübeck, Frederike, et al.
Pubblicazione: (2025)
di: Lübeck, Frederike, et al.
Pubblicazione: (2025)
Strategic Linear Contextual Bandits
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2024)
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2024)
Optimistic Task Inference for Behavior Foundation Models
di: Rupf, Thomas, et al.
Pubblicazione: (2025)
di: Rupf, Thomas, et al.
Pubblicazione: (2025)
Active Few-Shot Fine-Tuning
di: Hübotter, Jonas, et al.
Pubblicazione: (2024)
di: Hübotter, Jonas, et al.
Pubblicazione: (2024)
Transductive Active Learning: Theory and Applications
di: Hübotter, Jonas, et al.
Pubblicazione: (2024)
di: Hübotter, Jonas, et al.
Pubblicazione: (2024)
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
di: Shenfeld, Idan, et al.
Pubblicazione: (2023)
di: Shenfeld, Idan, et al.
Pubblicazione: (2023)
Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
di: Hübotter, Jonas, et al.
Pubblicazione: (2025)
di: Hübotter, Jonas, et al.
Pubblicazione: (2025)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
di: Hilel, Almog, et al.
Pubblicazione: (2025)
di: Hilel, Almog, et al.
Pubblicazione: (2025)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
di: Damani, Mehul, et al.
Pubblicazione: (2024)
di: Damani, Mehul, et al.
Pubblicazione: (2024)
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
di: Baur, Raphaël, et al.
Pubblicazione: (2026)
di: Baur, Raphaël, et al.
Pubblicazione: (2026)
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
di: Melikidze, Davit, et al.
Pubblicazione: (2026)
di: Melikidze, Davit, et al.
Pubblicazione: (2026)
Language Model Personalization via Reward Factorization
di: Shenfeld, Idan, et al.
Pubblicazione: (2025)
di: Shenfeld, Idan, et al.
Pubblicazione: (2025)
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
di: Ankile, Lars, et al.
Pubblicazione: (2024)
di: Ankile, Lars, et al.
Pubblicazione: (2024)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
di: Yang, Daniel, et al.
Pubblicazione: (2026)
di: Yang, Daniel, et al.
Pubblicazione: (2026)
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
di: Aminian, Gholamali, et al.
Pubblicazione: (2025)
di: Aminian, Gholamali, et al.
Pubblicazione: (2025)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
di: Puri, Isha, et al.
Pubblicazione: (2026)
di: Puri, Isha, et al.
Pubblicazione: (2026)
Zero-Shot Offline Imitation Learning via Optimal Transport
di: Rupf, Thomas, et al.
Pubblicazione: (2024)
di: Rupf, Thomas, et al.
Pubblicazione: (2024)
From Imitation to Refinement -- Residual RL for Precise Assembly
di: Ankile, Lars, et al.
Pubblicazione: (2024)
di: Ankile, Lars, et al.
Pubblicazione: (2024)
Value Augmented Sampling for Language Model Alignment and Personalization
di: Han, Seungwook, et al.
Pubblicazione: (2024)
di: Han, Seungwook, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Aligning Language Models from User Interactions
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2026) -
Majority Voting for Code Generation
di: Launer, Tim, et al.
Pubblicazione: (2026) -
Self-Distillation Enables Continual Learning
di: Shenfeld, Idan, et al.
Pubblicazione: (2026) -
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
di: Hübotter, Jonas, et al.
Pubblicazione: (2024) -
Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning
di: Otth, Matthias, et al.
Pubblicazione: (2025)