Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Hübotter, Jonas, Diaz-Bone, Leander, Hakimi, Ido, Krause, Andreas, Hardt, Moritz |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
by: Diaz-Bone, Leander, et al.
Published: (2025)
by: Diaz-Bone, Leander, et al.
Published: (2025)
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
by: Bertolissi, Ryo, et al.
Published: (2025)
by: Bertolissi, Ryo, et al.
Published: (2025)
Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning
by: Otth, Matthias, et al.
Published: (2025)
by: Otth, Matthias, et al.
Published: (2025)
Probabilistic Artificial Intelligence
by: Krause, Andreas, et al.
Published: (2025)
by: Krause, Andreas, et al.
Published: (2025)
Reinforcement Learning via Self-Distillation
by: Hübotter, Jonas, et al.
Published: (2026)
by: Hübotter, Jonas, et al.
Published: (2026)
Transductive Active Learning: Theory and Applications
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
by: Hübotter, Jonas, et al.
Published: (2025)
by: Hübotter, Jonas, et al.
Published: (2025)
Majority Voting for Code Generation
by: Launer, Tim, et al.
Published: (2026)
by: Launer, Tim, et al.
Published: (2026)
Active Few-Shot Fine-Tuning
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Thermodynamics of Reinforcement Learning Curricula
by: Adamczyk, Jacob, et al.
Published: (2026)
by: Adamczyk, Jacob, et al.
Published: (2026)
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
by: Melikidze, Davit, et al.
Published: (2026)
by: Melikidze, Davit, et al.
Published: (2026)
Test-time Offline Reinforcement Learning on Goal-related Experience
by: Bagatella, Marco, et al.
Published: (2025)
by: Bagatella, Marco, et al.
Published: (2025)
Train-before-Test Harmonizes Language Model Rankings
by: Zhang, Guanhua, et al.
Published: (2025)
by: Zhang, Guanhua, et al.
Published: (2025)
Aligning Language Models from User Interactions
by: Buening, Thomas Kleine, et al.
Published: (2026)
by: Buening, Thomas Kleine, et al.
Published: (2026)
Training on the Test Task Confounds Evaluation and Emergence
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
by: Yang, Daniel, et al.
Published: (2026)
by: Yang, Daniel, et al.
Published: (2026)
Computational Arbitrage in AI Model Markets
by: Olmedo, Ricardo, et al.
Published: (2026)
by: Olmedo, Ricardo, et al.
Published: (2026)
Offline Reinforcement Learning for Learning to Dispatch for Job Shop Scheduling
by: van Remmerden, Jesse, et al.
Published: (2024)
by: van Remmerden, Jesse, et al.
Published: (2024)
Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots
by: Li, Chenhao, et al.
Published: (2025)
by: Li, Chenhao, et al.
Published: (2025)
Simulation Priors for Data-Efficient Deep Learning
by: Treven, Lenart, et al.
Published: (2025)
by: Treven, Lenart, et al.
Published: (2025)
ReLA: Representation Learning and Aggregation for Job Scheduling with Reinforcement Learning
by: Kwan, Zhengyi, et al.
Published: (2026)
by: Kwan, Zhengyi, et al.
Published: (2026)
Learning Safety Constraints for Large Language Models
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
Distributionally Robust Model-based Reinforcement Learning with Large State Spaces
by: Ramesh, Shyam Sundhar, et al.
Published: (2023)
by: Ramesh, Shyam Sundhar, et al.
Published: (2023)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Sampling-Based Safe Reinforcement Learning
by: Vignola, Luca, et al.
Published: (2026)
by: Vignola, Luca, et al.
Published: (2026)
Contextual Bilevel Reinforcement Learning for Incentive Alignment
by: Thoma, Vinzenz, et al.
Published: (2024)
by: Thoma, Vinzenz, et al.
Published: (2024)
Target-Aligned Reinforcement Learning
by: Pleiss, Leonard S., et al.
Published: (2026)
by: Pleiss, Leonard S., et al.
Published: (2026)
Curricula for Learning Robust Policies with Factored State Representations in Changing Environments
by: Panayiotou, Panayiotis, et al.
Published: (2024)
by: Panayiotou, Panayiotis, et al.
Published: (2024)
Test-Time Tuned Language Models Enable End-to-end De Novo Molecular Structure Generation from MS/MS Spectra
by: Mismetti, Laura, et al.
Published: (2025)
by: Mismetti, Laura, et al.
Published: (2025)
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
by: Zhao, Chu, et al.
Published: (2026)
by: Zhao, Chu, et al.
Published: (2026)
What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time
by: Yan, Dong, et al.
Published: (2026)
by: Yan, Dong, et al.
Published: (2026)
Deep Reinforcement Learning Guided Improvement Heuristic for Job Shop Scheduling
by: Zhang, Cong, et al.
Published: (2022)
by: Zhang, Cong, et al.
Published: (2022)
Reinforcement Learning Teachers of Test Time Scaling
by: Cetin, Edoardo, et al.
Published: (2025)
by: Cetin, Edoardo, et al.
Published: (2025)
By Fair Means or Foul: Quantifying Collusion in a Market Simulation with Deep Reinforcement Learning
by: Schlechtinger, Michael, et al.
Published: (2024)
by: Schlechtinger, Michael, et al.
Published: (2024)
Diverse Projection Ensembles for Distributional Reinforcement Learning
by: Zanger, Moritz A., et al.
Published: (2023)
by: Zanger, Moritz A., et al.
Published: (2023)
Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
by: Sokota, Samuel, et al.
Published: (2025)
by: Sokota, Samuel, et al.
Published: (2025)
Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning
by: Han, Seungyub, et al.
Published: (2026)
by: Han, Seungyub, et al.
Published: (2026)
An Invitation to Deep Reinforcement Learning
by: Jaeger, Bernhard, et al.
Published: (2023)
by: Jaeger, Bernhard, et al.
Published: (2023)
Learning to Discover at Test Time
by: Yuksekgonul, Mert, et al.
Published: (2026)
by: Yuksekgonul, Mert, et al.
Published: (2026)
Similar Items
-
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
by: Diaz-Bone, Leander, et al.
Published: (2025) -
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
by: Hübotter, Jonas, et al.
Published: (2024) -
Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
by: Bertolissi, Ryo, et al.
Published: (2025) -
Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning
by: Otth, Matthias, et al.
Published: (2025) -
Probabilistic Artificial Intelligence
by: Krause, Andreas, et al.
Published: (2025)