General Intelligence Requires Reward-based Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Seungwook, Pari, Jyothish, Gershman, Samuel J., Agrawal, Pulkit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Collective Model Intelligence Requires Compatible Specialization
von: Pari, Jyothish, et al.
Veröffentlicht: (2024)
von: Pari, Jyothish, et al.
Veröffentlicht: (2024)
RL's Razor: Why Online Reinforcement Learning Forgets Less
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
von: Reuss, Moritz, et al.
Veröffentlicht: (2024)
von: Reuss, Moritz, et al.
Veröffentlicht: (2024)
Self-Adapting Language Models
von: Zweiger, Adam, et al.
Veröffentlicht: (2025)
von: Zweiger, Adam, et al.
Veröffentlicht: (2025)
Few-Shot Task Learning through Inverse Generative Modeling
von: Netanyahu, Aviv, et al.
Veröffentlicht: (2024)
von: Netanyahu, Aviv, et al.
Veröffentlicht: (2024)
Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective
von: Han, Seungwook, et al.
Veröffentlicht: (2024)
von: Han, Seungwook, et al.
Veröffentlicht: (2024)
Training Language Models via Neural Cellular Automata
von: Lee, Dan, et al.
Veröffentlicht: (2026)
von: Lee, Dan, et al.
Veröffentlicht: (2026)
Value Augmented Sampling for Language Model Alignment and Personalization
von: Han, Seungwook, et al.
Veröffentlicht: (2024)
von: Han, Seungwook, et al.
Veröffentlicht: (2024)
Language Model Personalization via Reward Factorization
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
H2LooP Spark Preview: Continual Pretraining of Large Language Models for Low-Level Embedded Systems Code
von: Singh, Amit, et al.
Veröffentlicht: (2026)
von: Singh, Amit, et al.
Veröffentlicht: (2026)
Fast weight programming and linear transformers: from machine learning to neurobiology
von: Irie, Kazuki, et al.
Veröffentlicht: (2025)
von: Irie, Kazuki, et al.
Veröffentlicht: (2025)
Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
von: Irie, Kazuki, et al.
Veröffentlicht: (2025)
von: Irie, Kazuki, et al.
Veröffentlicht: (2025)
Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules
von: Vastola, John J., et al.
Veröffentlicht: (2025)
von: Vastola, John J., et al.
Veröffentlicht: (2025)
A Variational Manifold Embedding Framework for Nonlinear Dimensionality Reduction
von: Vastola, John J., et al.
Veröffentlicht: (2025)
von: Vastola, John J., et al.
Veröffentlicht: (2025)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
von: Akyürek, Ekin, et al.
Veröffentlicht: (2024)
von: Akyürek, Ekin, et al.
Veröffentlicht: (2024)
Artificial intelligence for science: The easy and hard problems
von: Battleday, Ruairidh M., et al.
Veröffentlicht: (2024)
von: Battleday, Ruairidh M., et al.
Veröffentlicht: (2024)
Leveraging Manifold Embeddings for Enhanced Graph Transformer Representations and Learning
von: Jyothish, Ankit, et al.
Veröffentlicht: (2025)
von: Jyothish, Ankit, et al.
Veröffentlicht: (2025)
A circuit for predicting hierarchical structure in-context in Large Language Models
von: Saanum, Tankred, et al.
Veröffentlicht: (2025)
von: Saanum, Tankred, et al.
Veröffentlicht: (2025)
Successor-Predecessor Intrinsic Exploration
von: Yu, Changmin, et al.
Veröffentlicht: (2023)
von: Yu, Changmin, et al.
Veröffentlicht: (2023)
Pretraining Generative Flow Networks with Inexpensive Rewards for Molecular Graph Generation
von: Pandey, Mohit, et al.
Veröffentlicht: (2025)
von: Pandey, Mohit, et al.
Veröffentlicht: (2025)
Self-Distillation Enables Continual Learning
von: Shenfeld, Idan, et al.
Veröffentlicht: (2026)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2026)
Key-value memory in the brain
von: Gershman, Samuel J., et al.
Veröffentlicht: (2025)
von: Gershman, Samuel J., et al.
Veröffentlicht: (2025)
Regret Tail Characterization of Optimal Bandit Algorithms with Generic Rewards
von: Panda, Subhodip, et al.
Veröffentlicht: (2026)
von: Panda, Subhodip, et al.
Veröffentlicht: (2026)
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
von: Ankile, Lars, et al.
Veröffentlicht: (2024)
von: Ankile, Lars, et al.
Veröffentlicht: (2024)
Automatic Environment Shaping is the Next Frontier in RL
von: Park, Younghyo, et al.
Veröffentlicht: (2024)
von: Park, Younghyo, et al.
Veröffentlicht: (2024)
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
von: Shenfeld, Idan, et al.
Veröffentlicht: (2023)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2023)
FAST-Q: Fast-track Exploration with Adversarially Balanced State Representations for Counterfactual Action Estimation in Offline Reinforcement Learning
von: Agrawal, Pulkit, et al.
Veröffentlicht: (2025)
von: Agrawal, Pulkit, et al.
Veröffentlicht: (2025)
Going Beyond Heuristics by Imposing Policy Improvement as a Constraint
von: Lee, Chi-Chang, et al.
Veröffentlicht: (2025)
von: Lee, Chi-Chang, et al.
Veröffentlicht: (2025)
Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View
von: Heo, Gyuryang, et al.
Veröffentlicht: (2026)
von: Heo, Gyuryang, et al.
Veröffentlicht: (2026)
Fast MoE Inference via Predictive Prefetching and Expert Replication
von: Jyothish, Ankit, et al.
Veröffentlicht: (2026)
von: Jyothish, Ankit, et al.
Veröffentlicht: (2026)
Grokking as the Transition from Lazy to Rich Training Dynamics
von: Kumar, Tanishq, et al.
Veröffentlicht: (2023)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2023)
From Imitation to Refinement -- Residual RL for Precise Assembly
von: Ankile, Lars, et al.
Veröffentlicht: (2024)
von: Ankile, Lars, et al.
Veröffentlicht: (2024)
Random Latent Exploration for Deep Reinforcement Learning
von: Mahankali, Srinath, et al.
Veröffentlicht: (2024)
von: Mahankali, Srinath, et al.
Veröffentlicht: (2024)
GFlowNet Pretraining with Inexpensive Rewards
von: Pandey, Mohit, et al.
Veröffentlicht: (2024)
von: Pandey, Mohit, et al.
Veröffentlicht: (2024)
Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines
von: Carvalho, Wilka, et al.
Veröffentlicht: (2025)
von: Carvalho, Wilka, et al.
Veröffentlicht: (2025)
Bridging the Sim-to-Real Gap for Athletic Loco-Manipulation
von: Fey, Nolan, et al.
Veröffentlicht: (2025)
von: Fey, Nolan, et al.
Veröffentlicht: (2025)
SoftMimic: Learning Compliant Whole-body Control from Examples
von: Margolis, Gabriel B., et al.
Veröffentlicht: (2025)
von: Margolis, Gabriel B., et al.
Veröffentlicht: (2025)
Vegetable Peeling: A Case Study in Constrained Dexterous Manipulation
von: Chen, Tao, et al.
Veröffentlicht: (2024)
von: Chen, Tao, et al.
Veröffentlicht: (2024)
Explainable and Interpretable Forecasts on Non-Smooth Multivariate Time Series for Responsible Gameplay
von: Jagirdar, Hussain, et al.
Veröffentlicht: (2025)
von: Jagirdar, Hussain, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Collective Model Intelligence Requires Compatible Specialization
von: Pari, Jyothish, et al.
Veröffentlicht: (2024) -
RL's Razor: Why Online Reinforcement Learning Forgets Less
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025) -
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
von: Reuss, Moritz, et al.
Veröffentlicht: (2024) -
Self-Adapting Language Models
von: Zweiger, Adam, et al.
Veröffentlicht: (2025) -
Few-Shot Task Learning through Inverse Generative Modeling
von: Netanyahu, Aviv, et al.
Veröffentlicht: (2024)