Collective Model Intelligence Requires Compatible Specialization
Fuente:
arXiv
Saved in:
| Main Authors: | Pari, Jyothish, Jelassi, Samy, Agrawal, Pulkit |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
General Intelligence Requires Reward-based Pretraining
by: Han, Seungwook, et al.
Published: (2025)
by: Han, Seungwook, et al.
Published: (2025)
RL's Razor: Why Online Reinforcement Learning Forgets Less
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
by: Reuss, Moritz, et al.
Published: (2024)
by: Reuss, Moritz, et al.
Published: (2024)
Self-Adapting Language Models
by: Zweiger, Adam, et al.
Published: (2025)
by: Zweiger, Adam, et al.
Published: (2025)
Few-Shot Task Learning through Inverse Generative Modeling
by: Netanyahu, Aviv, et al.
Published: (2024)
by: Netanyahu, Aviv, et al.
Published: (2024)
How Does Overparameterization Affect Features?
by: Duzgun, Ahmet Cagri, et al.
Published: (2024)
by: Duzgun, Ahmet Cagri, et al.
Published: (2024)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
Q-Probe: A Lightweight Approach to Reward Maximization for Language Models
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
Language Model Personalization via Reward Factorization
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Leveraging Manifold Embeddings for Enhanced Graph Transformer Representations and Learning
by: Jyothish, Ankit, et al.
Published: (2025)
by: Jyothish, Ankit, et al.
Published: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
by: Jelassi, Samy, et al.
Published: (2026)
by: Jelassi, Samy, et al.
Published: (2026)
Self-Distillation Enables Continual Learning
by: Shenfeld, Idan, et al.
Published: (2026)
by: Shenfeld, Idan, et al.
Published: (2026)
H2LooP Spark Preview: Continual Pretraining of Large Language Models for Low-Level Embedded Systems Code
by: Singh, Amit, et al.
Published: (2026)
by: Singh, Amit, et al.
Published: (2026)
Training Language Models via Neural Cellular Automata
by: Lee, Dan, et al.
Published: (2026)
by: Lee, Dan, et al.
Published: (2026)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
by: Akyürek, Ekin, et al.
Published: (2024)
by: Akyürek, Ekin, et al.
Published: (2024)
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
Automatic Environment Shaping is the Next Frontier in RL
by: Park, Younghyo, et al.
Published: (2024)
by: Park, Younghyo, et al.
Published: (2024)
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
by: Shenfeld, Idan, et al.
Published: (2023)
by: Shenfeld, Idan, et al.
Published: (2023)
FAST-Q: Fast-track Exploration with Adversarially Balanced State Representations for Counterfactual Action Estimation in Offline Reinforcement Learning
by: Agrawal, Pulkit, et al.
Published: (2025)
by: Agrawal, Pulkit, et al.
Published: (2025)
Going Beyond Heuristics by Imposing Policy Improvement as a Constraint
by: Lee, Chi-Chang, et al.
Published: (2025)
by: Lee, Chi-Chang, et al.
Published: (2025)
Value Augmented Sampling for Language Model Alignment and Personalization
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
Fast MoE Inference via Predictive Prefetching and Expert Replication
by: Jyothish, Ankit, et al.
Published: (2026)
by: Jyothish, Ankit, et al.
Published: (2026)
From Imitation to Refinement -- Residual RL for Precise Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
Random Latent Exploration for Deep Reinforcement Learning
by: Mahankali, Srinath, et al.
Published: (2024)
by: Mahankali, Srinath, et al.
Published: (2024)
Bridging the Sim-to-Real Gap for Athletic Loco-Manipulation
by: Fey, Nolan, et al.
Published: (2025)
by: Fey, Nolan, et al.
Published: (2025)
SoftMimic: Learning Compliant Whole-body Control from Examples
by: Margolis, Gabriel B., et al.
Published: (2025)
by: Margolis, Gabriel B., et al.
Published: (2025)
Vegetable Peeling: A Case Study in Constrained Dexterous Manipulation
by: Chen, Tao, et al.
Published: (2024)
by: Chen, Tao, et al.
Published: (2024)
Explainable and Interpretable Forecasts on Non-Smooth Multivariate Time Series for Responsible Gameplay
by: Jagirdar, Hussain, et al.
Published: (2025)
by: Jagirdar, Hussain, et al.
Published: (2025)
Learning Multimodal Behaviors from Scratch with Diffusion Policy Gradient
by: Li, Zechu, et al.
Published: (2024)
by: Li, Zechu, et al.
Published: (2024)
Mixture of Parrots: Experts improve memorization more than reasoning
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Learning Force Control for Legged Manipulation
by: Portela, Tifanny, et al.
Published: (2024)
by: Portela, Tifanny, et al.
Published: (2024)
Separating Intrinsic Ambiguity from Estimation Uncertainty in Deep Generative Models for Linear Inverse Problems
by: Guo, Yuxin, et al.
Published: (2026)
by: Guo, Yuxin, et al.
Published: (2026)
Similar Items
-
General Intelligence Requires Reward-based Pretraining
by: Han, Seungwook, et al.
Published: (2025) -
RL's Razor: Why Online Reinforcement Learning Forgets Less
by: Shenfeld, Idan, et al.
Published: (2025) -
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
by: Reuss, Moritz, et al.
Published: (2024) -
Self-Adapting Language Models
by: Zweiger, Adam, et al.
Published: (2025) -
Few-Shot Task Learning through Inverse Generative Modeling
by: Netanyahu, Aviv, et al.
Published: (2024)