Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jelassi, Samy, Kwun, Mujin, Zhao, Rosie, Li, Yuanzhi, Fusi, Nicolo, Du, Yilun, Kakade, Sham M., Domingo-Enrich, Carles |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
by: Su, Huangyuan, et al.
Published: (2025)
by: Su, Huangyuan, et al.
Published: (2025)
SOAP: Improving and Stabilizing Shampoo using Adam
by: Vyas, Nikhil, et al.
Published: (2024)
by: Vyas, Nikhil, et al.
Published: (2024)
Q-Probe: A Lightweight Approach to Reward Maximization for Language Models
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
How Does Overparameterization Affect Features?
by: Duzgun, Ahmet Cagri, et al.
Published: (2024)
by: Duzgun, Ahmet Cagri, et al.
Published: (2024)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
Any-Order Flexible Length Masked Diffusion
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control
by: Domingo-Enrich, Carles, et al.
Published: (2026)
by: Domingo-Enrich, Carles, et al.
Published: (2026)
LOTION: Smoothing the Optimization Landscape for Quantized Training
by: Kwun, Mujin, et al.
Published: (2025)
by: Kwun, Mujin, et al.
Published: (2025)
A Taxonomy of Loss Functions for Stochastic Optimal Control
by: Domingo-Enrich, Carles
Published: (2024)
by: Domingo-Enrich, Carles
Published: (2024)
Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control
by: Domingo-Enrich, Carles, et al.
Published: (2024)
by: Domingo-Enrich, Carles, et al.
Published: (2024)
Deconstructing What Makes a Good Optimizer for Language Models
by: Zhao, Rosie, et al.
Published: (2024)
by: Zhao, Rosie, et al.
Published: (2024)
A unified perspective on fine-tuning and sampling with diffusion and flow models
by: Domingo-Enrich, Carles, et al.
Published: (2026)
by: Domingo-Enrich, Carles, et al.
Published: (2026)
Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models
by: Bergmeister, Andreas, et al.
Published: (2026)
by: Bergmeister, Andreas, et al.
Published: (2026)
Feature emergence via margin maximization: case studies in algebraic tasks
by: Morwani, Depen, et al.
Published: (2023)
by: Morwani, Depen, et al.
Published: (2023)
Mixture of Parrots: Experts improve memorization more than reasoning
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Adapting Language Models via Token Translation
by: Feng, Zhili, et al.
Published: (2024)
by: Feng, Zhili, et al.
Published: (2024)
Stochastic Optimal Control Matching
by: Domingo-Enrich, Carles, et al.
Published: (2023)
by: Domingo-Enrich, Carles, et al.
Published: (2023)
Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
by: Wang, Runqian, et al.
Published: (2025)
by: Wang, Runqian, et al.
Published: (2025)
Selective Underfitting in Diffusion Models
by: Song, Kiwhan, et al.
Published: (2025)
by: Song, Kiwhan, et al.
Published: (2025)
Cheap Permutation Testing
by: Domingo-Enrich, Carles, et al.
Published: (2025)
by: Domingo-Enrich, Carles, et al.
Published: (2025)
Compress Then Test: Powerful Kernel Testing in Near-linear Time
by: Domingo-Enrich, Carles, et al.
Published: (2023)
by: Domingo-Enrich, Carles, et al.
Published: (2023)
Value Gradient Guidance for Flow Matching Alignment
by: Liu, Zhen, et al.
Published: (2025)
by: Liu, Zhen, et al.
Published: (2025)
Random Scaling of Emergent Capabilities
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Self-Improving Language Models with Bidirectional Evolutionary Search
by: Xu, Guowei, et al.
Published: (2026)
by: Xu, Guowei, et al.
Published: (2026)
Collective Model Intelligence Requires Compatible Specialization
by: Pari, Jyothish, et al.
Published: (2024)
by: Pari, Jyothish, et al.
Published: (2024)
Beyond Implicit Bias: The Insignificance of SGD Noise in Online Learning
by: Vyas, Nikhil, et al.
Published: (2023)
by: Vyas, Nikhil, et al.
Published: (2023)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
by: Zhang, Hanlin, et al.
Published: (2026)
by: Zhang, Hanlin, et al.
Published: (2026)
IGDA: Interactive Graph Discovery through Large Language Model Agents
by: Havrilla, Alex, et al.
Published: (2025)
by: Havrilla, Alex, et al.
Published: (2025)
EvoLM: In Search of Lost Language Model Training Dynamics
by: Qi, Zhenting, et al.
Published: (2025)
by: Qi, Zhenting, et al.
Published: (2025)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
by: Jin, Jikai, et al.
Published: (2025)
by: Jin, Jikai, et al.
Published: (2025)
Rare Event Analysis via Stochastic Optimal Control
by: Du, Yuanqi, et al.
Published: (2026)
by: Du, Yuanqi, et al.
Published: (2026)
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
by: Kou, Yiwen, et al.
Published: (2024)
by: Kou, Yiwen, et al.
Published: (2024)
Cognitive models can reveal interpretable value trade-offs in language models
by: Murthy, Sonia K., et al.
Published: (2025)
by: Murthy, Sonia K., et al.
Published: (2025)
Peer-Predictive Self-Training for Language Model Reasoning
by: Feng, Shi, et al.
Published: (2026)
by: Feng, Shi, et al.
Published: (2026)
Similar Items
-
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026) -
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025) -
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024) -
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024) -
Characterization and Mitigation of Training Instabilities in Microscaling Formats
by: Su, Huangyuan, et al.
Published: (2025)