Q-Probe: A Lightweight Approach to Reward Maximization for Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Kenneth, Jelassi, Samy, Zhang, Hugh, Kakade, Sham, Wattenberg, Martin, Brandfonbrener, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Repeat After Me: Transformers are Better than State Space Models at Copying
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
Universal Length Generalization with Turing Programs
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
The Role of Sparsity for Length Generalization in Transformers
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
GQ-VAE: A gated quantized VAE for learning variable length tokens
von: Datta, Theo, et al.
Veröffentlicht: (2025)
von: Datta, Theo, et al.
Veröffentlicht: (2025)
Deconstructing What Makes a Good Optimizer for Language Models
von: Zhao, Rosie, et al.
Veröffentlicht: (2024)
von: Zhao, Rosie, et al.
Veröffentlicht: (2024)
Mixture of Parrots: Experts improve memorization more than reasoning
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2024)
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2024)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
von: Jelassi, Samy, et al.
Veröffentlicht: (2026)
von: Jelassi, Samy, et al.
Veröffentlicht: (2026)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
Collective Model Intelligence Requires Compatible Specialization
von: Pari, Jyothish, et al.
Veröffentlicht: (2024)
von: Pari, Jyothish, et al.
Veröffentlicht: (2024)
Scaling Reward Modeling without Human Supervision
von: Fan, Jingxuan, et al.
Veröffentlicht: (2026)
von: Fan, Jingxuan, et al.
Veröffentlicht: (2026)
How Does Overparameterization Affect Features?
von: Duzgun, Ahmet Cagri, et al.
Veröffentlicht: (2024)
von: Duzgun, Ahmet Cagri, et al.
Veröffentlicht: (2024)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
von: Qin, Tian, et al.
Veröffentlicht: (2025)
von: Qin, Tian, et al.
Veröffentlicht: (2025)
SOAP: Improving and Stabilizing Shampoo using Adam
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
von: Jin, Jikai, et al.
Veröffentlicht: (2025)
von: Jin, Jikai, et al.
Veröffentlicht: (2025)
Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn Planner
von: Li, Kenneth, et al.
Veröffentlicht: (2024)
von: Li, Kenneth, et al.
Veröffentlicht: (2024)
Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions
von: Lee, Andrew, et al.
Veröffentlicht: (2026)
von: Lee, Andrew, et al.
Veröffentlicht: (2026)
Eliminating Position Bias of Language Models: A Mechanistic Approach
von: Wang, Ziqi, et al.
Veröffentlicht: (2024)
von: Wang, Ziqi, et al.
Veröffentlicht: (2024)
Learning Hidden Markov Models Using Conditional Samples
von: Kakade, Sham M., et al.
Veröffentlicht: (2023)
von: Kakade, Sham M., et al.
Veröffentlicht: (2023)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
von: Song, Yuda, et al.
Veröffentlicht: (2024)
von: Song, Yuda, et al.
Veröffentlicht: (2024)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
von: Morwani, Depen, et al.
Veröffentlicht: (2025)
von: Morwani, Depen, et al.
Veröffentlicht: (2025)
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
von: Li, Kenneth, et al.
Veröffentlicht: (2023)
von: Li, Kenneth, et al.
Veröffentlicht: (2023)
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
von: Bansal, Rachit, et al.
Veröffentlicht: (2025)
von: Bansal, Rachit, et al.
Veröffentlicht: (2025)
When Bad Data Leads to Good Models
von: Li, Kenneth, et al.
Veröffentlicht: (2025)
von: Li, Kenneth, et al.
Veröffentlicht: (2025)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
von: Abreu, Natalie, et al.
Veröffentlicht: (2025)
von: Abreu, Natalie, et al.
Veröffentlicht: (2025)
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2024)
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2024)
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
von: Li, Kenneth, et al.
Veröffentlicht: (2024)
von: Li, Kenneth, et al.
Veröffentlicht: (2024)
Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion Training
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2026)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2026)
Selective Underfitting in Diffusion Models
von: Song, Kiwhan, et al.
Veröffentlicht: (2025)
von: Song, Kiwhan, et al.
Veröffentlicht: (2025)
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2025)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2025)
Shared Global and Local Geometry of Language Model Embeddings
von: Lee, Andrew, et al.
Veröffentlicht: (2025)
von: Lee, Andrew, et al.
Veröffentlicht: (2025)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
von: Kou, Yiwen, et al.
Veröffentlicht: (2024)
von: Kou, Yiwen, et al.
Veröffentlicht: (2024)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
von: Liu, Bingbin, et al.
Veröffentlicht: (2025)
von: Liu, Bingbin, et al.
Veröffentlicht: (2025)
Random Scaling of Emergent Capabilities
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
Relational Composition in Neural Networks: A Survey and Call to Action
von: Wattenberg, Martin, et al.
Veröffentlicht: (2024)
von: Wattenberg, Martin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Repeat After Me: Transformers are Better than State Space Models at Copying
von: Jelassi, Samy, et al.
Veröffentlicht: (2024) -
Universal Length Generalization with Turing Programs
von: Hou, Kaiying, et al.
Veröffentlicht: (2024) -
The Role of Sparsity for Length Generalization in Transformers
von: Golowich, Noah, et al.
Veröffentlicht: (2025) -
GQ-VAE: A gated quantized VAE for learning variable length tokens
von: Datta, Theo, et al.
Veröffentlicht: (2025) -
Deconstructing What Makes a Good Optimizer for Language Models
von: Zhao, Rosie, et al.
Veröffentlicht: (2024)