How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Higuchi, Rei, Kawata, Ryotaro, Wachi, Akifumi, Takakura, Shokichi, Miyaguchi, Kohei, Suzuki, Taiji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
von: Takakura, Shokichi, et al.
Veröffentlicht: (2026)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2026)
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
von: Wachi, Akifumi, et al.
Veröffentlicht: (2026)
von: Wachi, Akifumi, et al.
Veröffentlicht: (2026)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
von: Takakura, Shokichi, et al.
Veröffentlicht: (2024)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2024)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023)
A Provable Approach for End-to-End Safe Reinforcement Learning
von: Wachi, Akifumi, et al.
Veröffentlicht: (2025)
von: Wachi, Akifumi, et al.
Veröffentlicht: (2025)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
Stepwise Alignment for Constrained Language Model Policy Optimization
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024)
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024)
When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
Target Return Optimizer for Multi-Game Decision Transformer
von: Tatematsu, Kensuke, et al.
Veröffentlicht: (2025)
von: Tatematsu, Kensuke, et al.
Veröffentlicht: (2025)
FedDuA: Doubly Adaptive Federated Learning
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
Accelerating Differentially Private Federated Learning via Adaptive Extrapolation
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
Path Learning with Trajectory Advantage Regression
von: Miyaguchi, Kohei
Veröffentlicht: (2025)
von: Miyaguchi, Kohei
Veröffentlicht: (2025)
From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
Optimal Variance and Covariance Estimation under Differential Privacy in the Add-Remove Model and Beyond
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
A Survey of Constraint Formulations in Safe Reinforcement Learning
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024)
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024)
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
von: Tran, Thien Q., et al.
Veröffentlicht: (2025)
von: Tran, Thien Q., et al.
Veröffentlicht: (2025)
Long-term Safe Reinforcement Learning with Binary Feedback
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024)
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024)
Differentially Private Sampling from Distributions via Wasserstein Projection
von: Takakura, Shokichi, et al.
Veröffentlicht: (2026)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2026)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
Cross-Process Defect Attribution using Potential Loss Analysis
von: Idé, Tsuyoshi, et al.
Veröffentlicht: (2025)
von: Idé, Tsuyoshi, et al.
Veröffentlicht: (2025)
DPSQL+: A Differentially Private SQL Library with a Minimum Frequency Rule
von: Matsumoto, Tomoya, et al.
Veröffentlicht: (2026)
von: Matsumoto, Tomoya, et al.
Veröffentlicht: (2026)
Flipping-based Policy for Chance-Constrained Markov Decision Processes
von: Shen, Xun, et al.
Veröffentlicht: (2024)
von: Shen, Xun, et al.
Veröffentlicht: (2024)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
von: Awano, Ryoya, et al.
Veröffentlicht: (2026)
von: Awano, Ryoya, et al.
Veröffentlicht: (2026)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization Strategies
von: Yan, Runze, et al.
Veröffentlicht: (2025)
von: Yan, Runze, et al.
Veröffentlicht: (2025)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025)
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025)
Deep Two-Way Matrix Reordering for Relational Data Analysis
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression
von: Kim, Juno, et al.
Veröffentlicht: (2025)
von: Kim, Juno, et al.
Veröffentlicht: (2025)
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
von: Kudo, Mikoto, et al.
Veröffentlicht: (2026)
von: Kudo, Mikoto, et al.
Veröffentlicht: (2026)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
Wafer Defect Root Cause Analysis with Partial Trajectory Regression
von: Miyaguchi, Kohei, et al.
Veröffentlicht: (2025)
von: Miyaguchi, Kohei, et al.
Veröffentlicht: (2025)
Quantifying the Optimization and Generalization Advantages of Graph Neural Networks Over Multilayer Perceptrons
von: Huang, Wei, et al.
Veröffentlicht: (2023)
von: Huang, Wei, et al.
Veröffentlicht: (2023)
How do Transformers perform In-Context Autoregressive Learning?
von: Sander, Michael E., et al.
Veröffentlicht: (2024)
von: Sander, Michael E., et al.
Veröffentlicht: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
von: Takakura, Shokichi, et al.
Veröffentlicht: (2026) -
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
von: Wachi, Akifumi, et al.
Veröffentlicht: (2026) -
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
von: Takakura, Shokichi, et al.
Veröffentlicht: (2024) -
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023) -
A Provable Approach for End-to-End Safe Reinforcement Learning
von: Wachi, Akifumi, et al.
Veröffentlicht: (2025)