Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
Fuente:
arXiv
Saved in:
| Main Authors: | Takakura, Shokichi, Wachi, Akifumi, Higuchi, Rei, Miyaguchi, Kohei, Suzuki, Taiji |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
by: Higuchi, Rei, et al.
Published: (2026)
by: Higuchi, Rei, et al.
Published: (2026)
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
by: Wachi, Akifumi, et al.
Published: (2026)
by: Wachi, Akifumi, et al.
Published: (2026)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
by: Takakura, Shokichi, et al.
Published: (2024)
by: Takakura, Shokichi, et al.
Published: (2024)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
by: Takakura, Shokichi, et al.
Published: (2023)
by: Takakura, Shokichi, et al.
Published: (2023)
A Provable Approach for End-to-End Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2025)
by: Wachi, Akifumi, et al.
Published: (2025)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
by: Nishikawa, Naoki, et al.
Published: (2025)
by: Nishikawa, Naoki, et al.
Published: (2025)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
Stepwise Alignment for Constrained Language Model Policy Optimization
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
Accelerating Differentially Private Federated Learning via Adaptive Extrapolation
by: Takakura, Shokichi, et al.
Published: (2025)
by: Takakura, Shokichi, et al.
Published: (2025)
Differentially Private Sampling from Distributions via Wasserstein Projection
by: Takakura, Shokichi, et al.
Published: (2026)
by: Takakura, Shokichi, et al.
Published: (2026)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
by: Wakayama, Tomoya, et al.
Published: (2025)
by: Wakayama, Tomoya, et al.
Published: (2025)
Target Return Optimizer for Multi-Game Decision Transformer
by: Tatematsu, Kensuke, et al.
Published: (2025)
by: Tatematsu, Kensuke, et al.
Published: (2025)
FedDuA: Doubly Adaptive Federated Learning
by: Takakura, Shokichi, et al.
Published: (2025)
by: Takakura, Shokichi, et al.
Published: (2025)
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
by: Tran, Thien Q., et al.
Published: (2025)
by: Tran, Thien Q., et al.
Published: (2025)
Path Learning with Trajectory Advantage Regression
by: Miyaguchi, Kohei
Published: (2025)
by: Miyaguchi, Kohei
Published: (2025)
Optimal Variance and Covariance Estimation under Differential Privacy in the Add-Remove Model and Beyond
by: Takakura, Shokichi, et al.
Published: (2025)
by: Takakura, Shokichi, et al.
Published: (2025)
A Survey of Constraint Formulations in Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
DPSQL+: A Differentially Private SQL Library with a Minimum Frequency Rule
by: Matsumoto, Tomoya, et al.
Published: (2026)
by: Matsumoto, Tomoya, et al.
Published: (2026)
Long-term Safe Reinforcement Learning with Binary Feedback
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment
by: Kusaka, Shigeki, et al.
Published: (2025)
by: Kusaka, Shigeki, et al.
Published: (2025)
Sequence-Aware Inline Measurement Attribution for Good-Bad Wafer Diagnosis
by: Miyaguchi, Kohei, et al.
Published: (2025)
by: Miyaguchi, Kohei, et al.
Published: (2025)
Cross-Process Defect Attribution using Potential Loss Analysis
by: Idé, Tsuyoshi, et al.
Published: (2025)
by: Idé, Tsuyoshi, et al.
Published: (2025)
When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
by: Kawata, Ryotaro, et al.
Published: (2025)
by: Kawata, Ryotaro, et al.
Published: (2025)
Flipping-based Policy for Chance-Constrained Markov Decision Processes
by: Shen, Xun, et al.
Published: (2024)
by: Shen, Xun, et al.
Published: (2024)
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
by: Kudo, Mikoto, et al.
Published: (2026)
by: Kudo, Mikoto, et al.
Published: (2026)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
by: Awano, Ryoya, et al.
Published: (2026)
by: Awano, Ryoya, et al.
Published: (2026)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
by: Kawata, Ryotaro, et al.
Published: (2026)
by: Kawata, Ryotaro, et al.
Published: (2026)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
by: Nishikawa, Naoki, et al.
Published: (2024)
by: Nishikawa, Naoki, et al.
Published: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Test time training enhances in-context learning of nonlinear functions
by: Kuwataka, Kento, et al.
Published: (2025)
by: Kuwataka, Kento, et al.
Published: (2025)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
by: Oh, Junsoo, et al.
Published: (2025)
by: Oh, Junsoo, et al.
Published: (2025)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
by: Jiang, Jiarui, et al.
Published: (2025)
by: Jiang, Jiarui, et al.
Published: (2025)
Wafer Defect Root Cause Analysis with Partial Trajectory Regression
by: Miyaguchi, Kohei, et al.
Published: (2025)
by: Miyaguchi, Kohei, et al.
Published: (2025)
Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization Strategies
by: Yan, Runze, et al.
Published: (2025)
by: Yan, Runze, et al.
Published: (2025)
What is the Alignment Objective of GRPO?
by: Vojnovic, Milan, et al.
Published: (2025)
by: Vojnovic, Milan, et al.
Published: (2025)
Transformers are Minimax Optimal Nonparametric In-Context Learners
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Similar Items
-
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
by: Higuchi, Rei, et al.
Published: (2026) -
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
by: Wachi, Akifumi, et al.
Published: (2026) -
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
by: Takakura, Shokichi, et al.
Published: (2024) -
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
by: Takakura, Shokichi, et al.
Published: (2023) -
A Provable Approach for End-to-End Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2025)