Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Tomihari, Akiyoshi, Sato, Issei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective
by: Tomihari, Akiyoshi, et al.
Published: (2024)
by: Tomihari, Akiyoshi, et al.
Published: (2024)
Understanding Transformer Optimization via Gradient Heterogeneity
by: Tomihari, Akiyoshi, et al.
Published: (2025)
by: Tomihari, Akiyoshi, et al.
Published: (2025)
Learning Dynamics in RL Post-Training for Language Models
by: Tomihari, Akiyoshi
Published: (2026)
by: Tomihari, Akiyoshi
Published: (2026)
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
by: Tomihari, Akiyoshi, et al.
Published: (2025)
by: Tomihari, Akiyoshi, et al.
Published: (2025)
Top-Down Bayesian Posterior Sampling for Sum-Product Networks
by: Yokoi, Soma, et al.
Published: (2024)
by: Yokoi, Soma, et al.
Published: (2024)
On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding
by: Xu, Kevin, et al.
Published: (2024)
by: Xu, Kevin, et al.
Published: (2024)
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
by: Kajitsuka, Tokio, et al.
Published: (2023)
by: Kajitsuka, Tokio, et al.
Published: (2023)
Can Test-time Computation Mitigate Reproduction Bias in Neural Symbolic Regression?
by: Sato, Shun, et al.
Published: (2025)
by: Sato, Shun, et al.
Published: (2025)
Diagnosing and Mitigating System Bias in Self-Rewarding RL
by: Tan, Chuyi, et al.
Published: (2025)
by: Tan, Chuyi, et al.
Published: (2025)
End-to-End Training Induces Information Bottleneck through Layer-Role Differentiation: A Comparative Analysis with Layer-wise Training
by: Sakamoto, Keitaro, et al.
Published: (2024)
by: Sakamoto, Keitaro, et al.
Published: (2024)
Benign Overfitting in Token Selection of Attention Mechanism
by: Sakamoto, Keitaro, et al.
Published: (2024)
by: Sakamoto, Keitaro, et al.
Published: (2024)
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
by: Sakamoto, Keitaro, et al.
Published: (2025)
by: Sakamoto, Keitaro, et al.
Published: (2025)
Exploring Weight Balancing on Long-Tailed Recognition Problem
by: Hasegawa, Naoya, et al.
Published: (2023)
by: Hasegawa, Naoya, et al.
Published: (2023)
Multiplicative Logit Adjustment Approximates Neural-Collapse-Aware Decision Boundary Adjustment
by: Hasegawa, Naoya, et al.
Published: (2024)
by: Hasegawa, Naoya, et al.
Published: (2024)
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning
by: Chaudhary, Gaurav, et al.
Published: (2025)
by: Chaudhary, Gaurav, et al.
Published: (2025)
Max-pooling Network Revisited: Analyzing the Role of Semantic Probability in Multiple Instance Learning for Hallucination Detection
by: Fujikawa, Shota, et al.
Published: (2026)
by: Fujikawa, Shota, et al.
Published: (2026)
Fix Initial Codes and Iteratively Refine Textual Directions Toward Safe Multi-Turn Code Correction
by: Tanaka, Yuto, et al.
Published: (2026)
by: Tanaka, Yuto, et al.
Published: (2026)
Understanding Generalization in Physics Informed Models through Affine Variety Dimensions
by: Koshizuka, Takeshi, et al.
Published: (2025)
by: Koshizuka, Takeshi, et al.
Published: (2025)
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
by: Xu, Kevin, et al.
Published: (2025)
by: Xu, Kevin, et al.
Published: (2025)
Rethinking Associative Memory Mechanism in Induction Head
by: Wang, Shuo, et al.
Published: (2024)
by: Wang, Shuo, et al.
Published: (2024)
On the Optimal Memorization Capacity of Transformers
by: Kajitsuka, Tokio, et al.
Published: (2024)
by: Kajitsuka, Tokio, et al.
Published: (2024)
Bridging State and History Representations: Understanding Self-Predictive RL
by: Ni, Tianwei, et al.
Published: (2024)
by: Ni, Tianwei, et al.
Published: (2024)
A Formal Comparison Between Chain of Thought and Latent Thought
by: Xu, Kevin, et al.
Published: (2025)
by: Xu, Kevin, et al.
Published: (2025)
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Ensemble Distribution Distillation for Self-Supervised Human Activity Recognition
by: Nolan, Matthew, et al.
Published: (2025)
by: Nolan, Matthew, et al.
Published: (2025)
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
by: Yu, Xin, et al.
Published: (2026)
by: Yu, Xin, et al.
Published: (2026)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
by: Duan, Xintong, et al.
Published: (2025)
by: Duan, Xintong, et al.
Published: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
by: Zhu, Xuekai, et al.
Published: (2025)
by: Zhu, Xuekai, et al.
Published: (2025)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
by: Zhao, Siyan, et al.
Published: (2026)
by: Zhao, Siyan, et al.
Published: (2026)
Understanding the Expressivity and Trainability of Fourier Neural Operator: A Mean-Field Perspective
by: Koshizuka, Takeshi, et al.
Published: (2023)
by: Koshizuka, Takeshi, et al.
Published: (2023)
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
by: Muslimani, Calarina, et al.
Published: (2025)
by: Muslimani, Calarina, et al.
Published: (2025)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
by: Li, Gengsheng, et al.
Published: (2026)
by: Li, Gengsheng, et al.
Published: (2026)
Self-Distillation Enables Continual Learning
by: Shenfeld, Idan, et al.
Published: (2026)
by: Shenfeld, Idan, et al.
Published: (2026)
The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs
by: Zhao, Zibo, et al.
Published: (2026)
by: Zhao, Zibo, et al.
Published: (2026)
Mitigating Distribution Shift in Model-based Offline RL via Shifts-aware Reward Learning
by: Luo, Wang, et al.
Published: (2024)
by: Luo, Wang, et al.
Published: (2024)
When Sharpening Becomes Collapse: Sampling Bias and Semantic Coupling in RL with Verifiable Rewards
by: Fan, Mingyuan, et al.
Published: (2026)
by: Fan, Mingyuan, et al.
Published: (2026)
Reward Compatibility: A Framework for Inverse RL
by: Lazzati, Filippo, et al.
Published: (2025)
by: Lazzati, Filippo, et al.
Published: (2025)
Diversity-Rewarded CFG Distillation
by: Cideron, Geoffrey, et al.
Published: (2024)
by: Cideron, Geoffrey, et al.
Published: (2024)
Similar Items
-
Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective
by: Tomihari, Akiyoshi, et al.
Published: (2024) -
Understanding Transformer Optimization via Gradient Heterogeneity
by: Tomihari, Akiyoshi, et al.
Published: (2025) -
Learning Dynamics in RL Post-Training for Language Models
by: Tomihari, Akiyoshi
Published: (2026) -
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
by: Tomihari, Akiyoshi, et al.
Published: (2025) -
Top-Down Bayesian Posterior Sampling for Sum-Product Networks
by: Yokoi, Soma, et al.
Published: (2024)