Gespeichert in:
| Hauptverfasser: | Xie, Zeke, Xu, Zhiqiang, Zhang, Jingzhao, Sato, Issei, Sugiyama, Masashi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2020
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2011.11152 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding Transformer Optimization via Gradient Heterogeneity
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
von: Xu, Kevin, et al.
Veröffentlicht: (2025)
von: Xu, Kevin, et al.
Veröffentlicht: (2025)
A Formal Comparison Between Chain of Thought and Latent Thought
von: Xu, Kevin, et al.
Veröffentlicht: (2025)
von: Xu, Kevin, et al.
Veröffentlicht: (2025)
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
von: Chen, Lesi, et al.
Veröffentlicht: (2023)
von: Chen, Lesi, et al.
Veröffentlicht: (2023)
Fix Initial Codes and Iteratively Refine Textual Directions Toward Safe Multi-Turn Code Correction
von: Tanaka, Yuto, et al.
Veröffentlicht: (2026)
von: Tanaka, Yuto, et al.
Veröffentlicht: (2026)
Decoupled Weight Decay for Any $p$ Norm
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024)
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024)
Learning Robust Diffusion Models from Imprecise Supervision
von: Wu, Dong-Dong, et al.
Veröffentlicht: (2025)
von: Wu, Dong-Dong, et al.
Veröffentlicht: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
von: Cai, Xin-Qiang, et al.
Veröffentlicht: (2026)
von: Cai, Xin-Qiang, et al.
Veröffentlicht: (2026)
Random Masking Finds Winning Tickets for Parameter Efficient Fine-tuning
von: Xu, Jing, et al.
Veröffentlicht: (2024)
von: Xu, Jing, et al.
Veröffentlicht: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
Mano: Restriking Manifold Optimization for LLM Training
von: Gu, Yufei, et al.
Veröffentlicht: (2026)
von: Gu, Yufei, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
von: Ackermann, Johannes, et al.
Veröffentlicht: (2024)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2024)
GradientStabilizer:Fix the Norm, Not the Gradient
von: Huang, Tianjin, et al.
Veröffentlicht: (2025)
von: Huang, Tianjin, et al.
Veröffentlicht: (2025)
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples
von: Gao, Chengqian, et al.
Veröffentlicht: (2025)
von: Gao, Chengqian, et al.
Veröffentlicht: (2025)
Weak-to-Strong Diffusion with Reflection
von: Bai, Lichen, et al.
Veröffentlicht: (2025)
von: Bai, Lichen, et al.
Veröffentlicht: (2025)
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2025)
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2025)
On the Condition Number Dependency in Bilevel Optimization
von: Chen, Lesi, et al.
Veröffentlicht: (2025)
von: Chen, Lesi, et al.
Veröffentlicht: (2025)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
Towards Scalable Oversight via Partitioned Human Supervision
von: Yin, Ren, et al.
Veröffentlicht: (2025)
von: Yin, Ren, et al.
Veröffentlicht: (2025)
Low Rank Gradients and Where to Find Them
von: Sonthalia, Rishi, et al.
Veröffentlicht: (2025)
von: Sonthalia, Rishi, et al.
Veröffentlicht: (2025)
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
von: Chen, Hao, et al.
Veröffentlicht: (2023)
von: Chen, Hao, et al.
Veröffentlicht: (2023)
Reasoning Inconsistencies and How to Mitigate Them in Deep Learning
von: Arakelyan, Erik
Veröffentlicht: (2025)
von: Arakelyan, Erik
Veröffentlicht: (2025)
Fantastic Multi-Task Gradient Updates and How to Find Them In a Cone
von: Hassanpour, Negar, et al.
Veröffentlicht: (2025)
von: Hassanpour, Negar, et al.
Veröffentlicht: (2025)
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
von: Xu, Huangyu, et al.
Veröffentlicht: (2026)
von: Xu, Huangyu, et al.
Veröffentlicht: (2026)
Sharpness-Aware Black-Box Optimization
von: Ye, Feiyang, et al.
Veröffentlicht: (2024)
von: Ye, Feiyang, et al.
Veröffentlicht: (2024)
The Geometry of Multi-Task Grokking: Transverse Instability, Superposition, and Weight Decay Phase Structure
von: Xu, Yongzhong
Veröffentlicht: (2026)
von: Xu, Yongzhong
Veröffentlicht: (2026)
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
von: Fu, Xiaoliang, et al.
Veröffentlicht: (2026)
von: Fu, Xiaoliang, et al.
Veröffentlicht: (2026)
GNN Explanations that do not Explain and How to find Them
von: Azzolin, Steve, et al.
Veröffentlicht: (2026)
von: Azzolin, Steve, et al.
Veröffentlicht: (2026)
How to Square Tensor Networks and Circuits Without Squaring Them
von: Loconte, Lorenzo, et al.
Veröffentlicht: (2025)
von: Loconte, Lorenzo, et al.
Veröffentlicht: (2025)
Calibrated Language Models and How to Find Them with Label Smoothing
von: Huang, Jerry, et al.
Veröffentlicht: (2025)
von: Huang, Jerry, et al.
Veröffentlicht: (2025)
Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2026)
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2026)
Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
von: Zhang, Zhen-Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Zhen-Yu, et al.
Veröffentlicht: (2024)
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
von: Xu, Haoran, et al.
Veröffentlicht: (2025)
von: Xu, Haoran, et al.
Veröffentlicht: (2025)
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
von: Liu, Siyuan, et al.
Veröffentlicht: (2026)
von: Liu, Siyuan, et al.
Veröffentlicht: (2026)
Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation
von: Xu, Jie, et al.
Veröffentlicht: (2025)
von: Xu, Jie, et al.
Veröffentlicht: (2025)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)
Fantastic Copyrighted Beasts and How (Not) to Generate Them
von: He, Luxi, et al.
Veröffentlicht: (2024)
von: He, Luxi, et al.
Veröffentlicht: (2024)
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
von: Wen, Xuexiang, et al.
Veröffentlicht: (2026)
von: Wen, Xuexiang, et al.
Veröffentlicht: (2026)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
von: He, Di, et al.
Veröffentlicht: (2025)
von: He, Di, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding Transformer Optimization via Gradient Heterogeneity
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025) -
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
von: Xu, Kevin, et al.
Veröffentlicht: (2025) -
A Formal Comparison Between Chain of Thought and Latent Thought
von: Xu, Kevin, et al.
Veröffentlicht: (2025) -
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
von: Chen, Lesi, et al.
Veröffentlicht: (2023) -
Fix Initial Codes and Iteratively Refine Textual Directions Toward Safe Multi-Turn Code Correction
von: Tanaka, Yuto, et al.
Veröffentlicht: (2026)