Gespeichert in:
| Hauptverfasser: | Jiang, Yuxuan, Zhou, Ziming, Xu, Boyu, Liu, Beijie, Xu, Runhui, Huang, Peng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2506.14813 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
A Triple-Inertial Accelerated Alternating Optimization Method for Deep Learning Training
von: Yan, Chengcheng, et al.
Veröffentlicht: (2025)
von: Yan, Chengcheng, et al.
Veröffentlicht: (2025)
Training Proactive and Personalized LLM Agents
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
von: Luo, Haohao, et al.
Veröffentlicht: (2026)
von: Luo, Haohao, et al.
Veröffentlicht: (2026)
DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
Learning from Complexity: Exploring Dynamic Sample Pruning of Spatio-Temporal Training
von: Chen, Wei, et al.
Veröffentlicht: (2026)
von: Chen, Wei, et al.
Veröffentlicht: (2026)
Silent Neuron Theory and Plasticity Preservation for Deep Reinforcement Learning in Adaptive Video Streaming
von: He, Zhiqiang, et al.
Veröffentlicht: (2025)
von: He, Zhiqiang, et al.
Veröffentlicht: (2025)
Approximated Likelihood Ratio: A Forward-Only and Parallel Framework for Boosting Neural Network Training
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training
von: Gong, Xue, et al.
Veröffentlicht: (2026)
von: Gong, Xue, et al.
Veröffentlicht: (2026)
Efficient Deep Learning Board: Training Feedback Is Not All You Need
von: Gong, Lina, et al.
Veröffentlicht: (2024)
von: Gong, Lina, et al.
Veröffentlicht: (2024)
Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple Actions
von: Gan, Guo, et al.
Veröffentlicht: (2026)
von: Gan, Guo, et al.
Veröffentlicht: (2026)
Reinforcement Learning on Pre-Training Data
von: Li, Siheng, et al.
Veröffentlicht: (2025)
von: Li, Siheng, et al.
Veröffentlicht: (2025)
Neural Thermodynamic Laws for Large Language Model Training
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
CAdam: Confidence-Based Optimization for Online Learning
von: Wang, Shaowen, et al.
Veröffentlicht: (2024)
von: Wang, Shaowen, et al.
Veröffentlicht: (2024)
Proactive Gradient Conflict Mitigation in Multi-Task Learning: A Sparse Training Perspective
von: Zhang, Zhi, et al.
Veröffentlicht: (2024)
von: Zhang, Zhi, et al.
Veröffentlicht: (2024)
Low-redundancy Distillation for Continual Learning
von: Liu, RuiQi, et al.
Veröffentlicht: (2023)
von: Liu, RuiQi, et al.
Veröffentlicht: (2023)
Tools Fail: Detecting Silent Errors in Faulty Tools
von: Sun, Jimin, et al.
Veröffentlicht: (2024)
von: Sun, Jimin, et al.
Veröffentlicht: (2024)
Z-Error Loss for Training Neural Networks
von: Godin, Guillaume
Veröffentlicht: (2025)
von: Godin, Guillaume
Veröffentlicht: (2025)
A Sensitivity-Driven Expert Allocation Method in LoRA-MoE for Efficient Fine-Tuning
von: Xu, Junzhou, et al.
Veröffentlicht: (2025)
von: Xu, Junzhou, et al.
Veröffentlicht: (2025)
Native Fortran Implementation of TensorFlow-Trained Deep and Bayesian Neural Networks
von: Furlong, Aidan, et al.
Veröffentlicht: (2025)
von: Furlong, Aidan, et al.
Veröffentlicht: (2025)
Towards Faster Training of Diffusion Models: An Inspiration of A Consistency Phenomenon
von: Xu, Tianshuo, et al.
Veröffentlicht: (2024)
von: Xu, Tianshuo, et al.
Veröffentlicht: (2024)
Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks
von: Hu, Rui, et al.
Veröffentlicht: (2024)
von: Hu, Rui, et al.
Veröffentlicht: (2024)
On the Interplay Between Sparsity and Training in Deep Reinforcement Learning
von: Davelouis, Fatima, et al.
Veröffentlicht: (2025)
von: Davelouis, Fatima, et al.
Veröffentlicht: (2025)
Efficient Multi-Task Modeling through Automated Fusion of Trained Models
von: Zhou, Jingxuan, et al.
Veröffentlicht: (2025)
von: Zhou, Jingxuan, et al.
Veröffentlicht: (2025)
To Train or Not to Train: Balancing Efficiency and Training Cost in Deep Reinforcement Learning for Mobile Edge Computing
von: Boscaro, Maddalena, et al.
Veröffentlicht: (2024)
von: Boscaro, Maddalena, et al.
Veröffentlicht: (2024)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
Preparing Lessons for Progressive Training on Language Models
von: Pan, Yu, et al.
Veröffentlicht: (2024)
von: Pan, Yu, et al.
Veröffentlicht: (2024)
Symmetry-Aware Transformer Training for Automated Planning
von: Fritzsche, Markus, et al.
Veröffentlicht: (2025)
von: Fritzsche, Markus, et al.
Veröffentlicht: (2025)
ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning
von: Zhang, Junru, et al.
Veröffentlicht: (2025)
von: Zhang, Junru, et al.
Veröffentlicht: (2025)
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
von: Hou, Hongru, et al.
Veröffentlicht: (2026)
von: Hou, Hongru, et al.
Veröffentlicht: (2026)
Open-World Test-Time Training: Self-Training with Contrast Learning
von: Su, Houcheng, et al.
Veröffentlicht: (2024)
von: Su, Houcheng, et al.
Veröffentlicht: (2024)
POINT$^{2}$: A Polymer Informatics Training and Testing Database
von: Xu, Jiaxin, et al.
Veröffentlicht: (2025)
von: Xu, Jiaxin, et al.
Veröffentlicht: (2025)
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
von: Fluri, Lukas, et al.
Veröffentlicht: (2024)
von: Fluri, Lukas, et al.
Veröffentlicht: (2024)
Enhancing Deep Learning with Optimized Gradient Descent: Bridging Numerical Methods and Neural Network Training
von: Ma, Yuhan, et al.
Veröffentlicht: (2024)
von: Ma, Yuhan, et al.
Veröffentlicht: (2024)
Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
von: Chhabra, Anshuman, et al.
Veröffentlicht: (2024)
von: Chhabra, Anshuman, et al.
Veröffentlicht: (2024)
Online Training and Pruning of Deep Reinforcement Learning Networks
von: Guenter, Valentin Frank Ingmar, et al.
Veröffentlicht: (2025)
von: Guenter, Valentin Frank Ingmar, et al.
Veröffentlicht: (2025)
Exploring Dynamic Properties of Backdoor Training Through Information Bottleneck
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
Sparse Training for Federated Learning with Regularized Error Correction
von: Greidi, Ran, et al.
Veröffentlicht: (2023)
von: Greidi, Ran, et al.
Veröffentlicht: (2023)
Training Large Language Models to Reason via EM Policy Gradient
von: Xu, Tianbing
Veröffentlicht: (2025)
von: Xu, Tianbing
Veröffentlicht: (2025)
Ähnliche Einträge
-
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
von: Xu, Yuanda, et al.
Veröffentlicht: (2026) -
A Triple-Inertial Accelerated Alternating Optimization Method for Deep Learning Training
von: Yan, Chengcheng, et al.
Veröffentlicht: (2025) -
Training Proactive and Personalized LLM Agents
von: Sun, Weiwei, et al.
Veröffentlicht: (2025) -
IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
von: Luo, Haohao, et al.
Veröffentlicht: (2026) -
DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)