Saved in:
| Main Authors: | Zhai, Zhiyuan, Li, Bingcong, Xiao, Bingnan, Li, Ming, Wang, Xin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.14853 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FedVSSAM: Mitigating Flatness Incompatibility in Sharpness-Aware Federated Learning
by: Xiao, Bingnan, et al.
Published: (2026)
by: Xiao, Bingnan, et al.
Published: (2026)
Scalable Variational Bayesian Fine-Tuning of LLMs via Orthogonalized Low-Rank Adapters
by: Xiang, Haotian, et al.
Published: (2026)
by: Xiang, Haotian, et al.
Published: (2026)
Revisable by Design: A Theory of Streaming LLM Agent Execution
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
FLARE: A New Federated Learning Framework with Adjustable Learning Rates over Resource-Constrained Wireless Networks
by: Xiao, Bingnan, et al.
Published: (2024)
by: Xiao, Bingnan, et al.
Published: (2024)
Autoregressive Policy Optimization for Constrained Allocation Tasks
by: Winkel, David, et al.
Published: (2024)
by: Winkel, David, et al.
Published: (2024)
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
by: Alomrani, Mohammad Ali, et al.
Published: (2025)
by: Alomrani, Mohammad Ali, et al.
Published: (2025)
When to Ponder: Adaptive Compute Allocation for Code Generation via Test-Time Training
by: Sim, Gihyeon
Published: (2025)
by: Sim, Gihyeon
Published: (2025)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
by: Huang, Jiawei, et al.
Published: (2025)
by: Huang, Jiawei, et al.
Published: (2025)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
by: Li, Gang, et al.
Published: (2025)
by: Li, Gang, et al.
Published: (2025)
Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
ANCRe: Adaptive Neural Connection Reassignment for Efficient Depth Scaling
by: Zhang, Yilang, et al.
Published: (2026)
by: Zhang, Yilang, et al.
Published: (2026)
Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning
by: Wang, Peihao, et al.
Published: (2026)
by: Wang, Peihao, et al.
Published: (2026)
Muown: Row-Norm Control for Muon Optimization
by: Lion, Kai, et al.
Published: (2026)
by: Lion, Kai, et al.
Published: (2026)
Implicit Regularization of Sharpness-Aware Minimization for Scale-Invariant Problems
by: Li, Bingcong, et al.
Published: (2024)
by: Li, Bingcong, et al.
Published: (2024)
State Tuning: State-based Test-Time Scaling on RWKV-7
by: Xiao, Liu, et al.
Published: (2025)
by: Xiao, Liu, et al.
Published: (2025)
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
Targeted Tests for LLM Reasoning: An Audit-Constrained Protocol
by: Li, Hongmin
Published: (2026)
by: Li, Hongmin
Published: (2026)
Calibration-Aware Policy Optimization for Reasoning LLMs
by: Wang, Ziqi, et al.
Published: (2026)
by: Wang, Ziqi, et al.
Published: (2026)
Intelligent Resource Allocation Optimization for Cloud Computing via Machine Learning
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy Optimization
by: Zhu, Yuchen, et al.
Published: (2025)
by: Zhu, Yuchen, et al.
Published: (2025)
Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
by: Niu, Yifan, et al.
Published: (2025)
by: Niu, Yifan, et al.
Published: (2025)
Secure Resource Allocation via Constrained Deep Reinforcement Learning
by: Sun, Jianfei, et al.
Published: (2025)
by: Sun, Jianfei, et al.
Published: (2025)
ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
by: Behric, Lejs Deen, et al.
Published: (2025)
by: Behric, Lejs Deen, et al.
Published: (2025)
Meta-Learning with Versatile Loss Geometries for Fast Adaptation Using Mirror Descent
by: Zhang, Yilang, et al.
Published: (2023)
by: Zhang, Yilang, et al.
Published: (2023)
Learnable Loss Geometries with Mirror Descent for Scalable and Convergent Meta-Learning
by: Zhang, Yilang, et al.
Published: (2025)
by: Zhang, Yilang, et al.
Published: (2025)
VASSO: Variance Suppression for Sharpness-Aware Minimization
by: Li, Bingcong, et al.
Published: (2025)
by: Li, Bingcong, et al.
Published: (2025)
Preconditioned Sharpness-Aware Minimization: Unifying Analysis and a Novel Learning Algorithm
by: Zhang, Yilang, et al.
Published: (2025)
by: Zhang, Yilang, et al.
Published: (2025)
RefLoRA: Refactored Low-Rank Adaptation for Efficient Fine-Tuning of Large Models
by: Zhang, Yilang, et al.
Published: (2025)
by: Zhang, Yilang, et al.
Published: (2025)
$\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space
by: Wang, Peihao, et al.
Published: (2026)
by: Wang, Peihao, et al.
Published: (2026)
Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization
by: Hu, Kai, et al.
Published: (2024)
by: Hu, Kai, et al.
Published: (2024)
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
by: Wang, Junxiong, et al.
Published: (2025)
by: Wang, Junxiong, et al.
Published: (2025)
DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization
by: Li, Gang, et al.
Published: (2025)
by: Li, Gang, et al.
Published: (2025)
ATLAS: Adaptive Test-Time Latent Steering with External Verifiers for Enhancing LLMs Reasoning
by: Nguyen, Tuc, et al.
Published: (2026)
by: Nguyen, Tuc, et al.
Published: (2026)
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
by: Li, Hengli, et al.
Published: (2025)
by: Li, Hengli, et al.
Published: (2025)
GRPOformer: Advancing Hyperparameter Optimization via Group Relative Policy Optimization
by: Guo, Haoxin, et al.
Published: (2025)
by: Guo, Haoxin, et al.
Published: (2025)
Low-Rank Adaptation Redux for Large Models
by: Li, Bingcong, et al.
Published: (2026)
by: Li, Bingcong, et al.
Published: (2026)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
AGPO: Adaptive Group Policy Optimization with Dual Statistical Feedback
by: Hu, Miaobo, et al.
Published: (2026)
by: Hu, Miaobo, et al.
Published: (2026)
Similar Items
-
FedVSSAM: Mitigating Flatness Incompatibility in Sharpness-Aware Federated Learning
by: Xiao, Bingnan, et al.
Published: (2026) -
Scalable Variational Bayesian Fine-Tuning of LLMs via Orthogonalized Low-Rank Adapters
by: Xiang, Haotian, et al.
Published: (2026) -
Revisable by Design: A Theory of Streaming LLM Agent Execution
by: Zhai, Zhiyuan, et al.
Published: (2026) -
FLARE: A New Federated Learning Framework with Adjustable Learning Rates over Resource-Constrained Wireless Networks
by: Xiao, Bingnan, et al.
Published: (2024) -
Autoregressive Policy Optimization for Constrained Allocation Tasks
by: Winkel, David, et al.
Published: (2024)