Saved in:
| Main Authors: | Cai, Wenrui, Zhu, Defa, Liu, Qingjie, Min, Qiyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.22777 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ultra-Sparse Memory Network
by: Huang, Zihao, et al.
Published: (2024)
by: Huang, Zihao, et al.
Published: (2024)
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
by: Huang, Hongzhi, et al.
Published: (2025)
by: Huang, Hongzhi, et al.
Published: (2025)
Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts
by: Yuan, Yike, et al.
Published: (2025)
by: Yuan, Yike, et al.
Published: (2025)
Frac-Connections: Fractional Extension of Hyper-Connections
by: Zhu, Defa, et al.
Published: (2025)
by: Zhu, Defa, et al.
Published: (2025)
Hyper-Connections
by: Zhu, Defa, et al.
Published: (2024)
by: Zhu, Defa, et al.
Published: (2024)
UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning
by: Huang, Zihao, et al.
Published: (2025)
by: Huang, Zihao, et al.
Published: (2025)
High-Layer Attention Pruning with Rescaling
by: Liu, Songtao, et al.
Published: (2025)
by: Liu, Songtao, et al.
Published: (2025)
Achieving Margin Maximization Exponentially Fast via Progressive Norm Rescaling
by: Wang, Mingze, et al.
Published: (2023)
by: Wang, Mingze, et al.
Published: (2023)
An Iterative Algorithm for Rescaled Hyperbolic Functions Regression
by: Gao, Yeqi, et al.
Published: (2023)
by: Gao, Yeqi, et al.
Published: (2023)
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025)
by: Liao, Qichen, et al.
Published: (2025)
MimicNorm: Weight Mean and Last BN Layer Mimic the Dynamic of Batch Normalization
by: Fei, Wen, et al.
Published: (2020)
by: Fei, Wen, et al.
Published: (2020)
Thompson sampling: Precise arm-pull dynamics and adaptive inference
by: Han, Qiyang
Published: (2026)
by: Han, Qiyang
Published: (2026)
Temporal Score Rescaling for Temperature Sampling in Diffusion and Flow Models
by: Xu, Yanbo, et al.
Published: (2025)
by: Xu, Yanbo, et al.
Published: (2025)
Support Vector Machine Classifier with Rescaled Huberized Pinball Loss
by: Diao, Shibo
Published: (2025)
by: Diao, Shibo
Published: (2025)
Non-Vacuous Generalization Bounds: Can Rescaling Invariances Help?
by: Rouchouse, Damien, et al.
Published: (2025)
by: Rouchouse, Damien, et al.
Published: (2025)
Rescaled Influence Functions: Accurate Data Attribution in High Dimension
by: Rubinstein, Ittai, et al.
Published: (2025)
by: Rubinstein, Ittai, et al.
Published: (2025)
Recovering Plasticity of Neural Networks via Soft Weight Rescaling
by: Oh, Seungwon, et al.
Published: (2025)
by: Oh, Seungwon, et al.
Published: (2025)
Long-time dynamics and universality of nonconvex gradient descent
by: Han, Qiyang
Published: (2025)
by: Han, Qiyang
Published: (2025)
ConceptMoE: Adaptive Token-to-Concept Compression for Implicit Compute Allocation
by: Huang, Zihao, et al.
Published: (2026)
by: Huang, Zihao, et al.
Published: (2026)
STAR: Spectral Truncation and Rescale for Model Merging
by: Lee, Yu-Ang, et al.
Published: (2025)
by: Lee, Yu-Ang, et al.
Published: (2025)
Learning Patient-Specific Spatial Biomarker Dynamics via Operator Learning for Alzheimer's Disease Progression
by: Wang, Jindong, et al.
Published: (2025)
by: Wang, Jindong, et al.
Published: (2025)
Variance Control via Weight Rescaling in LLM Pre-training
by: Owen, Louis, et al.
Published: (2025)
by: Owen, Louis, et al.
Published: (2025)
Q-learning with Adjoint Matching
by: Li, Qiyang, et al.
Published: (2026)
by: Li, Qiyang, et al.
Published: (2026)
SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
HIPTrack: Visual Tracking with Historical Prompts
by: Cai, Wenrui, et al.
Published: (2023)
by: Cai, Wenrui, et al.
Published: (2023)
Learn Singularly Perturbed Solutions via Homotopy Dynamics
by: Chen, Chuqi, et al.
Published: (2025)
by: Chen, Chuqi, et al.
Published: (2025)
$α$-LoRA: Effective Fine-Tuning via Base Model Rescaling
by: Firdoussi, Aymane El, et al.
Published: (2025)
by: Firdoussi, Aymane El, et al.
Published: (2025)
Flow Q-Learning
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
A Rescaling-Invariant Lipschitz Bound Based on Path-Metrics for Modern ReLU Network Parameterizations
by: Gonon, Antoine, et al.
Published: (2024)
by: Gonon, Antoine, et al.
Published: (2024)
Stability Preserving Data-driven Models With Latent Dynamics
by: Luo, Yushuang, et al.
Published: (2022)
by: Luo, Yushuang, et al.
Published: (2022)
Newton Informed Neural Operator for Computing Multiple Solutions of Nonlinear Partials Differential Equations
by: Hao, Wenrui, et al.
Published: (2024)
by: Hao, Wenrui, et al.
Published: (2024)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
by: Mueller, Lion, et al.
Published: (2025)
by: Mueller, Lion, et al.
Published: (2025)
Task Expansion and Cross Refinement for Open-World Conditional Modeling
by: Brahmavar, Shreyas Bhat, et al.
Published: (2026)
by: Brahmavar, Shreyas Bhat, et al.
Published: (2026)
Reinforcement Learning with Action Chunking
by: Li, Qiyang, et al.
Published: (2025)
by: Li, Qiyang, et al.
Published: (2025)
Decoupled Q-Chunking
by: Li, Qiyang, et al.
Published: (2025)
by: Li, Qiyang, et al.
Published: (2025)
Gradient descent inference in empirical risk minimization
by: Han, Qiyang, et al.
Published: (2024)
by: Han, Qiyang, et al.
Published: (2024)
Normalizing Flows on Quotient Manifolds via Boundary Quotients
by: Ghanem, William, et al.
Published: (2025)
by: Ghanem, William, et al.
Published: (2025)
Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling for Reinforcement Learning
by: Li, Zichao, et al.
Published: (2026)
by: Li, Zichao, et al.
Published: (2026)
Kernelized Normalizing Constant Estimation: Bridging Bayesian Quadrature and Bayesian Optimization
by: Cai, Xu, et al.
Published: (2024)
by: Cai, Xu, et al.
Published: (2024)
Similar Items
-
Ultra-Sparse Memory Network
by: Huang, Zihao, et al.
Published: (2024) -
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
by: Huang, Hongzhi, et al.
Published: (2025) -
Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts
by: Yuan, Yike, et al.
Published: (2025) -
Frac-Connections: Fractional Extension of Hyper-Connections
by: Zhu, Defa, et al.
Published: (2025) -
Hyper-Connections
by: Zhu, Defa, et al.
Published: (2024)