Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized Models
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Yubin, Chen, Yixuan, Dong, Mingzhi, Yang, Xiaochen, Li, Dongsheng, Wang, Yujiang, Dick, Robert P., Lv, Qin, Zhao, Yingying, Yang, Fan, Lu, Tun, Gu, Ning, Shang, Li |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Denoising Reuse: Exploiting Inter-frame Motion Consistency for Efficient Video Latent Generation
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
Metis: Training LLMs with FP4 Quantization
by: Cao, Hengjie, et al.
Published: (2025)
by: Cao, Hengjie, et al.
Published: (2025)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
Faster Convergence of Local SGD for Over-Parameterized Models
by: Qin, Tiancheng, et al.
Published: (2022)
by: Qin, Tiancheng, et al.
Published: (2022)
Unbiased Collaborative Filtering with Fair Sampling
by: Liu, Jiahao, et al.
Published: (2025)
by: Liu, Jiahao, et al.
Published: (2025)
Frequency-aware Graph Signal Processing for Collaborative Filtering
by: Xia, Jiafeng, et al.
Published: (2024)
by: Xia, Jiafeng, et al.
Published: (2024)
A Comprehensive Summarization and Evaluation of Feature Refinement Modules for CTR Prediction
by: Wang, Fangye, et al.
Published: (2023)
by: Wang, Fangye, et al.
Published: (2023)
Oracle-guided Dynamic User Preference Modeling for Sequential Recommendation
by: Xia, Jiafeng, et al.
Published: (2024)
by: Xia, Jiafeng, et al.
Published: (2024)
FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning
by: Pan, Haihui, et al.
Published: (2026)
by: Pan, Haihui, et al.
Published: (2026)
GraphTransfer: A Generic Feature Fusion Framework for Collaborative Filtering
by: Xia, Jiafeng, et al.
Published: (2024)
by: Xia, Jiafeng, et al.
Published: (2024)
Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better
by: Zhao, Ji, et al.
Published: (2026)
by: Zhao, Ji, et al.
Published: (2026)
Sharp Generalization for Nonparametric Regression in Interpolation Space by Over-Parameterized Neural Networks Trained with Preconditioned Gradient Descent and Early Stopping
by: Yang, Yingzhen, et al.
Published: (2024)
by: Yang, Yingzhen, et al.
Published: (2024)
Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy
by: Huang, Zhendong, et al.
Published: (2026)
by: Huang, Zhendong, et al.
Published: (2026)
Improving Controllable Generation: Faster Training and Better Performance via $x_0$-Supervision
by: Sangare, Amadou S., et al.
Published: (2026)
by: Sangare, Amadou S., et al.
Published: (2026)
Knoop: Practical Enhancement of Knockoff with Over-Parameterization for Variable Selection
by: Zhang, Xiaochen, et al.
Published: (2025)
by: Zhang, Xiaochen, et al.
Published: (2025)
Faster and Better 3D Splatting via Group Training
by: Wang, Chengbo, et al.
Published: (2024)
by: Wang, Chengbo, et al.
Published: (2024)
AOTree: Aspect Order Tree-based Model for Explainable Recommendation
by: Zhao, Wenxin, et al.
Published: (2024)
by: Zhao, Wenxin, et al.
Published: (2024)
Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers
by: Chen, Anrui, et al.
Published: (2026)
by: Chen, Anrui, et al.
Published: (2026)
Improving Adaptivity via Over-Parameterization in Sequence Models
by: Li, Yicheng, et al.
Published: (2024)
by: Li, Yicheng, et al.
Published: (2024)
Improving LLM-powered Recommendations with Personalized Information
by: Liu, Jiahao, et al.
Published: (2025)
by: Liu, Jiahao, et al.
Published: (2025)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
by: Huang, Ruijun, et al.
Published: (2026)
by: Huang, Ruijun, et al.
Published: (2026)
LLM-Based User Simulation for Low-Knowledge Shilling Attacks on Recommender Systems
by: Gu, Shengkang, et al.
Published: (2025)
by: Gu, Shengkang, et al.
Published: (2025)
AgentCF++: Memory-enhanced LLM-based Agents for Popularity-aware Cross-domain Recommendations
by: Liu, Jiahao, et al.
Published: (2025)
by: Liu, Jiahao, et al.
Published: (2025)
Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models
by: Sadashivaiah, Vijay, et al.
Published: (2026)
by: Sadashivaiah, Vijay, et al.
Published: (2026)
Feature-Indexed Federated Recommendation with Residual-Quantized Codebooks
by: Han, Mingzhe, et al.
Published: (2026)
by: Han, Mingzhe, et al.
Published: (2026)
Bidirectional Knowledge Distillation for Enhancing Sequential Recommendation with Large Language Models
by: Wu, Jiongran, et al.
Published: (2025)
by: Wu, Jiongran, et al.
Published: (2025)
Tracking Meets LoRA: Faster Training, Larger Model, Stronger Performance
by: Lin, Liting, et al.
Published: (2024)
by: Lin, Liting, et al.
Published: (2024)
FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
by: Lv, Zhengyao, et al.
Published: (2024)
by: Lv, Zhengyao, et al.
Published: (2024)
Dispelling the Curse of Singularities in Neural Network Optimizations
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
Gradient Descent Finds Over-Parameterized Neural Networks with Sharp Generalization for Nonparametric Regression
by: Yang, Yingzhen, et al.
Published: (2024)
by: Yang, Yingzhen, et al.
Published: (2024)
Pointer Networks Trained Better via Evolutionary Algorithms
by: Zhong, Muyao, et al.
Published: (2023)
by: Zhong, Muyao, et al.
Published: (2023)
Faster Parameterized Vertex Multicut
by: Chu, Huairui, et al.
Published: (2026)
by: Chu, Huairui, et al.
Published: (2026)
Better World Models Can Lead to Better Post-Training Performance
by: Gupta, Prakhar, et al.
Published: (2025)
by: Gupta, Prakhar, et al.
Published: (2025)
Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics
by: Wang, Hai, et al.
Published: (2026)
by: Wang, Hai, et al.
Published: (2026)
Drift-Aware Continual Tokenization for Generative Recommendation
by: Feng, Yuebo, et al.
Published: (2026)
by: Feng, Yuebo, et al.
Published: (2026)
FedCIA: Federated Collaborative Information Aggregation for Privacy-Preserving Recommendation
by: Han, Mingzhe, et al.
Published: (2025)
by: Han, Mingzhe, et al.
Published: (2025)
Transparent and Controllable Recommendation Filtering via Multimodal Multi-Agent Collaboration
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)
by: Liu, Xin, et al.
Published: (2022)
Training-Driven Representational Geometry Modularization Predicts Brain Alignment in Language Models
by: Liu, Yixuan, et al.
Published: (2026)
by: Liu, Yixuan, et al.
Published: (2026)
Training Language Model to Critique for Better Refinement
by: Yu, Tianshu, et al.
Published: (2025)
by: Yu, Tianshu, et al.
Published: (2025)
Similar Items
-
Denoising Reuse: Exploiting Inter-frame Motion Consistency for Efficient Video Latent Generation
by: Wang, Chenyu, et al.
Published: (2024) -
Metis: Training LLMs with FP4 Quantization
by: Cao, Hengjie, et al.
Published: (2025) -
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026) -
Faster Convergence of Local SGD for Over-Parameterized Models
by: Qin, Tiancheng, et al.
Published: (2022) -
Unbiased Collaborative Filtering with Fair Sampling
by: Liu, Jiahao, et al.
Published: (2025)