Saved in:
| Main Authors: | Xiao, Jinying, Li, Ping, Nie, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2405.03228 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LNPT: Label-free Network Pruning and Training
by: Xiao, Jinying, et al.
Published: (2024)
by: Xiao, Jinying, et al.
Published: (2024)
SEVEN: Pruning Transformer Model by Reserving Sentinels
by: Xiao, Jinying, et al.
Published: (2024)
by: Xiao, Jinying, et al.
Published: (2024)
EMP: Enhance Memory in Data Pruning
by: Xiao, Jinying, et al.
Published: (2024)
by: Xiao, Jinying, et al.
Published: (2024)
TED: Training-Free Experience Distillation for Multimodal Reasoning
by: Yuan, Shuozhi, et al.
Published: (2026)
by: Yuan, Shuozhi, et al.
Published: (2026)
Incorruptible Neural Networks: Training Models that can Generalize to Large Internal Perturbations
by: Jacobson, Philip, et al.
Published: (2026)
by: Jacobson, Philip, et al.
Published: (2026)
xTED: Cross-Domain Adaptation via Diffusion-Based Trajectory Editing
by: Niu, Haoyi, et al.
Published: (2024)
by: Niu, Haoyi, et al.
Published: (2024)
TED: Related Party Transaction guided Tax Evasion Detection on Heterogeneous Graph
by: Xu, Yiming, et al.
Published: (2026)
by: Xu, Yiming, et al.
Published: (2026)
T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory Stitching
by: Pan, Zizheng, et al.
Published: (2024)
by: Pan, Zizheng, et al.
Published: (2024)
TED: Turn Emphasis with Dialogue Feature Attention for Emotion Recognition in Conversation
by: Ono, Junya, et al.
Published: (2025)
by: Ono, Junya, et al.
Published: (2025)
Outlier Smoothing with Closed-Form Rotations for W4A4 Large Language Model Quantization
by: Xiao, Jinying, et al.
Published: (2025)
by: Xiao, Jinying, et al.
Published: (2025)
Trainable Weight Averaging: Accelerating Training and Improving Generalization
by: Li, Tao, et al.
Published: (2022)
by: Li, Tao, et al.
Published: (2022)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
by: Bai, Yuyang, et al.
Published: (2026)
by: Bai, Yuyang, et al.
Published: (2026)
A Muon-Accelerated Algorithm for Low Separation Rank Tensor Generalized Linear Models
by: Liang, Xiao, et al.
Published: (2026)
by: Liang, Xiao, et al.
Published: (2026)
LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
Accelerating Reinforcement Learning Training Using Simulation Surrogate Models
by: Ghasemloo, Mohammadmahdi, et al.
Published: (2026)
by: Ghasemloo, Mohammadmahdi, et al.
Published: (2026)
Noise-Adaptive Layerwise Learning Rates: Accelerating Geometry-Aware Optimization for Deep Neural Network Training
by: Hao, Jie, et al.
Published: (2025)
by: Hao, Jie, et al.
Published: (2025)
Incorporating Inductive Biases to Energy-based Generative Models
by: Li, Yukun, et al.
Published: (2025)
by: Li, Yukun, et al.
Published: (2025)
Improving Model Fusion by Training-time Neuron Alignment with Fixed Neuron Anchors
by: Li, Zexi, et al.
Published: (2024)
by: Li, Zexi, et al.
Published: (2024)
A Training Data Recipe to Accelerate A* Search with Language Models
by: Gupta, Devaansh, et al.
Published: (2024)
by: Gupta, Devaansh, et al.
Published: (2024)
FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism
by: Wang, Yujie, et al.
Published: (2024)
by: Wang, Yujie, et al.
Published: (2024)
LLA: Enhancing Security and Privacy for Generative Models with Logic-Locked Accelerators
by: Li, You, et al.
Published: (2025)
by: Li, You, et al.
Published: (2025)
Memory Analysis on the Training Course of DeepSeek Models
by: Zhang, Ping, et al.
Published: (2025)
by: Zhang, Ping, et al.
Published: (2025)
Sharp Generalization for Nonparametric Regression in Interpolation Space by Over-Parameterized Neural Networks Trained with Preconditioned Gradient Descent and Early Stopping
by: Yang, Yingzhen, et al.
Published: (2024)
by: Yang, Yingzhen, et al.
Published: (2024)
Boost Post-Training Quantization via Null Space Optimization for Large Language Models
by: Zhao, Jiaqi, et al.
Published: (2025)
by: Zhao, Jiaqi, et al.
Published: (2025)
Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity
by: Zhou, Yucheng, et al.
Published: (2026)
by: Zhou, Yucheng, et al.
Published: (2026)
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse
by: Wang, Jinghui, et al.
Published: (2025)
by: Wang, Jinghui, et al.
Published: (2025)
Diversified Batch Selection for Training Acceleration
by: Hong, Feng, et al.
Published: (2024)
by: Hong, Feng, et al.
Published: (2024)
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
by: Li, Jiacheng, et al.
Published: (2026)
by: Li, Jiacheng, et al.
Published: (2026)
SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training
by: He, Zhongyu, et al.
Published: (2026)
by: He, Zhongyu, et al.
Published: (2026)
AB-Cache: Training-Free Acceleration of Diffusion Models via Adams-Bashforth Cached Feature Reuse
by: Yu, Zichao, et al.
Published: (2025)
by: Yu, Zichao, et al.
Published: (2025)
Predictive Feature Caching for Training-free Acceleration of Molecular Geometry Generation
by: Sommer, Johanna, et al.
Published: (2025)
by: Sommer, Johanna, et al.
Published: (2025)
Accelerating High-Throughput Catalyst Screening by Direct Generation of Equilibrium Adsorption Structures
by: Huo, Songze, et al.
Published: (2025)
by: Huo, Songze, et al.
Published: (2025)
ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
by: Zhang, Guangwei, et al.
Published: (2025)
by: Zhang, Guangwei, et al.
Published: (2025)
Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion Training
by: Kim, Jaeyeon, et al.
Published: (2026)
by: Kim, Jaeyeon, et al.
Published: (2026)
Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation
by: Nie, Dong
Published: (2026)
by: Nie, Dong
Published: (2026)
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
by: Chai, Huichao, et al.
Published: (2026)
by: Chai, Huichao, et al.
Published: (2026)
Accelerating Recommender Model Training by Dynamically Skipping Stale Embeddings
by: Maboud, Yassaman Ebrahimzadeh, et al.
Published: (2024)
by: Maboud, Yassaman Ebrahimzadeh, et al.
Published: (2024)
A Multi-Level Framework for Accelerating Training Transformer Models
by: Zou, Longwei, et al.
Published: (2024)
by: Zou, Longwei, et al.
Published: (2024)
Training-free Heterogeneous Model Merging
by: Xu, Zhengqi, et al.
Published: (2024)
by: Xu, Zhengqi, et al.
Published: (2024)
Internalizing World Models via Self-Play Finetuning for Agentic RL
by: Chen, Shiqi, et al.
Published: (2025)
by: Chen, Shiqi, et al.
Published: (2025)
Similar Items
-
LNPT: Label-free Network Pruning and Training
by: Xiao, Jinying, et al.
Published: (2024) -
SEVEN: Pruning Transformer Model by Reserving Sentinels
by: Xiao, Jinying, et al.
Published: (2024) -
EMP: Enhance Memory in Data Pruning
by: Xiao, Jinying, et al.
Published: (2024) -
TED: Training-Free Experience Distillation for Multimodal Reasoning
by: Yuan, Shuozhi, et al.
Published: (2026) -
Incorruptible Neural Networks: Training Models that can Generalize to Large Internal Perturbations
by: Jacobson, Philip, et al.
Published: (2026)