SCPL: Enhancing Neural Network Training Throughput with Decoupled Local Losses and Model Parallelism
Fuente:
arXiv
Guardado en:
| Autores principales: | Ho, Ming-Yao, Wang, Cheng-Kai, Lin, You-Teng, Chen, Hung-Hsuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
por: Huang, Zih-Hao, et al.
Publicado: (2025)
por: Huang, Zih-Hao, et al.
Publicado: (2025)
Dynamic DropConnect: Enhancing Neural Network Robustness through Adaptive Edge Dropping Strategies
por: Yang, Yuan-Chih, et al.
Publicado: (2025)
por: Yang, Yuan-Chih, et al.
Publicado: (2025)
Understanding Gradient Boosting Classifier: Training, Prediction, and the Role of $γ_j$
por: Chen, Hung-Hsuan
Publicado: (2024)
por: Chen, Hung-Hsuan
Publicado: (2024)
Energy Consumption in Parallel Neural Network Training
por: Huber, Philipp, et al.
Publicado: (2025)
por: Huber, Philipp, et al.
Publicado: (2025)
Z-Error Loss for Training Neural Networks
por: Godin, Guillaume
Publicado: (2025)
por: Godin, Guillaume
Publicado: (2025)
Decoupled Split Learning via Auxiliary Loss
por: Zihad, Anower, et al.
Publicado: (2026)
por: Zihad, Anower, et al.
Publicado: (2026)
Decoupling Search and Learning in Neural Net Training
por: Vegesna, Akshay, et al.
Publicado: (2025)
por: Vegesna, Akshay, et al.
Publicado: (2025)
How Fast Should a Model Commit to Supervision? Training Reasoning Models on the Tsallis Loss Continuum
por: Lin, Chu-Cheng, et al.
Publicado: (2026)
por: Lin, Chu-Cheng, et al.
Publicado: (2026)
Multivariate Beta Mixture Model: Probabilistic Clustering With Flexible Cluster Shapes
por: Hsu, Yung-Peng, et al.
Publicado: (2024)
por: Hsu, Yung-Peng, et al.
Publicado: (2024)
Decouple Graph Neural Networks: Train Multiple Simple GNNs Simultaneously Instead of One
por: Zhang, Hongyuan, et al.
Publicado: (2023)
por: Zhang, Hongyuan, et al.
Publicado: (2023)
Flexible Bivariate Beta Mixture Model: A Probabilistic Approach for Clustering Complex Data Structures
por: Hsu, Yung-Peng, et al.
Publicado: (2025)
por: Hsu, Yung-Peng, et al.
Publicado: (2025)
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
por: Chen, Hung-Hsuan
Publicado: (2026)
por: Chen, Hung-Hsuan
Publicado: (2026)
CeRA: Overcoming the Linear Ceiling of Low-Rank Adaptation via Capacity Expansion
por: Chen, Hung-Hsuan
Publicado: (2026)
por: Chen, Hung-Hsuan
Publicado: (2026)
ROOT: Robust Orthogonalized Optimizer for Neural Network Training
por: He, Wei, et al.
Publicado: (2025)
por: He, Wei, et al.
Publicado: (2025)
Interpretable by AI Mother Tongue: Native Symbolic Reasoning in Neural Models
por: Liu, Hung Ming
Publicado: (2025)
por: Liu, Hung Ming
Publicado: (2025)
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
por: Shin, Baekrok, et al.
Publicado: (2024)
por: Shin, Baekrok, et al.
Publicado: (2024)
Approximated Likelihood Ratio: A Forward-Only and Parallel Framework for Boosting Neural Network Training
por: Zhang, Zeliang, et al.
Publicado: (2024)
por: Zhang, Zeliang, et al.
Publicado: (2024)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
por: Li, Gang, et al.
Publicado: (2025)
por: Li, Gang, et al.
Publicado: (2025)
Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks
por: Hu, Rui, et al.
Publicado: (2024)
por: Hu, Rui, et al.
Publicado: (2024)
Local Control Networks (LCNs): Optimizing Flexibility in Neural Network Data Pattern Capture
por: Nguyen, Hy, et al.
Publicado: (2025)
por: Nguyen, Hy, et al.
Publicado: (2025)
Lightweight Geometric Adaptation for Training Physics-Informed Neural Networks
por: An, Kang, et al.
Publicado: (2026)
por: An, Kang, et al.
Publicado: (2026)
MixGCN: Scalable GCN Training by Mixture of Parallelism and Mixture of Accelerators
por: Wan, Cheng, et al.
Publicado: (2025)
por: Wan, Cheng, et al.
Publicado: (2025)
Taming Latent Diffusion Model for Neural Radiance Field Inpainting
por: Lin, Chieh Hubert, et al.
Publicado: (2024)
por: Lin, Chieh Hubert, et al.
Publicado: (2024)
Regime Change Hypothesis: Foundations for Decoupled Dynamics in Neural Network Training
por: Pérez-Corral, Cristian, et al.
Publicado: (2026)
por: Pérez-Corral, Cristian, et al.
Publicado: (2026)
Neural Network Plasticity and Loss Sharpness
por: Koster, Max, et al.
Publicado: (2024)
por: Koster, Max, et al.
Publicado: (2024)
Adversarial Training for Robust Coverage Network under Worst-case Facility Losses
por: Miao, Changhao, et al.
Publicado: (2026)
por: Miao, Changhao, et al.
Publicado: (2026)
TEESlice: Protecting Sensitive Neural Network Models in Trusted Execution Environments When Attackers have Pre-Trained Models
por: Li, Ding, et al.
Publicado: (2024)
por: Li, Ding, et al.
Publicado: (2024)
SP2RINT: Spatially-Decoupled Physics-Inspired Progressive Inverse Optimization for Scalable, PDE-Constrained Meta-Optical Neural Network Training
por: Ma, Pingchuan, et al.
Publicado: (2025)
por: Ma, Pingchuan, et al.
Publicado: (2025)
Parallelizing Node-Level Explainability in Graph Neural Networks
por: Llorente, Oscar, et al.
Publicado: (2026)
por: Llorente, Oscar, et al.
Publicado: (2026)
Enhancing Trustworthiness of Graph Neural Networks with Rank-Based Conformal Training
por: Wang, Ting, et al.
Publicado: (2025)
por: Wang, Ting, et al.
Publicado: (2025)
A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing
por: Seal, Sudip K., et al.
Publicado: (2025)
por: Seal, Sudip K., et al.
Publicado: (2025)
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
por: Huh, Minyoung, et al.
Publicado: (2024)
por: Huh, Minyoung, et al.
Publicado: (2024)
MemLoss: Enhancing Adversarial Training with Recycling Adversarial Examples
por: Mahdi, Soroush, et al.
Publicado: (2025)
por: Mahdi, Soroush, et al.
Publicado: (2025)
DeToNATION: Decoupled Torch Network-Aware Training on Interlinked Online Nodes
por: From, Mogens Henrik, et al.
Publicado: (2025)
por: From, Mogens Henrik, et al.
Publicado: (2025)
MoNTA: Accelerating Mixture-of-Experts Training with Network-Traffc-Aware Parallel Optimization
por: Guo, Jingming, et al.
Publicado: (2024)
por: Guo, Jingming, et al.
Publicado: (2024)
Heterophily-Agnostic Hypergraph Neural Networks with Riemannian Local Exchanger
por: Sun, Li, et al.
Publicado: (2026)
por: Sun, Li, et al.
Publicado: (2026)
Addressing Long-Tail Noisy Label Learning Problems: a Two-Stage Solution with Label Refurbishment Considering Label Rarity
por: Wu, Ying-Hsuan, et al.
Publicado: (2024)
por: Wu, Ying-Hsuan, et al.
Publicado: (2024)
Dynamic Universal Approximation Theory: Foundations for Parallelism in Neural Networks
por: Wang, Wei, et al.
Publicado: (2024)
por: Wang, Wei, et al.
Publicado: (2024)
Enhancing the Performance of Neural Networks Through Causal Discovery and Integration of Domain Knowledge
por: Zhang, Xiaoge, et al.
Publicado: (2023)
por: Zhang, Xiaoge, et al.
Publicado: (2023)
Locally Convex Global Loss Network for Decision-Focused Learning
por: Jeon, Haeun, et al.
Publicado: (2024)
por: Jeon, Haeun, et al.
Publicado: (2024)
Ejemplares similares
-
DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
por: Huang, Zih-Hao, et al.
Publicado: (2025) -
Dynamic DropConnect: Enhancing Neural Network Robustness through Adaptive Edge Dropping Strategies
por: Yang, Yuan-Chih, et al.
Publicado: (2025) -
Understanding Gradient Boosting Classifier: Training, Prediction, and the Role of $γ_j$
por: Chen, Hung-Hsuan
Publicado: (2024) -
Energy Consumption in Parallel Neural Network Training
por: Huber, Philipp, et al.
Publicado: (2025) -
Z-Error Loss for Training Neural Networks
por: Godin, Guillaume
Publicado: (2025)