AXLearn: Modular, Hardware-Agnostic Large Model Training
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Mark, Lan, Chang, Gunter, Tom, Peebles, John, Zhou, Hanzhi, Zou, Kelvin, Bangalore, Sneha, Chiu, Chung-Cheng, Du, Nan, Du, Xianzhi, Dufter, Philipp, Hou, Ruixuan, Huang, Haoshuo, Hwang, Dongseong, Kong, Xiang, Lei, Jinhao, Lei, Tao, Li, Meng, Li, Li, Lu, Jiarui, Lu, Zhiyun, Ma, Yiping, Qiu, David, Rathod, Vivek, Tong, Senyu, Tu, Zhucheng, Wang, Jianyu, Wang, Yongqiang, Wang, Zirui, Weers, Floris, Wiseman, Sam, Yin, Guoli, Zhang, Bowen, Zhou, Xiyou, Zhuo, Danyang, Leong, Cheng, Pang, Ruoming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
por: Wang, Chong, et al.
Publicado: (2026)
por: Wang, Chong, et al.
Publicado: (2026)
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
por: Findeis, Arduin, et al.
Publicado: (2025)
por: Findeis, Arduin, et al.
Publicado: (2025)
EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing
por: Sun, Haotian, et al.
Publicado: (2024)
por: Sun, Haotian, et al.
Publicado: (2024)
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
por: Arora, Siddhant, et al.
Publicado: (2025)
por: Arora, Siddhant, et al.
Publicado: (2025)
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
por: McKinzie, Brandon, et al.
Publicado: (2024)
por: McKinzie, Brandon, et al.
Publicado: (2024)
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
por: Jaiswal, Ajay, et al.
Publicado: (2025)
por: Jaiswal, Ajay, et al.
Publicado: (2025)
Multifunctional Metal–Phenolic Networks: A Game‐Changer in Food Preservation Technologies
por: Minyu Li, et al.
Publicado: (2025)
por: Minyu Li, et al.
Publicado: (2025)
A categorification for the partial-dual genus polynomial
por: Cheng, Zhiyun, et al.
Publicado: (2024)
por: Cheng, Zhiyun, et al.
Publicado: (2024)
The necessity of (co)unit in nearly Frobenius algebra
por: Cheng, Zhiyun, et al.
Publicado: (2024)
por: Cheng, Zhiyun, et al.
Publicado: (2024)
Instruction-Following Pruning for Large Language Models
por: Hou, Bairu, et al.
Publicado: (2025)
por: Hou, Bairu, et al.
Publicado: (2025)
Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training
por: Du, Xianzhi, et al.
Publicado: (2024)
por: Du, Xianzhi, et al.
Publicado: (2024)
A categorification for the signed chromatic polynomial
por: Cheng, Zhiyun, et al.
Publicado: (2021)
por: Cheng, Zhiyun, et al.
Publicado: (2021)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
por: Li, Yixiao, et al.
Publicado: (2025)
por: Li, Yixiao, et al.
Publicado: (2025)
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
por: Li, Yanghao, et al.
Publicado: (2025)
por: Li, Yanghao, et al.
Publicado: (2025)
Analysis of Bladder Cancer Staging Prediction Using Deep Residual Neural Network, Radiomics, and RNA-Seq from High-Definition CT Images
por: Yao Zhou, et al.
Publicado: (2024)
por: Yao Zhou, et al.
Publicado: (2024)
Hierarchical Multiple Kernel K-Means Algorithm Based on Sparse Connectivity
por: Wang, Lei, et al.
Publicado: (2024)
por: Wang, Lei, et al.
Publicado: (2024)
Temporal dynamics and network drivers of coral reef structural-functional relationships in the Nansha Islands, South China Sea.
por: Zhou, Yanyan, et al.
Publicado: (2026)
por: Zhou, Yanyan, et al.
Publicado: (2026)
Development and validation of a risk prediction model for lower limb lymphedema in postoperative cervical cancer patients
por: Zhiyue Li, et al.
Publicado: (2025)
por: Zhiyue Li, et al.
Publicado: (2025)
RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs
por: Zhou, Runlong, et al.
Publicado: (2025)
por: Zhou, Runlong, et al.
Publicado: (2025)
Constructing topological biquandles via skew braces
por: Cheng, Zhiyun
Publicado: (2024)
por: Cheng, Zhiyun
Publicado: (2024)
Partial-dual genus polynomial of graphs
por: Cheng, Zhiyun
Publicado: (2025)
por: Cheng, Zhiyun
Publicado: (2025)
The chord index, its definitions, applications and generalizations
por: Cheng, Zhiyun
Publicado: (2016)
por: Cheng, Zhiyun
Publicado: (2016)
Region crossing change on planar trivalent graphs
por: Cheng, Zhiyun
Publicado: (2022)
por: Cheng, Zhiyun
Publicado: (2022)
Intersection graph and writhe polynomial
por: Cheng, Zhiyun
Publicado: (2023)
por: Cheng, Zhiyun
Publicado: (2023)
Hypofractionated Radiation Therapy for Pain Relief of Patients With Spinal Metastasis: A Real‐World Analysis
por: Lu Sun, et al.
Publicado: (2026)
por: Lu Sun, et al.
Publicado: (2026)
FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information
por: Hwang, Dongseong
Publicado: (2024)
por: Hwang, Dongseong
Publicado: (2024)
The Influence of Ni 2 O 3 Doping Concentration on the Alternating Current Aging Characteristics of ZnO Varistor
por: Lei Wang, et al.
Publicado: (2025)
por: Lei Wang, et al.
Publicado: (2025)
A collision-oriented interacting particle system for Landau-type equations and the molecular chaos
por: Du, Kai, et al.
Publicado: (2024)
por: Du, Kai, et al.
Publicado: (2024)
Movement of sediments across a gently sloping muddy coast: Wave‐ and current‐supported gravity flows
por: Qian Yu, et al.
Publicado: (2024)
por: Qian Yu, et al.
Publicado: (2024)
Derived Weil Representation and Relative Langlands Duality
por: Fu, Haoshuo
Publicado: (2026)
por: Fu, Haoshuo
Publicado: (2026)
Revisiting Local PageRank Estimation on Undirected Graphs: Simple and Optimal
por: Wang, Hanzhi
Publicado: (2024)
por: Wang, Hanzhi
Publicado: (2024)
Quantitative cancer-immunity cycle modeling to optimize bevacizumab and atezolizumab combination therapy for advanced renal cell carcinoma
por: Du, Lei, et al.
Publicado: (2026)
por: Du, Lei, et al.
Publicado: (2026)
Symmetry Nonnegative Matrix Factorization Algorithm Based on Self-paced Learning
por: Wang, Lei, et al.
Publicado: (2024)
por: Wang, Lei, et al.
Publicado: (2024)
Task Adaptive Feature Distribution Based Network for Few-shot Fine-grained Target Classification
por: Li, Ping, et al.
Publicado: (2024)
por: Li, Ping, et al.
Publicado: (2024)
Relations between quantum metrology and criticality in general su(1, 1) systems
por: Zhang, Rui, et al.
Publicado: (2023)
por: Zhang, Rui, et al.
Publicado: (2023)
IVAC-P2L: Leveraging Irregular Repetition Priors for Improving Video Action Counting
por: Wang, Hang, et al.
Publicado: (2024)
por: Wang, Hang, et al.
Publicado: (2024)
Primary-Fine Decoupling for Action Generation in Robotic Imitation
por: Lei, Xiaohan, et al.
Publicado: (2026)
por: Lei, Xiaohan, et al.
Publicado: (2026)
Automatisierte Inhaltserschließung an der Bibliothek des Max-Planck-Instituts für Mathematik in den Naturwissenschaften
por: Weers, Beatrice Simon
Publicado: (2026)
por: Weers, Beatrice Simon
Publicado: (2026)
DBR: Divergence-Based Regularization for Debiasing Natural Language Understanding Models
por: Li, Zihao, et al.
Publicado: (2025)
por: Li, Zihao, et al.
Publicado: (2025)
Domain-Agnostic Mutual Prompting for Unsupervised Domain Adaptation
por: Du, Zhekai, et al.
Publicado: (2024)
por: Du, Zhekai, et al.
Publicado: (2024)
Ejemplares similares
-
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
por: Wang, Chong, et al.
Publicado: (2026) -
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
por: Findeis, Arduin, et al.
Publicado: (2025) -
EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing
por: Sun, Haotian, et al.
Publicado: (2024) -
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
por: Arora, Siddhant, et al.
Publicado: (2025) -
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
por: McKinzie, Brandon, et al.
Publicado: (2024)