UNIFERENCE: A Discrete Event Simulation Framework for Developing Distributed AI Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Eldenk, Doğaç, Xia, Stephen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Deploying Atmospheric and Oceanic AI Models on Chinese Hardware and Framework: Migration Strategies, Performance Optimization and Analysis
por: Sun, Yuze, et al.
Publicado: (2025)
por: Sun, Yuze, et al.
Publicado: (2025)
Edge-Cloud Collaborative Computing on Distributed Intelligence and Model Optimization: A Survey
por: Liu, Jing, et al.
Publicado: (2025)
por: Liu, Jing, et al.
Publicado: (2025)
Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
por: Liu, Xinyi, et al.
Publicado: (2025)
por: Liu, Xinyi, et al.
Publicado: (2025)
Game-Theoretic Deep Reinforcement Learning to Minimize Carbon Emissions and Energy Costs for AI Inference Workloads in Geo-Distributed Data Centers
por: Hogade, Ninad, et al.
Publicado: (2024)
por: Hogade, Ninad, et al.
Publicado: (2024)
Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
por: Zhang, Biyao, et al.
Publicado: (2025)
por: Zhang, Biyao, et al.
Publicado: (2025)
FedComLoc: Communication-Efficient Distributed Training of Sparse and Quantized Models
por: Yi, Kai, et al.
Publicado: (2024)
por: Yi, Kai, et al.
Publicado: (2024)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
por: Yao, Jinghan, et al.
Publicado: (2024)
por: Yao, Jinghan, et al.
Publicado: (2024)
Bayesian Federated Model Compression for Communication and Computation Efficiency
por: Xia, Chengyu, et al.
Publicado: (2024)
por: Xia, Chengyu, et al.
Publicado: (2024)
GLow -- A Novel, Flower-Based Simulated Gossip Learning Strategy
por: Belenguer, Aitor, et al.
Publicado: (2025)
por: Belenguer, Aitor, et al.
Publicado: (2025)
HSplitLoRA: A Heterogeneous Split Parameter-Efficient Fine-Tuning Framework for Large Language Models
por: Lin, Zheng, et al.
Publicado: (2025)
por: Lin, Zheng, et al.
Publicado: (2025)
Two-Timescale Model Caching and Resource Allocation for Edge-Enabled AI-Generated Content Services
por: Liu, Zhang, et al.
Publicado: (2024)
por: Liu, Zhang, et al.
Publicado: (2024)
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
por: Feng, Yicheng, et al.
Publicado: (2026)
por: Feng, Yicheng, et al.
Publicado: (2026)
Frontier: Simulating the Next Generation of LLM Inference Systems
por: Feng, Yicheng, et al.
Publicado: (2025)
por: Feng, Yicheng, et al.
Publicado: (2025)
Speeding up Policy Simulation in Supply Chain RL
por: Farias, Vivek, et al.
Publicado: (2024)
por: Farias, Vivek, et al.
Publicado: (2024)
Post-Deterministic Distributed Systems: A New Foundation for Trustworthy Autonomous Infrastructure
por: He, Jun, et al.
Publicado: (2026)
por: He, Jun, et al.
Publicado: (2026)
AB-Training: A Communication-Efficient Approach for Distributed Low-Rank Learning
por: Coquelin, Daniel, et al.
Publicado: (2024)
por: Coquelin, Daniel, et al.
Publicado: (2024)
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
por: Kadiyala, Divya Kiran, et al.
Publicado: (2022)
por: Kadiyala, Divya Kiran, et al.
Publicado: (2022)
Acceleration for Deep Reinforcement Learning using Parallel and Distributed Computing: A Survey
por: Liu, Zhihong, et al.
Publicado: (2024)
por: Liu, Zhihong, et al.
Publicado: (2024)
Adaptive Graph Pruning with Sudden-Events Evaluation for Traffic Prediction using Online Semi-Decentralized ST-GNNs
por: Kralj, Ivan, et al.
Publicado: (2025)
por: Kralj, Ivan, et al.
Publicado: (2025)
On the Fragility of Data Attribution When Learning Is Distributed
por: Gao, Xian, et al.
Publicado: (2026)
por: Gao, Xian, et al.
Publicado: (2026)
Optimal Transport Aggregation for Distributed Mixture-of-Experts
por: Chamroukhi, Faïcel, et al.
Publicado: (2023)
por: Chamroukhi, Faïcel, et al.
Publicado: (2023)
Trustworthiness of Stochastic Gradient Descent in Distributed Learning
por: Li, Hongyang, et al.
Publicado: (2024)
por: Li, Hongyang, et al.
Publicado: (2024)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
por: Deng, Xiumei, et al.
Publicado: (2025)
por: Deng, Xiumei, et al.
Publicado: (2025)
Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training
por: Brewer, Wesley, et al.
Publicado: (2025)
por: Brewer, Wesley, et al.
Publicado: (2025)
Laminar: A Scalable Asynchronous RL Post-Training Framework
por: Sheng, Guangming, et al.
Publicado: (2025)
por: Sheng, Guangming, et al.
Publicado: (2025)
Loss- and Reward-Weighting for Efficient Distributed Reinforcement Learning
por: Holen, Martin, et al.
Publicado: (2023)
por: Holen, Martin, et al.
Publicado: (2023)
Adaptive Consensus Gradients Aggregation for Scaled Distributed Training
por: Choukroun, Yoni, et al.
Publicado: (2024)
por: Choukroun, Yoni, et al.
Publicado: (2024)
Measuring Heterogeneity in Machine Learning with Distributed Energy Distance
por: Fan, Mengchen, et al.
Publicado: (2025)
por: Fan, Mengchen, et al.
Publicado: (2025)
Distributed Low-Communication Training with Decoupled Momentum Optimization
por: Nedelkoski, Sasho, et al.
Publicado: (2025)
por: Nedelkoski, Sasho, et al.
Publicado: (2025)
Trillion Parameter AI Serving Infrastructure for Scientific Discovery: A Survey and Vision
por: Hudson, Nathaniel, et al.
Publicado: (2024)
por: Hudson, Nathaniel, et al.
Publicado: (2024)
A Feature Engineering Approach for Business Impact-Oriented Failure Detection in Distributed Instant Payment Systems
por: Porcelli, Lorenzo
Publicado: (2025)
por: Porcelli, Lorenzo
Publicado: (2025)
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
por: Lu, Yunchi, et al.
Publicado: (2025)
por: Lu, Yunchi, et al.
Publicado: (2025)
FedCore: Straggler-Free Federated Learning with Distributed Coresets
por: Guo, Hongpeng, et al.
Publicado: (2024)
por: Guo, Hongpeng, et al.
Publicado: (2024)
DistShap: Scalable GNN Explanations with Distributed Shapley Values
por: Akkas, Selahattin, et al.
Publicado: (2025)
por: Akkas, Selahattin, et al.
Publicado: (2025)
Efficient Resource Scheduling for Distributed Infrastructures Using Negotiation Capabilities
por: Chu, Junjie, et al.
Publicado: (2024)
por: Chu, Junjie, et al.
Publicado: (2024)
Adaptive and Resource-efficient Agentic AI Systems for Mobile and Embedded Devices: A Survey
por: Liu, Sicong, et al.
Publicado: (2025)
por: Liu, Sicong, et al.
Publicado: (2025)
Efficient and Scalable Agentic AI with Heterogeneous Systems
por: Asgar, Zain, et al.
Publicado: (2025)
por: Asgar, Zain, et al.
Publicado: (2025)
DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
por: Huang, Zih-Hao, et al.
Publicado: (2025)
por: Huang, Zih-Hao, et al.
Publicado: (2025)
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
por: Ockerman, Seth, et al.
Publicado: (2025)
por: Ockerman, Seth, et al.
Publicado: (2025)
ATTENTION2D: Communication Efficient Distributed Self-Attention Mechanism
por: Elango, Venmugil
Publicado: (2025)
por: Elango, Venmugil
Publicado: (2025)
Ejemplares similares
-
Deploying Atmospheric and Oceanic AI Models on Chinese Hardware and Framework: Migration Strategies, Performance Optimization and Analysis
por: Sun, Yuze, et al.
Publicado: (2025) -
Edge-Cloud Collaborative Computing on Distributed Intelligence and Model Optimization: A Survey
por: Liu, Jing, et al.
Publicado: (2025) -
Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
por: Liu, Xinyi, et al.
Publicado: (2025) -
Game-Theoretic Deep Reinforcement Learning to Minimize Carbon Emissions and Energy Costs for AI Inference Workloads in Geo-Distributed Data Centers
por: Hogade, Ninad, et al.
Publicado: (2024) -
Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
por: Zhang, Biyao, et al.
Publicado: (2025)