Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lübbering, Max, Ruland, Timm, Rutmann, Richard, Stollenwerk, Felix, Fitzek, David, Fromm, Michael, Weber, Alexander, Sifa, Rafet, Flores-Herr, Nicolas, Köhler, Joachim, Ali, Mehdi |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
par: Liang, Wanchao, et autres
Publié: (2024)
par: Liang, Wanchao, et autres
Publié: (2024)
Comparative Evaluation of PyTorch, JAX, SciPy, and Neal for Solving QUBO Problems at Scale
par: Yang, Pei-Kun
Publié: (2025)
par: Yang, Pei-Kun
Publié: (2025)
Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
par: Nagaitsev, Kirill, et autres
Publié: (2025)
par: Nagaitsev, Kirill, et autres
Publié: (2025)
torch-sla: Differentiable Sparse Linear Algebra with Adjoint Solvers and Sparse Tensor Parallelism for PyTorch
par: Chi, Mingyuan, et autres
Publié: (2026)
par: Chi, Mingyuan, et autres
Publié: (2026)
WgPy: GPU-accelerated NumPy-like array library for web browsers
par: Hidaka, Masatoshi, et autres
Publié: (2025)
par: Hidaka, Masatoshi, et autres
Publié: (2025)
TorchGT: A Holistic System for Large-scale Graph Transformer Training
par: Zhang, Meng, et autres
Publié: (2024)
par: Zhang, Meng, et autres
Publié: (2024)
Sarus Suite: Cloud-native Containers for HPC
par: Madonna, Alberto, et autres
Publié: (2026)
par: Madonna, Alberto, et autres
Publié: (2026)
Towards cloud-native scientific workflow management
par: Orzechowski, Michal, et autres
Publié: (2024)
par: Orzechowski, Michal, et autres
Publié: (2024)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
par: Chazapis, Antony, et autres
Publié: (2024)
par: Chazapis, Antony, et autres
Publié: (2024)
Multiple Concurrent Proposers: Why and How
par: Garimidi, Pranav, et autres
Publié: (2025)
par: Garimidi, Pranav, et autres
Publié: (2025)
FIDRS: A Novel Framework for Integrated Distributed Reliable Systems
par: Gashti, Mehdi Zekriyapanah
Publié: (2025)
par: Gashti, Mehdi Zekriyapanah
Publié: (2025)
Scrutiny new framework in integrated distributed reliable systems
par: Gashti, Mehdi Zekriyapanah
Publié: (2025)
par: Gashti, Mehdi Zekriyapanah
Publié: (2025)
PyRQA -- Conducting Recurrence Quantification Analysis on Very Long Time Series Efficiently
par: Rawald, Tobias, et autres
Publié: (2024)
par: Rawald, Tobias, et autres
Publié: (2024)
TorchGWAS : GPU-accelerated GWAS for thousands of quantitative phenotypes
par: Zhao, Xingzhong, et autres
Publié: (2026)
par: Zhao, Xingzhong, et autres
Publié: (2026)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
par: Wu, Jingfeng, et autres
Publié: (2024)
par: Wu, Jingfeng, et autres
Publié: (2024)
Optimizing the Longhorn Cloud-native Software Defined Storage Engine for High Performance
par: Kampadais, Konstantinos, et autres
Publié: (2025)
par: Kampadais, Konstantinos, et autres
Publié: (2025)
Dflow, a Python framework for constructing cloud-native AI-for-Science workflows
par: Liu, Xinzijian, et autres
Publié: (2024)
par: Liu, Xinzijian, et autres
Publié: (2024)
Collaborative State Machines: A Better Programming Model for the Cloud-Edge-IoT Continuum
par: Etheredge, Marlon, et autres
Publié: (2025)
par: Etheredge, Marlon, et autres
Publié: (2025)
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
par: Ng, Nathan, et autres
Publié: (2025)
par: Ng, Nathan, et autres
Publié: (2025)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
par: Xu, Minxian, et autres
Publié: (2026)
par: Xu, Minxian, et autres
Publié: (2026)
Decentralized Learning Made Easy with DecentralizePy
par: Dhasade, Akash, et autres
Publié: (2023)
par: Dhasade, Akash, et autres
Publié: (2023)
Next Generation Cloud-native In-Memory Stores: From Redis to Valkey and Beyond
par: Rosensch"old, Carl-Johan Fauvelle Munck af, et autres
Publié: (2025)
par: Rosensch"old, Carl-Johan Fauvelle Munck af, et autres
Publié: (2025)
Visualizing Cloud-native Applications with KubeDiagrams
par: Merle, Philippe, et autres
Publié: (2025)
par: Merle, Philippe, et autres
Publié: (2025)
Exploration of Energy and Throughput Tradeoffs for Dataflow Networks
par: Karim, Abrarul, et autres
Publié: (2026)
par: Karim, Abrarul, et autres
Publié: (2026)
Toward Optimal-Complexity Hash-Based Asynchronous MVBA with Optimal Resilience
par: Komatovic, Jovan, et autres
Publié: (2024)
par: Komatovic, Jovan, et autres
Publié: (2024)
Modality Inflation: Energy Characterization and Optimization Opportunities for MLLM Inference
par: Moghadampanah, Mona, et autres
Publié: (2025)
par: Moghadampanah, Mona, et autres
Publié: (2025)
Analysis and Optimization of Wireless Multimodal Federated Learning on Modal Heterogeneity
par: Han, Xuefeng, et autres
Publié: (2025)
par: Han, Xuefeng, et autres
Publié: (2025)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
par: Tang, Xinru, et autres
Publié: (2025)
par: Tang, Xinru, et autres
Publié: (2025)
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
par: Hua, Qin, et autres
Publié: (2024)
par: Hua, Qin, et autres
Publié: (2024)
OMP4Py: a pure Python implementation of OpenMP
par: Piñeiro, César, et autres
Publié: (2024)
par: Piñeiro, César, et autres
Publié: (2024)
Multi-Modal Style Transfer-based Prompt Tuning for Efficient Federated Domain Generalization
par: Chen, Yuliang, et autres
Publié: (2026)
par: Chen, Yuliang, et autres
Publié: (2026)
DRackSim: Simulator for Rack-scale Memory Disaggregation
par: Puri, Amit, et autres
Publié: (2023)
par: Puri, Amit, et autres
Publié: (2023)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
par: Hu, Tiancheng, et autres
Publié: (2026)
par: Hu, Tiancheng, et autres
Publié: (2026)
Towards Cloud Efficiency with Large-scale Workload Characterization
par: Parayil, Anjaly, et autres
Publié: (2024)
par: Parayil, Anjaly, et autres
Publié: (2024)
Navigating the Energy Doldrums: Can We Exploit Energy-Price Volatility To Lower the Cost of Computing?
par: Arzt, Peter, et autres
Publié: (2025)
par: Arzt, Peter, et autres
Publié: (2025)
Tolerating Disasters with Hierarchical Consensus
par: Yahyaoui, Wassim, et autres
Publié: (2025)
par: Yahyaoui, Wassim, et autres
Publié: (2025)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
par: Mi, Liang, et autres
Publié: (2026)
par: Mi, Liang, et autres
Publié: (2026)
Portable, heterogeneous ensemble workflows at scale using libEnsemble
par: Hudson, Stephen, et autres
Publié: (2024)
par: Hudson, Stephen, et autres
Publié: (2024)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
par: Zhong, Yuchen, et autres
Publié: (2023)
par: Zhong, Yuchen, et autres
Publié: (2023)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
par: Xu, Zhihao, et autres
Publié: (2025)
par: Xu, Zhihao, et autres
Publié: (2025)
Documents similaires
-
TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
par: Liang, Wanchao, et autres
Publié: (2024) -
Comparative Evaluation of PyTorch, JAX, SciPy, and Neal for Solving QUBO Problems at Scale
par: Yang, Pei-Kun
Publié: (2025) -
Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
par: Nagaitsev, Kirill, et autres
Publié: (2025) -
torch-sla: Differentiable Sparse Linear Algebra with Adjoint Solvers and Sparse Tensor Parallelism for PyTorch
par: Chi, Mingyuan, et autres
Publié: (2026) -
WgPy: GPU-accelerated NumPy-like array library for web browsers
par: Hidaka, Masatoshi, et autres
Publié: (2025)