Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
Fuente:
arXiv
Saved in:
| Main Authors: | Lübbering, Max, Ruland, Timm, Rutmann, Richard, Stollenwerk, Felix, Fitzek, David, Fromm, Michael, Weber, Alexander, Sifa, Rafet, Flores-Herr, Nicolas, Köhler, Joachim, Ali, Mehdi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
by: Liang, Wanchao, et al.
Published: (2024)
by: Liang, Wanchao, et al.
Published: (2024)
Comparative Evaluation of PyTorch, JAX, SciPy, and Neal for Solving QUBO Problems at Scale
by: Yang, Pei-Kun
Published: (2025)
by: Yang, Pei-Kun
Published: (2025)
Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
by: Nagaitsev, Kirill, et al.
Published: (2025)
by: Nagaitsev, Kirill, et al.
Published: (2025)
torch-sla: Differentiable Sparse Linear Algebra with Adjoint Solvers and Sparse Tensor Parallelism for PyTorch
by: Chi, Mingyuan, et al.
Published: (2026)
by: Chi, Mingyuan, et al.
Published: (2026)
WgPy: GPU-accelerated NumPy-like array library for web browsers
by: Hidaka, Masatoshi, et al.
Published: (2025)
by: Hidaka, Masatoshi, et al.
Published: (2025)
TorchGT: A Holistic System for Large-scale Graph Transformer Training
by: Zhang, Meng, et al.
Published: (2024)
by: Zhang, Meng, et al.
Published: (2024)
Sarus Suite: Cloud-native Containers for HPC
by: Madonna, Alberto, et al.
Published: (2026)
by: Madonna, Alberto, et al.
Published: (2026)
Towards cloud-native scientific workflow management
by: Orzechowski, Michal, et al.
Published: (2024)
by: Orzechowski, Michal, et al.
Published: (2024)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
by: Chazapis, Antony, et al.
Published: (2024)
by: Chazapis, Antony, et al.
Published: (2024)
Multiple Concurrent Proposers: Why and How
by: Garimidi, Pranav, et al.
Published: (2025)
by: Garimidi, Pranav, et al.
Published: (2025)
FIDRS: A Novel Framework for Integrated Distributed Reliable Systems
by: Gashti, Mehdi Zekriyapanah
Published: (2025)
by: Gashti, Mehdi Zekriyapanah
Published: (2025)
Scrutiny new framework in integrated distributed reliable systems
by: Gashti, Mehdi Zekriyapanah
Published: (2025)
by: Gashti, Mehdi Zekriyapanah
Published: (2025)
PyRQA -- Conducting Recurrence Quantification Analysis on Very Long Time Series Efficiently
by: Rawald, Tobias, et al.
Published: (2024)
by: Rawald, Tobias, et al.
Published: (2024)
TorchGWAS : GPU-accelerated GWAS for thousands of quantitative phenotypes
by: Zhao, Xingzhong, et al.
Published: (2026)
by: Zhao, Xingzhong, et al.
Published: (2026)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
by: Wu, Jingfeng, et al.
Published: (2024)
by: Wu, Jingfeng, et al.
Published: (2024)
Optimizing the Longhorn Cloud-native Software Defined Storage Engine for High Performance
by: Kampadais, Konstantinos, et al.
Published: (2025)
by: Kampadais, Konstantinos, et al.
Published: (2025)
Dflow, a Python framework for constructing cloud-native AI-for-Science workflows
by: Liu, Xinzijian, et al.
Published: (2024)
by: Liu, Xinzijian, et al.
Published: (2024)
Collaborative State Machines: A Better Programming Model for the Cloud-Edge-IoT Continuum
by: Etheredge, Marlon, et al.
Published: (2025)
by: Etheredge, Marlon, et al.
Published: (2025)
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
by: Ng, Nathan, et al.
Published: (2025)
by: Ng, Nathan, et al.
Published: (2025)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
by: Xu, Minxian, et al.
Published: (2026)
by: Xu, Minxian, et al.
Published: (2026)
Decentralized Learning Made Easy with DecentralizePy
by: Dhasade, Akash, et al.
Published: (2023)
by: Dhasade, Akash, et al.
Published: (2023)
Next Generation Cloud-native In-Memory Stores: From Redis to Valkey and Beyond
by: Rosensch"old, Carl-Johan Fauvelle Munck af, et al.
Published: (2025)
by: Rosensch"old, Carl-Johan Fauvelle Munck af, et al.
Published: (2025)
Visualizing Cloud-native Applications with KubeDiagrams
by: Merle, Philippe, et al.
Published: (2025)
by: Merle, Philippe, et al.
Published: (2025)
Exploration of Energy and Throughput Tradeoffs for Dataflow Networks
by: Karim, Abrarul, et al.
Published: (2026)
by: Karim, Abrarul, et al.
Published: (2026)
Toward Optimal-Complexity Hash-Based Asynchronous MVBA with Optimal Resilience
by: Komatovic, Jovan, et al.
Published: (2024)
by: Komatovic, Jovan, et al.
Published: (2024)
Modality Inflation: Energy Characterization and Optimization Opportunities for MLLM Inference
by: Moghadampanah, Mona, et al.
Published: (2025)
by: Moghadampanah, Mona, et al.
Published: (2025)
Analysis and Optimization of Wireless Multimodal Federated Learning on Modal Heterogeneity
by: Han, Xuefeng, et al.
Published: (2025)
by: Han, Xuefeng, et al.
Published: (2025)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
by: Tang, Xinru, et al.
Published: (2025)
by: Tang, Xinru, et al.
Published: (2025)
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
by: Hua, Qin, et al.
Published: (2024)
by: Hua, Qin, et al.
Published: (2024)
OMP4Py: a pure Python implementation of OpenMP
by: Piñeiro, César, et al.
Published: (2024)
by: Piñeiro, César, et al.
Published: (2024)
Multi-Modal Style Transfer-based Prompt Tuning for Efficient Federated Domain Generalization
by: Chen, Yuliang, et al.
Published: (2026)
by: Chen, Yuliang, et al.
Published: (2026)
DRackSim: Simulator for Rack-scale Memory Disaggregation
by: Puri, Amit, et al.
Published: (2023)
by: Puri, Amit, et al.
Published: (2023)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)
by: Parayil, Anjaly, et al.
Published: (2024)
Navigating the Energy Doldrums: Can We Exploit Energy-Price Volatility To Lower the Cost of Computing?
by: Arzt, Peter, et al.
Published: (2025)
by: Arzt, Peter, et al.
Published: (2025)
Tolerating Disasters with Hierarchical Consensus
by: Yahyaoui, Wassim, et al.
Published: (2025)
by: Yahyaoui, Wassim, et al.
Published: (2025)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
by: Mi, Liang, et al.
Published: (2026)
by: Mi, Liang, et al.
Published: (2026)
Portable, heterogeneous ensemble workflows at scale using libEnsemble
by: Hudson, Stephen, et al.
Published: (2024)
by: Hudson, Stephen, et al.
Published: (2024)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
by: Zhong, Yuchen, et al.
Published: (2023)
by: Zhong, Yuchen, et al.
Published: (2023)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
by: Xu, Zhihao, et al.
Published: (2025)
by: Xu, Zhihao, et al.
Published: (2025)
Similar Items
-
TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
by: Liang, Wanchao, et al.
Published: (2024) -
Comparative Evaluation of PyTorch, JAX, SciPy, and Neal for Solving QUBO Problems at Scale
by: Yang, Pei-Kun
Published: (2025) -
Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
by: Nagaitsev, Kirill, et al.
Published: (2025) -
torch-sla: Differentiable Sparse Linear Algebra with Adjoint Solvers and Sparse Tensor Parallelism for PyTorch
by: Chi, Mingyuan, et al.
Published: (2026) -
WgPy: GPU-accelerated NumPy-like array library for web browsers
by: Hidaka, Masatoshi, et al.
Published: (2025)