MegaFold: System-Level Optimizations for Accelerating Protein Structure Prediction Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | La, Hoa, Gupta, Ahan, Morehead, Alex, Cheng, Jianlin, Zhang, Minjia |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
par: Gupta, Ahan, et autres
Publié: (2026)
par: Gupta, Ahan, et autres
Publié: (2026)
Fold-CP: A Context Parallelism Framework for Biomolecular Modeling
par: Lin, Dejun, et autres
Publié: (2026)
par: Lin, Dejun, et autres
Publié: (2026)
Optimizations on Graph-Level for Domain Specific Computations in Julia and Application to QED
par: Reinhard, Anton, et autres
Publié: (2025)
par: Reinhard, Anton, et autres
Publié: (2025)
APACE: AlphaFold2 and advanced computing as a service for accelerated discovery in biophysics
par: Park, Hyun, et autres
Publié: (2023)
par: Park, Hyun, et autres
Publié: (2023)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
par: Rahimi, Ghazal, et autres
Publié: (2026)
par: Rahimi, Ghazal, et autres
Publié: (2026)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
par: Wang, Tuowei, et autres
Publié: (2024)
par: Wang, Tuowei, et autres
Publié: (2024)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
par: Wang, Yuxin, et autres
Publié: (2024)
par: Wang, Yuxin, et autres
Publié: (2024)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
par: Chen, David, et autres
Publié: (2026)
par: Chen, David, et autres
Publié: (2026)
Taming Cold Starts: Proactive Serverless Scheduling with Model Predictive Control
par: Nguyen, Chanh, et autres
Publié: (2025)
par: Nguyen, Chanh, et autres
Publié: (2025)
PASTA: A Modular Program Analysis Tool Framework for Accelerators
par: Lin, Mao, et autres
Publié: (2026)
par: Lin, Mao, et autres
Publié: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
par: Zhang, Li, et autres
Publié: (2025)
par: Zhang, Li, et autres
Publié: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
par: Ather, Hammad, et autres
Publié: (2024)
par: Ather, Hammad, et autres
Publié: (2024)
Accelerating Gaussian beam tracing method with dynamic parallelism on graphics processing units
par: Sheng, Zhang, et autres
Publié: (2025)
par: Sheng, Zhang, et autres
Publié: (2025)
Ridgeline: A 2D Roofline Model for Distributed Systems
par: Checconi, Fabio, et autres
Publié: (2022)
par: Checconi, Fabio, et autres
Publié: (2022)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
par: Nicusan, Andrei-Leonard, et autres
Publié: (2025)
par: Nicusan, Andrei-Leonard, et autres
Publié: (2025)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
par: Lyu, Hanzheng, et autres
Publié: (2024)
par: Lyu, Hanzheng, et autres
Publié: (2024)
Cloud Resource Allocation with Convex Optimization
par: Boghani, Shayan, et autres
Publié: (2025)
par: Boghani, Shayan, et autres
Publié: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
par: Zhao, Xuanlei, et autres
Publié: (2024)
par: Zhao, Xuanlei, et autres
Publié: (2024)
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
par: Veeravalli, Bharadwaj
Publié: (2026)
par: Veeravalli, Bharadwaj
Publié: (2026)
Inductive Loop Analysis for Practical HPC Application Optimization
par: Schaad, Philipp, et autres
Publié: (2025)
par: Schaad, Philipp, et autres
Publié: (2025)
Staging Blocked Evaluation over Structured Sparse Matrices
par: Das, Pratyush, et autres
Publié: (2024)
par: Das, Pratyush, et autres
Publié: (2024)
Disaggregated Design for GPU-Based Volumetric Data Structures
par: Meneghin, Massimiliano, et autres
Publié: (2025)
par: Meneghin, Massimiliano, et autres
Publié: (2025)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
par: Dutt, Anurag, et autres
Publié: (2025)
par: Dutt, Anurag, et autres
Publié: (2025)
EfiMon: A Process Analyser for Granular Power Consumption Prediction
par: León-Vega, Luis G., et autres
Publié: (2024)
par: León-Vega, Luis G., et autres
Publié: (2024)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
par: Islam, Tanzima Z., et autres
Publié: (2024)
par: Islam, Tanzima Z., et autres
Publié: (2024)
An Online Probabilistic Distributed Tracing System
par: Toslali, M., et autres
Publié: (2024)
par: Toslali, M., et autres
Publié: (2024)
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
par: Tariq, Syed Salauddin Mohammad, et autres
Publié: (2025)
par: Tariq, Syed Salauddin Mohammad, et autres
Publié: (2025)
Vectorization of Gradient Boosting of Decision Trees Prediction in the CatBoost Library for RISC-V Processors
par: Kozinov, Evgeny, et autres
Publié: (2024)
par: Kozinov, Evgeny, et autres
Publié: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
par: Karfakis, George, et autres
Publié: (2025)
par: Karfakis, George, et autres
Publié: (2025)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
par: Ng, Nathan, et autres
Publié: (2026)
par: Ng, Nathan, et autres
Publié: (2026)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
par: Proaño, Andrès Rubio, et autres
Publié: (2024)
par: Proaño, Andrès Rubio, et autres
Publié: (2024)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
par: Scheinert, Dominik, et autres
Publié: (2023)
par: Scheinert, Dominik, et autres
Publié: (2023)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
par: Zhang, Yaozheng, et autres
Publié: (2025)
par: Zhang, Yaozheng, et autres
Publié: (2025)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
par: Ramesh, Risshab Srinivas
Publié: (2024)
par: Ramesh, Risshab Srinivas
Publié: (2024)
Operational Strategies for Non-Disruptive Scheduling Transitions in Production HPC Systems
par: MacLachlan, Glen, et autres
Publié: (2026)
par: MacLachlan, Glen, et autres
Publié: (2026)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
par: León-Vega, Luis G., et autres
Publié: (2024)
par: León-Vega, Luis G., et autres
Publié: (2024)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
par: Xu, Jingwei, et autres
Publié: (2025)
par: Xu, Jingwei, et autres
Publié: (2025)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
par: Rashid, Md Hasanur, et autres
Publié: (2026)
par: Rashid, Md Hasanur, et autres
Publié: (2026)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
par: Mao, Ying, et autres
Publié: (2020)
par: Mao, Ying, et autres
Publié: (2020)
DIAL: Decentralized I/O AutoTuning via Learned Client-side Local Metrics for Parallel File System
par: Rashid, Md Hasanur, et autres
Publié: (2026)
par: Rashid, Md Hasanur, et autres
Publié: (2026)
Documents similaires
-
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
par: Gupta, Ahan, et autres
Publié: (2026) -
Fold-CP: A Context Parallelism Framework for Biomolecular Modeling
par: Lin, Dejun, et autres
Publié: (2026) -
Optimizations on Graph-Level for Domain Specific Computations in Julia and Application to QED
par: Reinhard, Anton, et autres
Publié: (2025) -
APACE: AlphaFold2 and advanced computing as a service for accelerated discovery in biophysics
par: Park, Hyun, et autres
Publié: (2023) -
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
par: Rahimi, Ghazal, et autres
Publié: (2026)