Deploying Atmospheric and Oceanic AI Models on Chinese Hardware and Framework: Migration Strategies, Performance Optimization and Analysis
Fuente:
arXiv
Guardado en:
| Autores principales: | Sun, Yuze, Luo, Wentao, Xiang, Yanfei, Pan, Jiancheng, Li, Jiahao, Zhang, Quan, Huang, Xiaomeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
por: Patke, Archit, et al.
Publicado: (2025)
por: Patke, Archit, et al.
Publicado: (2025)
On the Impact of White-box Deployment Strategies for Edge AI on Latency and Model Performance
por: Singh, Jaskirat, et al.
Publicado: (2024)
por: Singh, Jaskirat, et al.
Publicado: (2024)
Distributed Locking: Performance Analysis and Optimization Strategies
por: Rodriguez, Andre, et al.
Publicado: (2025)
por: Rodriguez, Andre, et al.
Publicado: (2025)
Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization
por: Zhou, Yangjie, et al.
Publicado: (2024)
por: Zhou, Yangjie, et al.
Publicado: (2024)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
por: Li, Haley, et al.
Publicado: (2026)
por: Li, Haley, et al.
Publicado: (2026)
Benchmarking Compound AI Applications for Hardware-Software Co-Design
por: Samuthrsindh, Paramuth, et al.
Publicado: (2026)
por: Samuthrsindh, Paramuth, et al.
Publicado: (2026)
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
por: Shubham, et al.
Publicado: (2024)
por: Shubham, et al.
Publicado: (2024)
An Optimized Error-controlled MPI Collective Framework Integrated with Lossy Compression
por: Huang, Jiajun, et al.
Publicado: (2023)
por: Huang, Jiajun, et al.
Publicado: (2023)
FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
por: Guo, Xianwen, et al.
Publicado: (2025)
por: Guo, Xianwen, et al.
Publicado: (2025)
Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
por: Reisecker, Maximilian, et al.
Publicado: (2025)
por: Reisecker, Maximilian, et al.
Publicado: (2025)
Performance Evaluation of Automated Multi-Service Deployment in Edge-Cloud Environments with the CODECO Toolkit
por: Koukis, Georgios, et al.
Publicado: (2026)
por: Koukis, Georgios, et al.
Publicado: (2026)
Agentic AI-Driven UAV Network Deployment: An LLM-Enhanced Exact Potential Game Approach
por: Tang, Xin, et al.
Publicado: (2026)
por: Tang, Xin, et al.
Publicado: (2026)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
por: Arima, Eishi, et al.
Publicado: (2024)
por: Arima, Eishi, et al.
Publicado: (2024)
Examining MPI and its Extensions for Asynchronous Multithreaded Communication
por: Yan, Jiakun, et al.
Publicado: (2025)
por: Yan, Jiakun, et al.
Publicado: (2025)
CIR: Lightweight Container Image for Cross-Platform Deployment
por: Li, Fengzhi, et al.
Publicado: (2026)
por: Li, Fengzhi, et al.
Publicado: (2026)
ML-ECS: A Collaborative Multimodal Learning Framework for Edge-Cloud Synergies
por: Liu, Yuze, et al.
Publicado: (2026)
por: Liu, Yuze, et al.
Publicado: (2026)
Fog Device-as-a-Service (FDaaS): A Framework for Service Deployment in Public Fog Environments
por: Battula, Sudheer Kumar, et al.
Publicado: (2023)
por: Battula, Sudheer Kumar, et al.
Publicado: (2023)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
por: Tschand, Arya, et al.
Publicado: (2025)
por: Tschand, Arya, et al.
Publicado: (2025)
HyperParallel: A Supernode-Affinity AI Framework
por: Zhang, Xin, et al.
Publicado: (2026)
por: Zhang, Xin, et al.
Publicado: (2026)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
por: Sun, Desen, et al.
Publicado: (2025)
por: Sun, Desen, et al.
Publicado: (2025)
TaPS: A Performance Evaluation Suite for Task-based Execution Frameworks
por: Pauloski, J. Gregory, et al.
Publicado: (2024)
por: Pauloski, J. Gregory, et al.
Publicado: (2024)
Generating Bindings in MPICH
por: Zhou, Hui, et al.
Publicado: (2024)
por: Zhou, Hui, et al.
Publicado: (2024)
KubeDSM: A Kubernetes-based Dynamic Scheduling and Migration Framework for Cloud-Assisted Edge Clusters
por: Pashaeehir, Amirhossein, et al.
Publicado: (2025)
por: Pashaeehir, Amirhossein, et al.
Publicado: (2025)
Dynamic Client Clustering, Bandwidth Allocation, and Workload Optimization for Semi-synchronous Federated Learning
por: Yu, Liangkun, et al.
Publicado: (2024)
por: Yu, Liangkun, et al.
Publicado: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
por: He, Yiyuan, et al.
Publicado: (2025)
por: He, Yiyuan, et al.
Publicado: (2025)
ProvDeploy: Provenance-oriented Containerization of High Performance Computing Scientific Workflows
por: Kunstmann, Liliane, et al.
Publicado: (2024)
por: Kunstmann, Liliane, et al.
Publicado: (2024)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
por: Islam, Tanzima Z., et al.
Publicado: (2024)
por: Islam, Tanzima Z., et al.
Publicado: (2024)
B-PASTE: Beam-Aware Pattern-Guided Speculative Execution for Resource-Constrained LLM Agents
por: Song, Yanfei
Publicado: (2026)
por: Song, Yanfei
Publicado: (2026)
Dynamic Resource Manager for Automating Deployments in the Computing Continuum
por: Samani, Zahra Najafabadi, et al.
Publicado: (2024)
por: Samani, Zahra Najafabadi, et al.
Publicado: (2024)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
por: Carl, Natalie, et al.
Publicado: (2025)
por: Carl, Natalie, et al.
Publicado: (2025)
Distributed Inference Performance Optimization for LLMs on CPUs
por: He, Pujiang, et al.
Publicado: (2024)
por: He, Pujiang, et al.
Publicado: (2024)
MoFa: A Unified Performance Modeling Framework for LLM Pretraining
por: Zhao, Lu, et al.
Publicado: (2025)
por: Zhao, Lu, et al.
Publicado: (2025)
MPI Progress For All
por: Zhou, Hui, et al.
Publicado: (2024)
por: Zhou, Hui, et al.
Publicado: (2024)
Frustrated with MPI+Threads? Try MPIxThreads!
por: Zhou, Hui, et al.
Publicado: (2024)
por: Zhou, Hui, et al.
Publicado: (2024)
Implementing True MPI Sessions and Evaluating MPI Initialization Scalability
por: Zhou, Hui, et al.
Publicado: (2026)
por: Zhou, Hui, et al.
Publicado: (2026)
Green by Design: Constraint-Based Adaptive Deployment in the Cloud Continuum
por: D'Iapico, Andrea, et al.
Publicado: (2026)
por: D'Iapico, Andrea, et al.
Publicado: (2026)
In-Transit Data Transport Strategies for Coupled AI-Simulation Workflow Patterns
por: Tummalapalli, Harikrishna, et al.
Publicado: (2025)
por: Tummalapalli, Harikrishna, et al.
Publicado: (2025)
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
por: Yang, Xiang, et al.
Publicado: (2022)
por: Yang, Xiang, et al.
Publicado: (2022)
SW-TNC : Reaching the Most Complex Random Quantum Circuit via Tensor Network Contraction
por: Chen, Yaojian, et al.
Publicado: (2025)
por: Chen, Yaojian, et al.
Publicado: (2025)
SlimEdge: Performance and Device Aware Distributed DNN Deployment on Resource-Constrained Edge Hardware
por: Kumar, Mahadev Sunil, et al.
Publicado: (2025)
por: Kumar, Mahadev Sunil, et al.
Publicado: (2025)
Ejemplares similares
-
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
por: Patke, Archit, et al.
Publicado: (2025) -
On the Impact of White-box Deployment Strategies for Edge AI on Latency and Model Performance
por: Singh, Jaskirat, et al.
Publicado: (2024) -
Distributed Locking: Performance Analysis and Optimization Strategies
por: Rodriguez, Andre, et al.
Publicado: (2025) -
Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization
por: Zhou, Yangjie, et al.
Publicado: (2024) -
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
por: Li, Haley, et al.
Publicado: (2026)