Multi-Resolution Model Fusion for Accelerating the Convolutional Neural Network Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Kewei, Lee, Claire Songhyun, Lee, Sunwoo, Gupta, Vishu, Balewski, Jan, Sim, Alex, Nugent, Peter, Agrawal, Ankit, Choudhary, Alok, Wu, Kesheng, Liao, Wei-keng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Parallel Data Object Creation: Towards Scalable Metadata Management in High-Performance I/O Library
von: Li, Youjia, et al.
Veröffentlicht: (2025)
von: Li, Youjia, et al.
Veröffentlicht: (2025)
Improving Slow Transfer Predictions: Generative Methods Compared
von: Kim, Jacob Taegon, et al.
Veröffentlicht: (2025)
von: Kim, Jacob Taegon, et al.
Veröffentlicht: (2025)
Embracing Federated Learning: Enabling Weak Client Participation via Partial Model Training
von: Lee, Sunwoo, et al.
Veröffentlicht: (2024)
von: Lee, Sunwoo, et al.
Veröffentlicht: (2024)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
von: Li, Shiju, et al.
Veröffentlicht: (2025)
von: Li, Shiju, et al.
Veröffentlicht: (2025)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
Composing Distributed Computations Through Task and Kernel Fusion
von: Yadav, Rohan, et al.
Veröffentlicht: (2024)
von: Yadav, Rohan, et al.
Veröffentlicht: (2024)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Accelerating Particle-Mesh Algorithms with FPGAs and OmpSs@OpenCL
von: Guidotti, Nicolas Lee
Veröffentlicht: (2025)
von: Guidotti, Nicolas Lee
Veröffentlicht: (2025)
A Real-Time, Auto-Regression Method for In-Situ Feature Extraction in Hydrodynamics Simulations
von: Yan, Kewei, et al.
Veröffentlicht: (2025)
von: Yan, Kewei, et al.
Veröffentlicht: (2025)
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
von: Lee, Wonbeom, et al.
Veröffentlicht: (2024)
von: Lee, Wonbeom, et al.
Veröffentlicht: (2024)
Towards Scalable GPU-Accelerated SNN Training via Temporal Fusion
von: Li, Yanchen, et al.
Veröffentlicht: (2024)
von: Li, Yanchen, et al.
Veröffentlicht: (2024)
Accelerating Compound LLM Training Workloads with Maestro
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
von: Tang, Yupeng, et al.
Veröffentlicht: (2023)
von: Tang, Yupeng, et al.
Veröffentlicht: (2023)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
City-Scale Visibility Graph Analysis via GPU-Accelerated HyperBall
von: Hodge, Alex, et al.
Veröffentlicht: (2026)
von: Hodge, Alex, et al.
Veröffentlicht: (2026)
Distributed Matrix-Based Sampling for Graph Neural Network Training
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
Nezha: Breaking Multi-Rail Network Barriers for Distributed DNN Training
von: Yu, Enda, et al.
Veröffentlicht: (2024)
von: Yu, Enda, et al.
Veröffentlicht: (2024)
Accelerating Depthwise Separable Convolutions on Ultra-Low-Power Devices
von: Daghero, Francesco, et al.
Veröffentlicht: (2024)
von: Daghero, Francesco, et al.
Veröffentlicht: (2024)
Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
von: Park, Seongyeon, et al.
Veröffentlicht: (2024)
von: Park, Seongyeon, et al.
Veröffentlicht: (2024)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
von: Fang, Jiahao, et al.
Veröffentlicht: (2024)
von: Fang, Jiahao, et al.
Veröffentlicht: (2024)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
SCARIF: Towards Carbon Modeling of Cloud Servers with Accelerators
von: Ji, Shixin, et al.
Veröffentlicht: (2024)
von: Ji, Shixin, et al.
Veröffentlicht: (2024)
Pagoda: An Energy and Time Roofline Study for DNN Workloads on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Future-Proofing IoT: Unleashing the Power of AWS Greengrass in Propelling Smart Devices to New Heights
von: Kokkula, Sahasra, et al.
Veröffentlicht: (2024)
von: Kokkula, Sahasra, et al.
Veröffentlicht: (2024)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025)
Distributed Convolutional Neural Network Training on Mobile and Edge Clusters
von: Rama, Pranav, et al.
Veröffentlicht: (2024)
von: Rama, Pranav, et al.
Veröffentlicht: (2024)
GPT-FL: Generative Pre-trained Model-Assisted Federated Learning
von: Zhang, Tuo, et al.
Veröffentlicht: (2023)
von: Zhang, Tuo, et al.
Veröffentlicht: (2023)
Cross-region Model Training with Communication-Computation Overlapping and Delay Compensation
von: Zhu, Ying, et al.
Veröffentlicht: (2025)
von: Zhu, Ying, et al.
Veröffentlicht: (2025)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
von: Song, Jaeyong, et al.
Veröffentlicht: (2026)
von: Song, Jaeyong, et al.
Veröffentlicht: (2026)
An Edge-based WiFi Fingerprinting Indoor Localization Using Convolutional Neural Network and Convolutional Auto-Encoder
von: Kargar-Barzi, Amin, et al.
Veröffentlicht: (2023)
von: Kargar-Barzi, Amin, et al.
Veröffentlicht: (2023)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025)
von: Park, Jay H., et al.
Veröffentlicht: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
Unleashing Multicore Strength for Efficient Execution of Transactions
von: Ravish, Ankit, et al.
Veröffentlicht: (2024)
von: Ravish, Ankit, et al.
Veröffentlicht: (2024)
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
von: Lee, Haeun, et al.
Veröffentlicht: (2025)
von: Lee, Haeun, et al.
Veröffentlicht: (2025)
ClusterFusion++: Expanding Cluster-Level Fusion to Full Transformer-Block Decoding
von: Jin, ChiHeng, et al.
Veröffentlicht: (2026)
von: Jin, ChiHeng, et al.
Veröffentlicht: (2026)
Photon: Federated LLM Pre-Training
von: Sani, Lorenzo, et al.
Veröffentlicht: (2024)
von: Sani, Lorenzo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Parallel Data Object Creation: Towards Scalable Metadata Management in High-Performance I/O Library
von: Li, Youjia, et al.
Veröffentlicht: (2025) -
Improving Slow Transfer Predictions: Generative Methods Compared
von: Kim, Jacob Taegon, et al.
Veröffentlicht: (2025) -
Embracing Federated Learning: Enabling Weak Client Participation via Partial Model Training
von: Lee, Sunwoo, et al.
Veröffentlicht: (2024) -
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
von: Li, Shiju, et al.
Veröffentlicht: (2025) -
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)