Gespeichert in:
| Hauptverfasser: | Wang, Hansheng, Shi, Lu, duan, Zhekai, Wu, Panruo, Guo, Liwei, Zhang, Shaoshuai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2410.02170 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
von: Wang, Hansheng, et al.
Veröffentlicht: (2025)
von: Wang, Hansheng, et al.
Veröffentlicht: (2025)
Pipelet: Practical Streamlined Blockchain Protocol
von: Karihaloo, Vivek, et al.
Veröffentlicht: (2024)
von: Karihaloo, Vivek, et al.
Veröffentlicht: (2024)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
von: Reisecker, Maximilian, et al.
Veröffentlicht: (2025)
von: Reisecker, Maximilian, et al.
Veröffentlicht: (2025)
PRISM: Processing-In-Memory Sparse MTTKRP for Tensor Decomposition Acceleration
von: Pacheco, Daniel, et al.
Veröffentlicht: (2026)
von: Pacheco, Daniel, et al.
Veröffentlicht: (2026)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
Communication-Efficient Model Aggregation with Layer Divergence Feedback in Federated Learning
von: Wang, Liwei, et al.
Veröffentlicht: (2024)
von: Wang, Liwei, et al.
Veröffentlicht: (2024)
Experimental Evaluation of Distributed k-Core Decomposition
von: Guo, Bin, et al.
Veröffentlicht: (2024)
von: Guo, Bin, et al.
Veröffentlicht: (2024)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026)
von: Liang, Yan, et al.
Veröffentlicht: (2026)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
von: Feng, Weigang, et al.
Veröffentlicht: (2025)
von: Feng, Weigang, et al.
Veröffentlicht: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
Federated k-Core Decomposition: A Secure Distributed Approach
von: Guo, Bin, et al.
Veröffentlicht: (2024)
von: Guo, Bin, et al.
Veröffentlicht: (2024)
Exploiting Multicast for Accelerating Collective Communication
von: Xu, Chao, et al.
Veröffentlicht: (2026)
von: Xu, Chao, et al.
Veröffentlicht: (2026)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
von: Zhang, Qiao, et al.
Veröffentlicht: (2025)
von: Zhang, Qiao, et al.
Veröffentlicht: (2025)
Accelerating Biclique Counting on GPU
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
Investigating Sharding Advancements, Methodologies, and Adoption Potential in Hedera
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
Accelerating Sparse DNNs Based on Tiled GEMM
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
von: Wu, Tian, et al.
Veröffentlicht: (2025)
von: Wu, Tian, et al.
Veröffentlicht: (2025)
Federated Learning Using Coupled Tensor Train Decomposition
von: Zhang, Xiangtao, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangtao, et al.
Veröffentlicht: (2024)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
From Symmetric to Asymmetric Asynchronous Byzantine Consensus
von: Cachin, Christian, et al.
Veröffentlicht: (2020)
von: Cachin, Christian, et al.
Veröffentlicht: (2020)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
HMTRace: Hardware-Assisted Memory-Tagging based Dynamic Data Race Detection
von: Shastri, Jaidev, et al.
Veröffentlicht: (2024)
von: Shastri, Jaidev, et al.
Veröffentlicht: (2024)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
GPZ: GPU-Accelerated Lossy Compressor for Particle Data
von: Li, Ruoyu, et al.
Veröffentlicht: (2025)
von: Li, Ruoyu, et al.
Veröffentlicht: (2025)
Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations
von: Daas, Hussam Al, et al.
Veröffentlicht: (2024)
von: Daas, Hussam Al, et al.
Veröffentlicht: (2024)
Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization
von: Zhou, Yangjie, et al.
Veröffentlicht: (2024)
von: Zhou, Yangjie, et al.
Veröffentlicht: (2024)
SCARIF: Towards Carbon Modeling of Cloud Servers with Accelerators
von: Ji, Shixin, et al.
Veröffentlicht: (2024)
von: Ji, Shixin, et al.
Veröffentlicht: (2024)
Enhancing ASIC Technology Mapping via Parallel Supergate Computing
von: Cai, Ye, et al.
Veröffentlicht: (2024)
von: Cai, Ye, et al.
Veröffentlicht: (2024)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
Benchmarking Compound AI Applications for Hardware-Software Co-Design
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026)
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
von: Wang, Hansheng, et al.
Veröffentlicht: (2025) -
Pipelet: Practical Streamlined Blockchain Protocol
von: Karihaloo, Vivek, et al.
Veröffentlicht: (2024) -
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025) -
Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
von: Reisecker, Maximilian, et al.
Veröffentlicht: (2025) -
PRISM: Processing-In-Memory Sparse MTTKRP for Tensor Decomposition Acceleration
von: Pacheco, Daniel, et al.
Veröffentlicht: (2026)