Gespeichert in:
| Hauptverfasser: | Kulkarni, Sudhanshu, Loring, Burlen, Bethel, E. Wes |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2402.01843 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Fast Fourier Transforms on the Tenstorrent Wormhole
von: Brown, Nick, et al.
Veröffentlicht: (2025)
von: Brown, Nick, et al.
Veröffentlicht: (2025)
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
von: Lu, Yishun, et al.
Veröffentlicht: (2026)
von: Lu, Yishun, et al.
Veröffentlicht: (2026)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
Towards a Testbed for Scalable FaaS Platforms
von: Schirmer, Trever, et al.
Veröffentlicht: (2025)
von: Schirmer, Trever, et al.
Veröffentlicht: (2025)
Transforming Lock-free Linked Lists into Distributed Lock-free Linked Lists
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2025)
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2025)
AI-coupled HPC Workflow Applications, Middleware and Performance
von: Brewer, Wes, et al.
Veröffentlicht: (2024)
von: Brewer, Wes, et al.
Veröffentlicht: (2024)
Towards Efficient and Scalable Distributed Vector Search with RDMA
von: Zhi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhi, Xiangyu, et al.
Veröffentlicht: (2025)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
von: Remke, Stefan, et al.
Veröffentlicht: (2024)
von: Remke, Stefan, et al.
Veröffentlicht: (2024)
Towards Fine-Grained Scalability for Stateful Stream Processing Systems
von: Qing, Yunfan, et al.
Veröffentlicht: (2025)
von: Qing, Yunfan, et al.
Veröffentlicht: (2025)
emucxl: an emulation framework for CXL-based disaggregated memory applications
von: Gond, Raja, et al.
Veröffentlicht: (2024)
von: Gond, Raja, et al.
Veröffentlicht: (2024)
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
Tolerance to Asynchrony of an Algorithm for Gathering Myopic Robots on an Infinite Triangular Grid
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2023)
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2023)
Fully Lattice-Linear Algorithms
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2022)
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2022)
Tolerance to Asynchrony in Algorithms for Multiplication and Modulo
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2023)
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2023)
Approximated Coded Computing: Towards Fast, Private and Secure Distributed Machine Learning
von: Qiu, Houming, et al.
Veröffentlicht: (2024)
von: Qiu, Houming, et al.
Veröffentlicht: (2024)
FPTC: A Fast Parallel Transform-based Codec for Efficient Asymmetric Signal Compression
von: Mechels, Ben, et al.
Veröffentlicht: (2026)
von: Mechels, Ben, et al.
Veröffentlicht: (2026)
FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
von: Zhang, Haoyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyuan, et al.
Veröffentlicht: (2025)
CloudFix: Automated Policy Repair for Cloud Access Control Policies Using Large Language Models
von: Hall, Bethel, et al.
Veröffentlicht: (2025)
von: Hall, Bethel, et al.
Veröffentlicht: (2025)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
Parallel Data Object Creation: Towards Scalable Metadata Management in High-Performance I/O Library
von: Li, Youjia, et al.
Veröffentlicht: (2025)
von: Li, Youjia, et al.
Veröffentlicht: (2025)
Asynchronous Checkpoint for Eventually Consistent Databases
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2025)
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2025)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
von: Curless, Brian, et al.
Veröffentlicht: (2025)
von: Curless, Brian, et al.
Veröffentlicht: (2025)
Distributing Context-Aware Shared Memory Data Structures: A Case Study on Singly-Linked Lists
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2024)
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2024)
Characterizing Production GPU Workloads using System-wide Telemetry Data
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
von: Pennino, Diego, et al.
Veröffentlicht: (2024)
von: Pennino, Diego, et al.
Veröffentlicht: (2024)
Scalable and Performant Data Loading
von: Hira, Moto, et al.
Veröffentlicht: (2025)
von: Hira, Moto, et al.
Veröffentlicht: (2025)
Fast-HotStuff: A Fast and Resilient HotStuff Protocol
von: Jalalzai, Mohammad M., et al.
Veröffentlicht: (2020)
von: Jalalzai, Mohammad M., et al.
Veröffentlicht: (2020)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
von: Kulkarni, Apurv Deepak, et al.
Veröffentlicht: (2025)
von: Kulkarni, Apurv Deepak, et al.
Veröffentlicht: (2025)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
Pilotfish: Distributed Execution for Scalable Blockchains
von: Kniep, Quentin, et al.
Veröffentlicht: (2024)
von: Kniep, Quentin, et al.
Veröffentlicht: (2024)
Scalable Maxflow Processing for Dynamic Graphs
von: Kannappan, Shruthi, et al.
Veröffentlicht: (2025)
von: Kannappan, Shruthi, et al.
Veröffentlicht: (2025)
Robust and Scalable Renaming with Subquadratic Bits
von: Bai, Sirui, et al.
Veröffentlicht: (2025)
von: Bai, Sirui, et al.
Veröffentlicht: (2025)
Fault-Tolerant Decentralized Distributed Asynchronous Federated Learning with Adaptive Termination Detection
von: Akkinepally, Phani Sahasra, et al.
Veröffentlicht: (2025)
von: Akkinepally, Phani Sahasra, et al.
Veröffentlicht: (2025)
A Fast Confirmation Rule (aka Fast Synchronous Finality) for the Ethereum Consensus Protocol
von: Asgaonkar, Aditya, et al.
Veröffentlicht: (2024)
von: Asgaonkar, Aditya, et al.
Veröffentlicht: (2024)
FastGraph: Optimized GPU-Enabled Algorithms for Fast Graph Building and Message Passing
von: Agarwal, Aarush, et al.
Veröffentlicht: (2025)
von: Agarwal, Aarush, et al.
Veröffentlicht: (2025)
TD-Orch: Scalable Load-Balancing for Distributed Systems with Applications to Graph Processing
von: Zhao, Yiwei, et al.
Veröffentlicht: (2025)
von: Zhao, Yiwei, et al.
Veröffentlicht: (2025)
FLeeC: a Fast Lock-Free Application Cache
von: Costa, André J., et al.
Veröffentlicht: (2024)
von: Costa, André J., et al.
Veröffentlicht: (2024)
Wilkins: HPC In Situ Workflows Made Easy
von: Yildiz, Orcun, et al.
Veröffentlicht: (2024)
von: Yildiz, Orcun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Exploring Fast Fourier Transforms on the Tenstorrent Wormhole
von: Brown, Nick, et al.
Veröffentlicht: (2025) -
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
von: Wu, Shixun, et al.
Veröffentlicht: (2024) -
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
von: Lu, Yishun, et al.
Veröffentlicht: (2026) -
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
von: Wu, Shixun, et al.
Veröffentlicht: (2024) -
Towards a Testbed for Scalable FaaS Platforms
von: Schirmer, Trever, et al.
Veröffentlicht: (2025)