DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Xu, Liu, Ke, Chang, Yisong, Zhang, Ke, Chen, Mingyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
por: Vo, Huynh Q. N., et al.
Publicado: (2025)
por: Vo, Huynh Q. N., et al.
Publicado: (2025)
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
por: Call, Aaron, et al.
Publicado: (2025)
por: Call, Aaron, et al.
Publicado: (2025)
CLAASIC: a Cortex-Inspired Hardware Accelerator
por: Puente, Valentin, et al.
Publicado: (2016)
por: Puente, Valentin, et al.
Publicado: (2016)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
por: Afzal, Ayesha, et al.
Publicado: (2026)
por: Afzal, Ayesha, et al.
Publicado: (2026)
ForgetMeNot: Understanding and Modeling the Impact of Forever Chemicals Toward Sustainable Large-Scale Computing
por: Roy, Rohan Basu, et al.
Publicado: (2025)
por: Roy, Rohan Basu, et al.
Publicado: (2025)
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
por: Tithi, Jesmin Jahan, et al.
Publicado: (2025)
por: Tithi, Jesmin Jahan, et al.
Publicado: (2025)
Managed-Retention Memory: A New Class of Memory for the AI Era
por: Legtchenko, Sergey, et al.
Publicado: (2025)
por: Legtchenko, Sergey, et al.
Publicado: (2025)
Carbon Connect: An Ecosystem for Sustainable Computing
por: Lee, Benjamin C., et al.
Publicado: (2024)
por: Lee, Benjamin C., et al.
Publicado: (2024)
TreeVQA: A Tree-Structured Execution Framework for Shot Reduction in Variational Quantum Algorithms
por: Hou, Yuewen, et al.
Publicado: (2025)
por: Hou, Yuewen, et al.
Publicado: (2025)
Architecting Distributed Quantum Computers: Design Insights from Resource Estimation
por: Filippov, Dmitry, et al.
Publicado: (2025)
por: Filippov, Dmitry, et al.
Publicado: (2025)
Reference Architecture of a Quantum-Centric Supercomputer
por: Seelam, Seetharami, et al.
Publicado: (2026)
por: Seelam, Seetharami, et al.
Publicado: (2026)
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
por: Ortega, Cristobal, et al.
Publicado: (2024)
por: Ortega, Cristobal, et al.
Publicado: (2024)
Transforming the Hybrid Cloud for Emerging AI Workloads
por: Chen, Deming, et al.
Publicado: (2024)
por: Chen, Deming, et al.
Publicado: (2024)
WaferLLM: Large Language Model Inference at Wafer Scale
por: He, Congjie, et al.
Publicado: (2025)
por: He, Congjie, et al.
Publicado: (2025)
Harnessing the Full Potential of RRAMs through Scalable and Distributed In-Memory Computing with Integrated Error Correction
por: Vo, Huynh Q. N., et al.
Publicado: (2025)
por: Vo, Huynh Q. N., et al.
Publicado: (2025)
Efficient Optimization Accelerator Framework for Multistate Ising Problems
por: Garg, Chirag, et al.
Publicado: (2025)
por: Garg, Chirag, et al.
Publicado: (2025)
PASS: An Asynchronous Probabilistic Processor for Next Generation Intelligence
por: Patel, Saavan, et al.
Publicado: (2024)
por: Patel, Saavan, et al.
Publicado: (2024)
COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators
por: Park, Jihoon, et al.
Publicado: (2025)
por: Park, Jihoon, et al.
Publicado: (2025)
Experience Deploying Containerized GenAI Services at an HPC Center
por: Beltre, Angel M., et al.
Publicado: (2025)
por: Beltre, Angel M., et al.
Publicado: (2025)
Datapath Combinational Equivalence Checking With Hybrid Sweeping Engines and Parallelization
por: Chen, Zhihan, et al.
Publicado: (2024)
por: Chen, Zhihan, et al.
Publicado: (2024)
Pooling Engram Conditional Memory in Large Language Models using CXL
por: Ma, Ruiyang, et al.
Publicado: (2026)
por: Ma, Ruiyang, et al.
Publicado: (2026)
CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling
por: Xu, Dong, et al.
Publicado: (2026)
por: Xu, Dong, et al.
Publicado: (2026)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
por: Wang, Xi, et al.
Publicado: (2024)
por: Wang, Xi, et al.
Publicado: (2024)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
por: Ma, Ke, et al.
Publicado: (2025)
por: Ma, Ke, et al.
Publicado: (2025)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
por: Li, Ming, et al.
Publicado: (2024)
por: Li, Ming, et al.
Publicado: (2024)
Highly Versatile FPGA-Implemented Cyber Coherent Ising Machine
por: Aonishi, Toru, et al.
Publicado: (2024)
por: Aonishi, Toru, et al.
Publicado: (2024)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
por: Zhang, Chen, et al.
Publicado: (2026)
por: Zhang, Chen, et al.
Publicado: (2026)
A Programming Model for Disaggregated Memory over CXL
por: Assa, Gal, et al.
Publicado: (2024)
por: Assa, Gal, et al.
Publicado: (2024)
On-Package Memory with Universal Chiplet Interconnect Express (UCIe): A Low Power, High Bandwidth, Low Latency and Low Cost Approach
por: Sharma, Debendra Das, et al.
Publicado: (2025)
por: Sharma, Debendra Das, et al.
Publicado: (2025)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
por: Qu, Huanyu, et al.
Publicado: (2025)
por: Qu, Huanyu, et al.
Publicado: (2025)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
por: Zou, An, et al.
Publicado: (2025)
por: Zou, An, et al.
Publicado: (2025)
Parendi: Thousand-Way Parallel RTL Simulation
por: Emami, Mahyar, et al.
Publicado: (2024)
por: Emami, Mahyar, et al.
Publicado: (2024)
SLA Conceptual Model for IoT Applications
por: Alqahtani, Awatif, et al.
Publicado: (2024)
por: Alqahtani, Awatif, et al.
Publicado: (2024)
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
por: Wang, Xi, et al.
Publicado: (2025)
por: Wang, Xi, et al.
Publicado: (2025)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
por: Wang, Yuqing, et al.
Publicado: (2025)
por: Wang, Yuqing, et al.
Publicado: (2025)
LACIN: Linearly Arranged Complete Interconnection Networks
por: Beivide, Ramón, et al.
Publicado: (2026)
por: Beivide, Ramón, et al.
Publicado: (2026)
Flex-PE: Flexible and SIMD Multi-Precision Processing Element for AI Workloads
por: Lokhande, Mukul, et al.
Publicado: (2024)
por: Lokhande, Mukul, et al.
Publicado: (2024)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
por: Liu, Xingyu, et al.
Publicado: (2025)
por: Liu, Xingyu, et al.
Publicado: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
por: Zhang, Yichao, et al.
Publicado: (2026)
por: Zhang, Yichao, et al.
Publicado: (2026)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
Ejemplares similares
-
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
por: Vo, Huynh Q. N., et al.
Publicado: (2025) -
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
por: Call, Aaron, et al.
Publicado: (2025) -
CLAASIC: a Cortex-Inspired Hardware Accelerator
por: Puente, Valentin, et al.
Publicado: (2016) -
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
por: Afzal, Ayesha, et al.
Publicado: (2026) -
ForgetMeNot: Understanding and Modeling the Impact of Forever Chemicals Toward Sustainable Large-Scale Computing
por: Roy, Rohan Basu, et al.
Publicado: (2025)