FRED: Flexible REduction-Distribution Interconnect and Communication Implementation for Wafer-Scale Distributed Training of DNN Models
Fuente:
arXiv
Saved in:
| Main Authors: | Rashidi, Saeed, Won, William, Srinivasan, Sudarshan, Gupta, Puneet, Krishna, Tushar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Link Quality Aware Pathfinding for Chiplet Interconnects
by: Yen, Aaron, et al.
Published: (2026)
by: Yen, Aaron, et al.
Published: (2026)
PipeOrgan: Efficient Inter-operation Pipelining with Flexible Spatial Organization and Interconnects
by: Garg, Raveesh, et al.
Published: (2024)
by: Garg, Raveesh, et al.
Published: (2024)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
by: Raj, Ritik, et al.
Published: (2025)
by: Raj, Ritik, et al.
Published: (2025)
Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale Integration
by: Feng, Yinxiao, et al.
Published: (2024)
by: Feng, Yinxiao, et al.
Published: (2024)
Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
Efficient VQ-QAT and Mixed Vector/Linear quantized Neural Networks
by: Gou, Terry, et al.
Published: (2026)
by: Gou, Terry, et al.
Published: (2026)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
by: Yu, Zhewen, et al.
Published: (2024)
by: Yu, Zhewen, et al.
Published: (2024)
Performance Modeling and Workload Analysis of Distributed Large Language Model Training and Inference
by: Kundu, Joyjit, et al.
Published: (2024)
by: Kundu, Joyjit, et al.
Published: (2024)
Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding
by: Iff, Patrick, et al.
Published: (2026)
by: Iff, Patrick, et al.
Published: (2026)
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
by: Won, William, et al.
Published: (2021)
by: Won, William, et al.
Published: (2021)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
by: Wang, Zhao, et al.
Published: (2021)
by: Wang, Zhao, et al.
Published: (2021)
Performance Analysis of DNN Inference/Training with Convolution and non-Convolution Operations
by: Esmaeilzadeh, Hadi, et al.
Published: (2023)
by: Esmaeilzadeh, Hadi, et al.
Published: (2023)
Exploration of Activation Fault Reliability in Quantized Systolic Array-Based DNN Accelerators
by: Taheri, Mahdi, et al.
Published: (2024)
by: Taheri, Mahdi, et al.
Published: (2024)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
DAISM: Digital Approximate In-SRAM Multiplier-based Accelerator for DNN Training and Inference
by: Sonnino, Lorenzo, et al.
Published: (2023)
by: Sonnino, Lorenzo, et al.
Published: (2023)
HaShiFlex: A High-Throughput Hardened Shifter DNN Accelerator with Fine-Tuning Flexibility
by: Herbst, Jonathan, et al.
Published: (2025)
by: Herbst, Jonathan, et al.
Published: (2025)
CIMPool: Scalable Neural Network Acceleration for Compute-In-Memory using Weight Pools
by: Li, Shurui, et al.
Published: (2025)
by: Li, Shurui, et al.
Published: (2025)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
by: Leon, Vasileios, et al.
Published: (2025)
by: Leon, Vasileios, et al.
Published: (2025)
WAGONN: Weight Bit Agglomeration in Crossbar Arrays for Reduced Impact of Interconnect Resistance on DNN Inference Accuracy
by: Victor, Jeffry, et al.
Published: (2024)
by: Victor, Jeffry, et al.
Published: (2024)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
by: Luo, Shuqing, et al.
Published: (2026)
by: Luo, Shuqing, et al.
Published: (2026)
Noise-Resilient Quantum Aggregation on NISQ for Federated ADAS Learning
by: Kabgere, Chethana Prasad, et al.
Published: (2025)
by: Kabgere, Chethana Prasad, et al.
Published: (2025)
DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators
by: Hong, Charles, et al.
Published: (2025)
by: Hong, Charles, et al.
Published: (2025)
Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
by: Zhu, Jingchen, et al.
Published: (2024)
by: Zhu, Jingchen, et al.
Published: (2024)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
FILCO: Flexible Composing Architecture with Real-Time Reconfigurability for DNN Acceleration
by: Chen, Xingzhen, et al.
Published: (2026)
by: Chen, Xingzhen, et al.
Published: (2026)
Mirage: An RNS-Based Photonic Accelerator for DNN Training
by: Demirkiran, Cansu, et al.
Published: (2023)
by: Demirkiran, Cansu, et al.
Published: (2023)
Leveraging Highly Approximated Multipliers in DNN Inference
by: Zervakis, Georgios, et al.
Published: (2024)
by: Zervakis, Georgios, et al.
Published: (2024)
ARCS: Autoregressive Circuit Synthesis with Topology-Aware Graph Attention and Spec Conditioning
by: Pathak, Tushar Dhananjay
Published: (2026)
by: Pathak, Tushar Dhananjay
Published: (2026)
YAP+: Pad-Layout-Aware Yield Modeling and Simulation for Hybrid Bonding
by: Chen, Zhichao, et al.
Published: (2025)
by: Chen, Zhichao, et al.
Published: (2025)
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
by: Cherezova, Natalia, et al.
Published: (2025)
by: Cherezova, Natalia, et al.
Published: (2025)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
H3DFact: Heterogeneous 3D Integrated CIM for Factorization with Holographic Perceptual Representations
by: Wan, Zishen, et al.
Published: (2024)
by: Wan, Zishen, et al.
Published: (2024)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
by: Zheng, Size, et al.
Published: (2024)
by: Zheng, Size, et al.
Published: (2024)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
by: Tong, Jianming, et al.
Published: (2024)
by: Tong, Jianming, et al.
Published: (2024)
AxMoE: Characterizing the Impact of Approximate Multipliers on Mixture-of-Experts DNN Architectures
by: Shende, Omkar B, et al.
Published: (2026)
by: Shende, Omkar B, et al.
Published: (2026)
NeFT: Negative Feedback Training to Improve Robustness of Compute-In-Memory DNN Accelerators
by: Qin, Yifan, et al.
Published: (2023)
by: Qin, Yifan, et al.
Published: (2023)
Similar Items
-
Link Quality Aware Pathfinding for Chiplet Interconnects
by: Yen, Aaron, et al.
Published: (2026) -
PipeOrgan: Efficient Inter-operation Pipelining with Flexible Spatial Organization and Interconnects
by: Garg, Raveesh, et al.
Published: (2024) -
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
by: Raj, Ritik, et al.
Published: (2025) -
Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale Integration
by: Feng, Yinxiao, et al.
Published: (2024) -
Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
by: Ramachandran, Akshat, et al.
Published: (2024)