FlexLink: Boosting your NVLink Bandwidth by 27% without accuracy concern
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shen, Ao, Zhang, Rui, Zhao, Junping |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Flex-TPU: A Flexible TPU with Runtime Reconfigurable Dataflow Architecture
par: Elbtity, Mohammed, et autres
Publié: (2024)
par: Elbtity, Mohammed, et autres
Publié: (2024)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
par: Zhang, Yichao, et autres
Publié: (2026)
par: Zhang, Yichao, et autres
Publié: (2026)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
par: Shen, Diyou, et autres
Publié: (2025)
par: Shen, Diyou, et autres
Publié: (2025)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
par: Liu, Fangxin, et autres
Publié: (2026)
par: Liu, Fangxin, et autres
Publié: (2026)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
par: Zhu, Wenbin, et autres
Publié: (2025)
par: Zhu, Wenbin, et autres
Publié: (2025)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
par: Zhao, Yiren, et autres
Publié: (2026)
par: Zhao, Yiren, et autres
Publié: (2026)
Advancing AI-assisted Hardware Design with Hierarchical Decentralized Training and Personalized Inference-Time Optimization
par: Chen, Hao Mark, et autres
Publié: (2025)
par: Chen, Hao Mark, et autres
Publié: (2025)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
par: Li, Jonathan, et autres
Publié: (2025)
par: Li, Jonathan, et autres
Publié: (2025)
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
par: Wang, Tinglue, et autres
Publié: (2025)
par: Wang, Tinglue, et autres
Publié: (2025)
The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
par: Wang, Zixuan, et autres
Publié: (2024)
par: Wang, Zixuan, et autres
Publié: (2024)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
par: Li, Bohan, et autres
Publié: (2026)
par: Li, Bohan, et autres
Publié: (2026)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
par: Zhao, Dan, et autres
Publié: (2024)
par: Zhao, Dan, et autres
Publié: (2024)
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
par: Zhou, Zhongchun, et autres
Publié: (2025)
par: Zhou, Zhongchun, et autres
Publié: (2025)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
par: Stojkovic, Jovan, et autres
Publié: (2025)
par: Stojkovic, Jovan, et autres
Publié: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
par: Pan, Yudong, et autres
Publié: (2026)
par: Pan, Yudong, et autres
Publié: (2026)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
par: Stojkovic, Jovan, et autres
Publié: (2024)
par: Stojkovic, Jovan, et autres
Publié: (2024)
NPU Design for Diffusion Language Model Inference
par: Lou, Binglei, et autres
Publié: (2026)
par: Lou, Binglei, et autres
Publié: (2026)
HPU: High-Bandwidth Processing Unit for Scalable, Cost-effective LLM Inference via GPU Co-processing
par: Rhee, Myunghyun, et autres
Publié: (2025)
par: Rhee, Myunghyun, et autres
Publié: (2025)
Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
par: Bergman, Shai, et autres
Publié: (2025)
par: Bergman, Shai, et autres
Publié: (2025)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
par: Qin, Ruoyu, et autres
Publié: (2024)
par: Qin, Ruoyu, et autres
Publié: (2024)
Investigating Memory Failure Prediction Across CPU Architectures
par: Yu, Qiao, et autres
Publié: (2024)
par: Yu, Qiao, et autres
Publié: (2024)
CUTEv2: Unified and Configurable Matrix Extension for Diverse CPU Architectures with Minimal Design Overhead
par: Ye, Jinpeng, et autres
Publié: (2026)
par: Ye, Jinpeng, et autres
Publié: (2026)
PIM-Opt: Demystifying Distributed Optimization Algorithms on a Real-World Processing-In-Memory System
par: Rhyner, Steve, et autres
Publié: (2024)
par: Rhyner, Steve, et autres
Publié: (2024)
Serving Large Language Models on Huawei CloudMatrix384
par: Zuo, Pengfei, et autres
Publié: (2025)
par: Zuo, Pengfei, et autres
Publié: (2025)
Power Stabilization for AI Training Datacenters
par: Choukse, Esha, et autres
Publié: (2025)
par: Choukse, Esha, et autres
Publié: (2025)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
par: DeBole, Michael V., et autres
Publié: (2025)
par: DeBole, Michael V., et autres
Publié: (2025)
PiKV: KV Cache Management System for Mixture of Experts
par: Liu, Dong, et autres
Publié: (2025)
par: Liu, Dong, et autres
Publié: (2025)
Modernizing Amdahl's Law: How AI Scaling Laws Shape Computer Architecture
par: Lu, Chien-Ping
Publié: (2026)
par: Lu, Chien-Ping
Publié: (2026)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
par: Chen, Jiesong, et autres
Publié: (2026)
par: Chen, Jiesong, et autres
Publié: (2026)
Co-design of a novel CMOS highly parallel, low-power, multi-chip neural network accelerator
par: Hokenmaier, W, et autres
Publié: (2024)
par: Hokenmaier, W, et autres
Publié: (2024)
Exploring energy consumption of AI frameworks on a 64-core RV64 Server CPU
par: Malenza, Giulio, et autres
Publié: (2025)
par: Malenza, Giulio, et autres
Publié: (2025)
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
par: Canakci, Burcu, et autres
Publié: (2025)
par: Canakci, Burcu, et autres
Publié: (2025)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
par: Yüzügüler, Ahmet Caner, et autres
Publié: (2025)
par: Yüzügüler, Ahmet Caner, et autres
Publié: (2025)
Strict Partitioning for Sporadic Rigid Gang Tasks
par: Sun, Binqi, et autres
Publié: (2024)
par: Sun, Binqi, et autres
Publié: (2024)
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
par: Dorairaj, Nij, et autres
Publié: (2026)
par: Dorairaj, Nij, et autres
Publié: (2026)
Debunking the CUDA Myth Towards GPU-based AI Systems
par: Lee, Yunjae, et autres
Publié: (2024)
par: Lee, Yunjae, et autres
Publié: (2024)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
par: Kubwimana, Benjamin, et autres
Publié: (2025)
par: Kubwimana, Benjamin, et autres
Publié: (2025)
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
par: Peccia, Federico Nicolas, et autres
Publié: (2024)
par: Peccia, Federico Nicolas, et autres
Publié: (2024)
Improving AI Efficiency in Data Centres by Power Dynamic Response
par: Marinoni, Andrea, et autres
Publié: (2025)
par: Marinoni, Andrea, et autres
Publié: (2025)
Forge-UGC: FX optimization and register-graph engine for universal graph compiler
par: Kumar, Satyam, et autres
Publié: (2026)
par: Kumar, Satyam, et autres
Publié: (2026)
Documents similaires
-
Flex-TPU: A Flexible TPU with Runtime Reconfigurable Dataflow Architecture
par: Elbtity, Mohammed, et autres
Publié: (2024) -
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
par: Zhang, Yichao, et autres
Publié: (2026) -
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
par: Shen, Diyou, et autres
Publié: (2025) -
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
par: Liu, Fangxin, et autres
Publié: (2026) -
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
par: Zhu, Wenbin, et autres
Publié: (2025)