The Case for Co-Designing Model Architectures with Hardware
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Anthony, Quentin, Hatef, Jacob, Narayanan, Deepak, Biderman, Stella, Bekman, Stas, Yin, Junqi, Shafi, Aamir, Subramoni, Hari, Panda, Dhabaleswar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
von: Xu, Lang, et al.
Veröffentlicht: (2025)
von: Xu, Lang, et al.
Veröffentlicht: (2025)
Demystifying the Communication Characteristics for Distributed Transformer Models
von: Anthony, Quentin, et al.
Veröffentlicht: (2024)
von: Anthony, Quentin, et al.
Veröffentlicht: (2024)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
von: Xu, Lang, et al.
Veröffentlicht: (2024)
von: Xu, Lang, et al.
Veröffentlicht: (2024)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)
Characterizing Communication Patterns in Distributed Large Language Model Inference
von: Xu, Lang, et al.
Veröffentlicht: (2025)
von: Xu, Lang, et al.
Veröffentlicht: (2025)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
Comparative Study of Large Language Model Architectures on Frontier
von: Yin, Junqi, et al.
Veröffentlicht: (2024)
von: Yin, Junqi, et al.
Veröffentlicht: (2024)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
von: Lian, Xinyu, et al.
Veröffentlicht: (2024)
von: Lian, Xinyu, et al.
Veröffentlicht: (2024)
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving
von: Gao, Shouwei, et al.
Veröffentlicht: (2026)
von: Gao, Shouwei, et al.
Veröffentlicht: (2026)
Benchmarking Compound AI Applications for Hardware-Software Co-Design
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026)
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026)
Comparing the Run-time Behavior of Modern PDES Engines on Alternative Hardware Architectures
von: Marotta, Romolo, et al.
Veröffentlicht: (2025)
von: Marotta, Romolo, et al.
Veröffentlicht: (2025)
MAC-Attention: a Match-Amend-Complete Scheme for Fast and Accurate Attention Computation
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)
Towards Optimal Deterministic LOCAL Algorithms on Trees
von: Brandt, Sebastian, et al.
Veröffentlicht: (2025)
von: Brandt, Sebastian, et al.
Veröffentlicht: (2025)
Hybrid Cloud Architectures for Research Computing: Applications and Use Cases
von: Stiensmeier, Xaver, et al.
Veröffentlicht: (2026)
von: Stiensmeier, Xaver, et al.
Veröffentlicht: (2026)
IntentContinuum: Using LLMs to Support Intent-Based Computing Across the Compute Continuum
von: Akbari, Negin, et al.
Veröffentlicht: (2025)
von: Akbari, Negin, et al.
Veröffentlicht: (2025)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
von: Wu, Siyu, et al.
Veröffentlicht: (2025)
von: Wu, Siyu, et al.
Veröffentlicht: (2025)
KernelFoundry: Hardware-aware evolutionary GPU kernel optimization
von: Wiedemann, Nina, et al.
Veröffentlicht: (2026)
von: Wiedemann, Nina, et al.
Veröffentlicht: (2026)
The Case for ABI Interoperability in a Fault Tolerant MPI
von: Xu, Yao, et al.
Veröffentlicht: (2025)
von: Xu, Yao, et al.
Veröffentlicht: (2025)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
ALBERTA: ALgorithm-Based Error Resilience in Transformer Architectures
von: Liu, Haoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Haoxuan, et al.
Veröffentlicht: (2023)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
von: Shyam, Vasu, et al.
Veröffentlicht: (2026)
von: Shyam, Vasu, et al.
Veröffentlicht: (2026)
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
von: Wang, Hansheng, et al.
Veröffentlicht: (2024)
von: Wang, Hansheng, et al.
Veröffentlicht: (2024)
DRackSim: Simulator for Rack-scale Memory Disaggregation
von: Puri, Amit, et al.
Veröffentlicht: (2023)
von: Puri, Amit, et al.
Veröffentlicht: (2023)
Distributed Resource Selection for Self-Organising Cloud-Edge Systems
von: Renau, Quentin, et al.
Veröffentlicht: (2025)
von: Renau, Quentin, et al.
Veröffentlicht: (2025)
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
von: Shubham, et al.
Veröffentlicht: (2024)
von: Shubham, et al.
Veröffentlicht: (2024)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
von: Patke, Archit, et al.
Veröffentlicht: (2025)
von: Patke, Archit, et al.
Veröffentlicht: (2025)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026)
von: Liang, Yan, et al.
Veröffentlicht: (2026)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
HMTRace: Hardware-Assisted Memory-Tagging based Dynamic Data Race Detection
von: Shastri, Jaidev, et al.
Veröffentlicht: (2024)
von: Shastri, Jaidev, et al.
Veröffentlicht: (2024)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
Modernizing an Operational Real-time Tsunami Simulator to Support Diverse Hardware Platforms
von: Takahashi, Keichi, et al.
Veröffentlicht: (2024)
von: Takahashi, Keichi, et al.
Veröffentlicht: (2024)
Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
von: Reisecker, Maximilian, et al.
Veröffentlicht: (2025)
von: Reisecker, Maximilian, et al.
Veröffentlicht: (2025)
An Efficient Approach for Energy Conservation in Cloud Computing Environment
von: Pande, Sohan Kumar, et al.
Veröffentlicht: (2025)
von: Pande, Sohan Kumar, et al.
Veröffentlicht: (2025)
Hardware-Aware Reformulation of Convolutions for Efficient Execution on Specialized AI Hardware: A Case Study on NVIDIA Tensor Cores
von: Bikshandi, Ganesh
Veröffentlicht: (2026)
von: Bikshandi, Ganesh
Veröffentlicht: (2026)
Bridging Simulation and Silicon: A Study of RISC-V Hardware and FireSim Simulation
von: Barai, Atanu, et al.
Veröffentlicht: (2025)
von: Barai, Atanu, et al.
Veröffentlicht: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
von: Xu, Lang, et al.
Veröffentlicht: (2025) -
Demystifying the Communication Characteristics for Distributed Transformer Models
von: Anthony, Quentin, et al.
Veröffentlicht: (2024) -
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
von: Yao, Jinghan, et al.
Veröffentlicht: (2024) -
Accelerating Large Language Model Training with Hybrid GPU-based Compression
von: Xu, Lang, et al.
Veröffentlicht: (2024) -
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)