Enabling full-speed random access to the entire memory on the A100 GPU
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866929348583358464 |
|---|---|
| author | Walker, Alden |
| author_facet | Walker, Alden |
| contents | We describe some features of the A100 memory architecture. In particular, we give a technique to reverse-engineer some hardware layout information. Using this information, we show how to avoid TLB issues to obtain full-speed random HBM access to the entire memory, as long as we constrain any particular thread to a reduced access window of less than 64GB. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_11425 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Enabling full-speed random access to the entire memory on the A100 GPU Walker, Alden Performance Hardware Architecture C.4; B.3.3 We describe some features of the A100 memory architecture. In particular, we give a technique to reverse-engineer some hardware layout information. Using this information, we show how to avoid TLB issues to obtain full-speed random HBM access to the entire memory, as long as we constrain any particular thread to a reduced access window of less than 64GB. |
| title | Enabling full-speed random access to the entire memory on the A100 GPU |
| topic | Performance Hardware Architecture C.4; B.3.3 |
| url | https://arxiv.org/abs/2405.11425 |