UpDown: Programmable fine-grained Events for Scalable Performance on Irregular Applications
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913452502548480 |
|---|---|
| author | Rajasukumar, Andronicus Su, Jiya Yuqing Wang Su, Tianshuo Nourian, Marziyeh Diaz, Jose M Monsalve Zhang, Tianchi Ding, Jianru Wang, Wenyi Zhang, Ziyi Jeje, Moubarak Hoffmann, Henry Li, Yanjing Chien, Andrew A. |
| author_facet | Rajasukumar, Andronicus Su, Jiya Yuqing Wang Su, Tianshuo Nourian, Marziyeh Diaz, Jose M Monsalve Zhang, Tianchi Ding, Jianru Wang, Wenyi Zhang, Ziyi Jeje, Moubarak Hoffmann, Henry Li, Yanjing Chien, Andrew A. |
| contents | Applications with irregular data structures, data-dependent control flows and fine-grained data transfers (e.g., real-world graph computations) perform poorly on cache-based systems. We propose the UpDown accelerator that supports fine-grained execution with novel architecture mechanisms - lightweight threading, event-driven scheduling, efficient ultra-short threads, and split-transaction DRAM access with software-controlled synchronization. These hardware primitives support software programmable events, enabling high performance on diverse data structures and algorithms. UpDown also supports scalable performance; hardware replication enables programs to scale up performance. Evaluation results show UpDown's flexibility and scalability enable it to outperform CPUs on graph mining and analytics computations by up to 116-195x geomean speedup and more than 4x speedup over prior accelerators. We show that UpDown generates high memory parallelism (~4.6x over CPU) required for memory intensive graph computations. We present measurements that attribute the performance of UpDown (23x architectural advantage) to its individual architectural mechanisms. Finally, we also analyze the area and power cost of UpDown's mechanisms for software programmability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_20773 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | UpDown: Programmable fine-grained Events for Scalable Performance on Irregular Applications Rajasukumar, Andronicus Su, Jiya Yuqing Wang Su, Tianshuo Nourian, Marziyeh Diaz, Jose M Monsalve Zhang, Tianchi Ding, Jianru Wang, Wenyi Zhang, Ziyi Jeje, Moubarak Hoffmann, Henry Li, Yanjing Chien, Andrew A. Hardware Architecture Applications with irregular data structures, data-dependent control flows and fine-grained data transfers (e.g., real-world graph computations) perform poorly on cache-based systems. We propose the UpDown accelerator that supports fine-grained execution with novel architecture mechanisms - lightweight threading, event-driven scheduling, efficient ultra-short threads, and split-transaction DRAM access with software-controlled synchronization. These hardware primitives support software programmable events, enabling high performance on diverse data structures and algorithms. UpDown also supports scalable performance; hardware replication enables programs to scale up performance. Evaluation results show UpDown's flexibility and scalability enable it to outperform CPUs on graph mining and analytics computations by up to 116-195x geomean speedup and more than 4x speedup over prior accelerators. We show that UpDown generates high memory parallelism (~4.6x over CPU) required for memory intensive graph computations. We present measurements that attribute the performance of UpDown (23x architectural advantage) to its individual architectural mechanisms. Finally, we also analyze the area and power cost of UpDown's mechanisms for software programmability. |
| title | UpDown: Programmable fine-grained Events for Scalable Performance on Irregular Applications |
| topic | Hardware Architecture |
| url | https://arxiv.org/abs/2407.20773 |