Flex-TPU: A Flexible TPU with Runtime Reconfigurable Dataflow Architecture
Fuente:
arXiv
Guardado en:
| Autores principales: | Elbtity, Mohammed, Chandarana, Peyton, Zand, Ramtin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
por: Shi, Tianyao, et al.
Publicado: (2025)
por: Shi, Tianyao, et al.
Publicado: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
por: Odema, Mohanad, et al.
Publicado: (2024)
por: Odema, Mohanad, et al.
Publicado: (2024)
Performance and Power: Systematic Evaluation of AI Workloads on Accelerators with CARAML
por: John, Chelsea Maria, et al.
Publicado: (2024)
por: John, Chelsea Maria, et al.
Publicado: (2024)
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
por: Panigrahy, Deepak, et al.
Publicado: (2026)
por: Panigrahy, Deepak, et al.
Publicado: (2026)
DABench-LLM: Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators for LLMs
por: Hu, Ziyu, et al.
Publicado: (2025)
por: Hu, Ziyu, et al.
Publicado: (2025)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
por: Choudhary, Mansi, et al.
Publicado: (2025)
por: Choudhary, Mansi, et al.
Publicado: (2025)
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
por: Yang, Peiming, et al.
Publicado: (2025)
por: Yang, Peiming, et al.
Publicado: (2025)
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
por: Giannoula, Christina, et al.
Publicado: (2024)
por: Giannoula, Christina, et al.
Publicado: (2024)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
por: Kumar, Deepak, et al.
Publicado: (2025)
por: Kumar, Deepak, et al.
Publicado: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
por: Lei, Jianlong, et al.
Publicado: (2026)
por: Lei, Jianlong, et al.
Publicado: (2026)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
por: Li, Jiaxi, et al.
Publicado: (2025)
por: Li, Jiaxi, et al.
Publicado: (2025)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
por: Renney, Harri, et al.
Publicado: (2026)
por: Renney, Harri, et al.
Publicado: (2026)
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
por: Odema, Mohanad, et al.
Publicado: (2024)
por: Odema, Mohanad, et al.
Publicado: (2024)
RedFuser: An Automatic Operator Fusion Framework for Cascaded Reductions on AI Accelerators
por: Tang, Xinsheng, et al.
Publicado: (2026)
por: Tang, Xinsheng, et al.
Publicado: (2026)
COMET: Neural Cost Model Explanation Framework
por: Chaudhary, Isha, et al.
Publicado: (2023)
por: Chaudhary, Isha, et al.
Publicado: (2023)
FlexLink: Boosting your NVLink Bandwidth by 27% without accuracy concern
por: Shen, Ao, et al.
Publicado: (2025)
por: Shen, Ao, et al.
Publicado: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
por: Szafarczyk, Robert, et al.
Publicado: (2025)
por: Szafarczyk, Robert, et al.
Publicado: (2025)
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
por: Adamopoulos, Dionysios, et al.
Publicado: (2025)
por: Adamopoulos, Dionysios, et al.
Publicado: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
por: Fan, Ruibo, et al.
Publicado: (2026)
por: Fan, Ruibo, et al.
Publicado: (2026)
Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry
por: Chen, Ziji, et al.
Publicado: (2025)
por: Chen, Ziji, et al.
Publicado: (2025)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
por: Luo, Weile, et al.
Publicado: (2025)
por: Luo, Weile, et al.
Publicado: (2025)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
por: Chen, Qian, et al.
Publicado: (2024)
por: Chen, Qian, et al.
Publicado: (2024)
INR-Arch: A Dataflow Architecture and Compiler for Arbitrary-Order Gradient Computations in Implicit Neural Representation Processing
por: Abi-Karam, Stefan, et al.
Publicado: (2023)
por: Abi-Karam, Stefan, et al.
Publicado: (2023)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
por: Guo, Licheng, et al.
Publicado: (2022)
por: Guo, Licheng, et al.
Publicado: (2022)
ARCAS: Adaptive Runtime System for Chiplet-Aware Scheduling
por: Fogli, Alessandro, et al.
Publicado: (2025)
por: Fogli, Alessandro, et al.
Publicado: (2025)
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge
por: Grailoo, M., et al.
Publicado: (2026)
por: Grailoo, M., et al.
Publicado: (2026)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
por: Li, Jonathan, et al.
Publicado: (2025)
por: Li, Jonathan, et al.
Publicado: (2025)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
por: Zhu, Yu, et al.
Publicado: (2025)
por: Zhu, Yu, et al.
Publicado: (2025)
GigaAPI for GPU Parallelization
por: Suvarna, M., et al.
Publicado: (2025)
por: Suvarna, M., et al.
Publicado: (2025)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
por: Ibeid, Huda, et al.
Publicado: (2025)
por: Ibeid, Huda, et al.
Publicado: (2025)
UPMEM Unleashed: Software Secrets for Speed
por: Chmielewski, Krystian, et al.
Publicado: (2025)
por: Chmielewski, Krystian, et al.
Publicado: (2025)
Experimental Assessment of Containers Running on Top of Virtual Machines
por: Aqasizade, Hossein, et al.
Publicado: (2024)
por: Aqasizade, Hossein, et al.
Publicado: (2024)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
por: Chakraborty, Abhinaba, et al.
Publicado: (2025)
por: Chakraborty, Abhinaba, et al.
Publicado: (2025)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
por: Wang, Xi, et al.
Publicado: (2024)
por: Wang, Xi, et al.
Publicado: (2024)
Parallelizing a modern GPU simulator
por: Huerta, Rodrigo, et al.
Publicado: (2025)
por: Huerta, Rodrigo, et al.
Publicado: (2025)
Can Asymmetric Tile Buffering Be Beneficial?
por: Wang, Chengyue, et al.
Publicado: (2025)
por: Wang, Chengyue, et al.
Publicado: (2025)
Exploiting long vectors with a CFD code: a co-design show case
por: Blancafort, Marc, et al.
Publicado: (2024)
por: Blancafort, Marc, et al.
Publicado: (2024)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
por: Qararyah, Fareed, et al.
Publicado: (2024)
por: Qararyah, Fareed, et al.
Publicado: (2024)
Simopt -- Simulation pass for Speculative Optimisation of FPGA-CAD flow
por: Wadhwa, Eashan, et al.
Publicado: (2024)
por: Wadhwa, Eashan, et al.
Publicado: (2024)
Ejemplares similares
-
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
por: Shi, Tianyao, et al.
Publicado: (2025) -
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
por: Vellaisamy, Prabhu, et al.
Publicado: (2025) -
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
por: Odema, Mohanad, et al.
Publicado: (2024) -
Performance and Power: Systematic Evaluation of AI Workloads on Accelerators with CARAML
por: John, Chelsea Maria, et al.
Publicado: (2024) -
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
por: Panigrahy, Deepak, et al.
Publicado: (2026)