N-TORC: Native Tensor Optimizer for Real-time Constraints
Fuente:
arXiv
Guardado en:
| Autores principales: | Singh, Suyash Vardhan, Ahmad, Iftakhar, Andrews, David, Huang, Miaoqing, Downey, Austin R. J., Bakos, Jason D. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
por: Kabir, Ehsan, et al.
Publicado: (2024)
por: Kabir, Ehsan, et al.
Publicado: (2024)
A Runtime-Adaptive Transformer Neural Network Accelerator on FPGAs
por: Kabir, Ehsan, et al.
Publicado: (2024)
por: Kabir, Ehsan, et al.
Publicado: (2024)
ProTEA: Programmable Transformer Encoder Acceleration on FPGA
por: Kabir, Ehsan, et al.
Publicado: (2024)
por: Kabir, Ehsan, et al.
Publicado: (2024)
IMAGine: An In-Memory Accelerated GEMV Engine Overlay
por: Kabir, MD Arafat, et al.
Publicado: (2024)
por: Kabir, MD Arafat, et al.
Publicado: (2024)
The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM Accelerators
por: Kabir, MD Arafat, et al.
Publicado: (2024)
por: Kabir, MD Arafat, et al.
Publicado: (2024)
Switchable Single/Dual Edge Registers for Pipeline Architecture
por: Singh, Suyash Vardhan, et al.
Publicado: (2024)
por: Singh, Suyash Vardhan, et al.
Publicado: (2024)
FADiff: Fusion-Aware Differentiable Optimization for DNN Scheduling on Tensor Accelerators
por: Jia, Shuao, et al.
Publicado: (2025)
por: Jia, Shuao, et al.
Publicado: (2025)
CrossNAS: A Cross-Layer Neural Architecture Search Framework for PIM Systems
por: Amin, Md Hasibul, et al.
Publicado: (2025)
por: Amin, Md Hasibul, et al.
Publicado: (2025)
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
por: Wu, Qizhe, et al.
Publicado: (2024)
por: Wu, Qizhe, et al.
Publicado: (2024)
TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
por: Bertuletti, Marco, et al.
Publicado: (2026)
por: Bertuletti, Marco, et al.
Publicado: (2026)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
por: Ye, Hanchen, et al.
Publicado: (2025)
por: Ye, Hanchen, et al.
Publicado: (2025)
Big-PERCIVAL: Exploring the Native Use of 64-Bit Posit Arithmetic in Scientific Computing
por: Mallasén, David, et al.
Publicado: (2023)
por: Mallasén, David, et al.
Publicado: (2023)
SKYLIGHT: A Scalable Hundred-Channel 3D Photonic In-Memory Tensor Core Architecture for Real-time AI Inference
por: Zhang, Meng, et al.
Publicado: (2026)
por: Zhang, Meng, et al.
Publicado: (2026)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
por: Huang, Sixiao, et al.
Publicado: (2025)
por: Huang, Sixiao, et al.
Publicado: (2025)
Holistic Optimization Framework for FPGA Accelerators
por: Pouget, Stéphane, et al.
Publicado: (2025)
por: Pouget, Stéphane, et al.
Publicado: (2025)
GTA: a new General Tensor Accelerator with Better Area Efficiency and Data Reuse
por: Ai, Chenyang, et al.
Publicado: (2024)
por: Ai, Chenyang, et al.
Publicado: (2024)
Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
por: Zhou, Weiyu, et al.
Publicado: (2025)
por: Zhou, Weiyu, et al.
Publicado: (2025)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
por: Taka, Endri, et al.
Publicado: (2025)
por: Taka, Endri, et al.
Publicado: (2025)
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
por: Zheng, Keran, et al.
Publicado: (2025)
por: Zheng, Keran, et al.
Publicado: (2025)
Accelerating Detailed Routing Convergence through Offline Reinforcement Learning
por: Khan, Afsara, et al.
Publicado: (2025)
por: Khan, Afsara, et al.
Publicado: (2025)
Error Checking for Sparse Systolic Tensor Arrays
por: Peltekis, Christodoulos, et al.
Publicado: (2024)
por: Peltekis, Christodoulos, et al.
Publicado: (2024)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
por: Zhou, Zhe, et al.
Publicado: (2024)
por: Zhou, Zhe, et al.
Publicado: (2024)
Open-source Stand-Alone Versatile Tensor Accelerator
por: Faure-Gignoux, Anthony, et al.
Publicado: (2025)
por: Faure-Gignoux, Anthony, et al.
Publicado: (2025)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
por: Shin, Yongwon, et al.
Publicado: (2024)
por: Shin, Yongwon, et al.
Publicado: (2024)
ATLAAS: Automatic Tensor-Level Abstraction of Accelerator Semantics
por: Gao, Ruijie, et al.
Publicado: (2026)
por: Gao, Ruijie, et al.
Publicado: (2026)
Real Time FPGA Based Transformers & VLMs for Vision Tasks: SOTA Designs and Optimizations
por: Sali, Safa Mohammed, et al.
Publicado: (2025)
por: Sali, Safa Mohammed, et al.
Publicado: (2025)
Linear Complexity Fermionic Simulation on Quantum Devices with Hardware Connectivity Constraints
por: Gao, Xiangyu, et al.
Publicado: (2026)
por: Gao, Xiangyu, et al.
Publicado: (2026)
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
por: Guan, Yue, et al.
Publicado: (2026)
por: Guan, Yue, et al.
Publicado: (2026)
Tailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Capacity
por: Xue, Zi Yu, et al.
Publicado: (2023)
por: Xue, Zi Yu, et al.
Publicado: (2023)
Tensor Memory Engine: On-the-fly Data Reorganization for Ideal Locality
por: Hoornaert, Denis, et al.
Publicado: (2026)
por: Hoornaert, Denis, et al.
Publicado: (2026)
AME-PIM: Can Memory be Your Next Tensor Accelerator?
por: Venieri, Emanuele, et al.
Publicado: (2026)
por: Venieri, Emanuele, et al.
Publicado: (2026)
PHAROS: Pipelined Heterogeneous Accelerators for Real-time Safety-critical Systems With Deadline Compliance
por: Ji, Shixin, et al.
Publicado: (2026)
por: Ji, Shixin, et al.
Publicado: (2026)
No Redundancy, No Stall: Lightweight Streaming 3D Gaussian Splatting for Real-time Rendering
por: Wei, Linye, et al.
Publicado: (2025)
por: Wei, Linye, et al.
Publicado: (2025)
Device-Level Optimization Techniques for Solid-State Drives: A Survey
por: Ren, Tianyu, et al.
Publicado: (2025)
por: Ren, Tianyu, et al.
Publicado: (2025)
Real-time Object Detection and Associated Hardware Accelerators Targeting Autonomous Vehicles: A Review
por: Sali, Safa, et al.
Publicado: (2025)
por: Sali, Safa, et al.
Publicado: (2025)
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
por: Lu, Jinming, et al.
Publicado: (2025)
por: Lu, Jinming, et al.
Publicado: (2025)
TeAAL: A Declarative Framework for Modeling Sparse Tensor Accelerators
por: Nayak, Nandeeka, et al.
Publicado: (2023)
por: Nayak, Nandeeka, et al.
Publicado: (2023)
Accelerating Sparse Graph Neural Networks with Tensor Core Optimization
por: Wu, Ka Wai
Publicado: (2024)
por: Wu, Ka Wai
Publicado: (2024)
Real Time FPGA Based CNNs for Detection, Classification, and Tracking in Autonomous Systems: State of the Art Designs and Optimizations
por: Sali, Safa Mohammed, et al.
Publicado: (2025)
por: Sali, Safa Mohammed, et al.
Publicado: (2025)
PyPIM: Integrating Digital Processing-in-Memory from Microarchitectural Design to Python Tensors
por: Leitersdorf, Orian, et al.
Publicado: (2023)
por: Leitersdorf, Orian, et al.
Publicado: (2023)
Ejemplares similares
-
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
por: Kabir, Ehsan, et al.
Publicado: (2024) -
A Runtime-Adaptive Transformer Neural Network Accelerator on FPGAs
por: Kabir, Ehsan, et al.
Publicado: (2024) -
ProTEA: Programmable Transformer Encoder Acceleration on FPGA
por: Kabir, Ehsan, et al.
Publicado: (2024) -
IMAGine: An In-Memory Accelerated GEMV Engine Overlay
por: Kabir, MD Arafat, et al.
Publicado: (2024) -
The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM Accelerators
por: Kabir, MD Arafat, et al.
Publicado: (2024)