A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine
Fuente:
arXiv
Saved in:
| Main Authors: | Sapkas, M., Triossi, A., Zanetti, M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge
by: Grailoo, M., et al.
Published: (2026)
by: Grailoo, M., et al.
Published: (2026)
PPU: Design and Implementation of a Pipelined Full Posit Processing Unit
by: Rossi, Federico, et al.
Published: (2023)
by: Rossi, Federico, et al.
Published: (2023)
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
by: Shen, Siyuan, et al.
Published: (2025)
by: Shen, Siyuan, et al.
Published: (2025)
Latency and Privacy-Aware Resource Allocation in Vehicular Edge Computing
by: Ahmadvand, Hossein, et al.
Published: (2025)
by: Ahmadvand, Hossein, et al.
Published: (2025)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
by: Zhou, Fang, et al.
Published: (2026)
by: Zhou, Fang, et al.
Published: (2026)
PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
by: Le, Truong-Thanh, et al.
Published: (2026)
by: Le, Truong-Thanh, et al.
Published: (2026)
CXL and the Return of Scale-Up Database Engines
by: Lerner, Alberto, et al.
Published: (2024)
by: Lerner, Alberto, et al.
Published: (2024)
Motion-to-Motion Latency Measurement Framework for Connected and Autonomous Vehicle Teleoperation
by: Provost, François, et al.
Published: (2025)
by: Provost, François, et al.
Published: (2025)
An Experimental Study of Low-Latency Video Streaming over 5G
by: Khan, Imran, et al.
Published: (2024)
by: Khan, Imran, et al.
Published: (2024)
DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs
by: Chen, Mingkai, et al.
Published: (2024)
by: Chen, Mingkai, et al.
Published: (2024)
CARINA: Carbon-Aware Execution of Recurrent Industrial Analytics
by: Farooq, Muhammad Umar
Published: (2026)
by: Farooq, Muhammad Umar
Published: (2026)
Latency Based Tiling
by: Cashman, Jack
Published: (2025)
by: Cashman, Jack
Published: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
by: Kong, Linghao, et al.
Published: (2026)
by: Kong, Linghao, et al.
Published: (2026)
Dynamic Precision Math Engine for Linear Algebra and Trigonometry Acceleration on Xtensa LX6 Microcontrollers
by: Preciado, Elian Alfonso Lopez
Published: (2026)
by: Preciado, Elian Alfonso Lopez
Published: (2026)
Uncertainty Quantification as a Complementary Latent Health Indicator for Remaining Useful Life Prediction on Turbofan Engines
by: Thil, Lucas, et al.
Published: (2025)
by: Thil, Lucas, et al.
Published: (2025)
The Multiserver-Job Stochastic Recurrence Equation for Cloud Computing Performance Evaluation
by: Baccelli, Francois, et al.
Published: (2026)
by: Baccelli, Francois, et al.
Published: (2026)
LoPace: A Lossless Optimized Prompt Accurate Compression Engine for Large Language Model Applications
by: Ulla, Aman
Published: (2026)
by: Ulla, Aman
Published: (2026)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
by: Papavasileiou, Ioannis, et al.
Published: (2026)
by: Papavasileiou, Ioannis, et al.
Published: (2026)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
by: Wang, Haoxin, et al.
Published: (2025)
by: Wang, Haoxin, et al.
Published: (2025)
Analysis and Evaluation of Using Microsecond-Latency Memory for In-Memory Indices and Caches in SSD-Based Key-Value Stores
by: Bando, Yosuke, et al.
Published: (2025)
by: Bando, Yosuke, et al.
Published: (2025)
Characterizing Machine Learning Force Fields as Emerging Molecular Dynamics Workloads on Graphics Processing Units
by: De Alwis, Udari, et al.
Published: (2026)
by: De Alwis, Udari, et al.
Published: (2026)
Unikernels vs. Containers: A Runtime-Level Performance Comparison for Resource-Constrained Edge Workloads
by: Dinh-Tuan, Hai
Published: (2025)
by: Dinh-Tuan, Hai
Published: (2025)
Light Differentiable Logic Gate Networks
by: Rüttgers, Lukas, et al.
Published: (2025)
by: Rüttgers, Lukas, et al.
Published: (2025)
Recurrent CircuitSAT Sampling for Sequential Circuits
by: Ardakani, Arash, et al.
Published: (2025)
by: Ardakani, Arash, et al.
Published: (2025)
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
by: Qararyah, Fareed, et al.
Published: (2025)
by: Qararyah, Fareed, et al.
Published: (2025)
Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
by: Almurshed, Osama, et al.
Published: (2025)
by: Almurshed, Osama, et al.
Published: (2025)
Implementing and Evaluating E2LSH on Storage
by: Nakanishi, Yu, et al.
Published: (2024)
by: Nakanishi, Yu, et al.
Published: (2024)
Rethinking Temporal Models for TinyML: LSTM versus 1D-CNN in Resource-Constrained Devices
by: Saha, Bidyut, et al.
Published: (2026)
by: Saha, Bidyut, et al.
Published: (2026)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
Mind the Gap: Removing the Discretization Gap in Differentiable Logic Gate Networks
by: Yousefi, Shakir, et al.
Published: (2025)
by: Yousefi, Shakir, et al.
Published: (2025)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
by: Jeffery, Andrew, et al.
Published: (2024)
by: Jeffery, Andrew, et al.
Published: (2024)
SysOM-AI: Continuous Cross-Layer Performance Diagnosis for Production AI Training
by: Zheng, Yusheng, et al.
Published: (2026)
by: Zheng, Yusheng, et al.
Published: (2026)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
by: Chu, Kexin, et al.
Published: (2026)
by: Chu, Kexin, et al.
Published: (2026)
Impact of AI-Triage on Radiologist Report Turnaround Time: Real-World Time-Savings and Insights from Model Predictions
by: Thompson, Yee Lam Elim, et al.
Published: (2025)
by: Thompson, Yee Lam Elim, et al.
Published: (2025)
Parallel Implementations Assessment of a Spatial-Spectral Classifier for Hyperspectral Clinical Applications
by: Lazcano, Raquel, et al.
Published: (2024)
by: Lazcano, Raquel, et al.
Published: (2024)
A Study on Inference Latency for Vision Transformers on Mobile Devices
by: Li, Zhuojin, et al.
Published: (2025)
by: Li, Zhuojin, et al.
Published: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025)
by: Chen, Feiyang, et al.
Published: (2025)
A Comparative Study and Implementation of Key Derivation Functions Standardized by NIST and IEEE
by: Chen, Abel C. H.
Published: (2025)
by: Chen, Abel C. H.
Published: (2025)
WCDT: Systematic WCET Optimization for Decision Tree Implementations
by: Hölscher, Nils, et al.
Published: (2025)
by: Hölscher, Nils, et al.
Published: (2025)
RWKV-edge: Deeply Compressed RWKV for Resource-Constrained Devices
by: Choe, Wonkyo, et al.
Published: (2024)
by: Choe, Wonkyo, et al.
Published: (2024)
Similar Items
-
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge
by: Grailoo, M., et al.
Published: (2026) -
PPU: Design and Implementation of a Pipelined Full Posit Processing Unit
by: Rossi, Federico, et al.
Published: (2023) -
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
by: Shen, Siyuan, et al.
Published: (2025) -
Latency and Privacy-Aware Resource Allocation in Vehicular Edge Computing
by: Ahmadvand, Hossein, et al.
Published: (2025) -
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
by: Zhou, Fang, et al.
Published: (2026)