Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Ming-Yen, Waqar, Faaiq, Yang, Hanchen, Karim, Muhammed Ahosan Ul, Simka, Harsono, Yu, Shimeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CMOS+X: Stacking Persistent Embedded Memories based on Oxide Transistors upon GPGPU Platforms
by: Waqar, Faaiq, et al.
Published: (2025)
by: Waqar, Faaiq, et al.
Published: (2025)
Monolithic 3D FPGAs Utilizing Back-End-of-Line Configuration Memories
by: Waqar, Faaiq, et al.
Published: (2025)
by: Waqar, Faaiq, et al.
Published: (2025)
Optimization and Benchmarking of Monolithically Stackable Gain Cell Memory for Last-Level Cache
by: Waqar, Faaiq, et al.
Published: (2025)
by: Waqar, Faaiq, et al.
Published: (2025)
Time Domain Near Memory Computing Engine
by: Antal, Sarthak, et al.
Published: (2026)
by: Antal, Sarthak, et al.
Published: (2026)
A WASM-Subset Stack Architecture for Low-cost FPGAs using Open-Source EDA Flows
by: Chakrabarti, Aradhya
Published: (2025)
by: Chakrabarti, Aradhya
Published: (2025)
GainSight: A Unified Framework for Data Lifetime Profiling and Heterogeneous Memory Composition
by: Li, Peijing, et al.
Published: (2025)
by: Li, Peijing, et al.
Published: (2025)
Improved Prefetching Techniques for Linked Data Structures
by: Maruszewski, Nikola Vuk
Published: (2025)
by: Maruszewski, Nikola Vuk
Published: (2025)
Nonvolatile Charge-Domain Attention with HZO Ferroelectric Capacitors: A Simulation-Based Device-to-System Evaluation
by: Abouagour, Faris
Published: (2026)
by: Abouagour, Faris
Published: (2026)
ECOLogic: Enabling Circular, Obfuscated, and Adaptive Logic via eFPGA-Augmented SoCs
by: Tashdid, Ishraq, et al.
Published: (2025)
by: Tashdid, Ishraq, et al.
Published: (2025)
A Customized Memory-aware Architecture for Biological Sequence Alignment
by: Akbari, Nasrin, et al.
Published: (2025)
by: Akbari, Nasrin, et al.
Published: (2025)
Scalable Digital Compute-in-Memory Ising Machines for Robustness Verification of Binary Neural Networks
by: Vadlamani, Madhav, et al.
Published: (2026)
by: Vadlamani, Madhav, et al.
Published: (2026)
HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration
by: Hsu, Po-Kai, et al.
Published: (2026)
by: Hsu, Po-Kai, et al.
Published: (2026)
Apparate: Evading Memory Hierarchy with GodSpeed Wireless-on-Chip
by: GS, Nitesh Narayana, et al.
Published: (2024)
by: GS, Nitesh Narayana, et al.
Published: (2024)
TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations
by: Sedukhin, Stanislav, et al.
Published: (2025)
by: Sedukhin, Stanislav, et al.
Published: (2025)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025)
by: Liu, Lian, et al.
Published: (2025)
MEDEA: A Design-Time Multi-Objective Manager for Energy-Efficient DNN Inference on Heterogeneous Ultra-Low Power Platforms
by: Taji, Hossein, et al.
Published: (2025)
by: Taji, Hossein, et al.
Published: (2025)
MASIM: An Efficient Multi-Array Scheduler for In-Memory SIMD Computation
by: Qian, Xingyue, et al.
Published: (2024)
by: Qian, Xingyue, et al.
Published: (2024)
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
by: Parameshwara, Arya
Published: (2025)
by: Parameshwara, Arya
Published: (2025)
A Comparative Study of Digital Memristor-Based Processing-In-Memory from a Device and Reliability Perspective
by: Neuner, Thomas, et al.
Published: (2026)
by: Neuner, Thomas, et al.
Published: (2026)
RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
by: Yildirim, Muhammed, et al.
Published: (2025)
by: Yildirim, Muhammed, et al.
Published: (2025)
XL-HD: Extended Learning in Hyperdimensional Computing via Deterministic Projections for In-Memory Accelerators
by: Moon, Sabrina Hassan, et al.
Published: (2026)
by: Moon, Sabrina Hassan, et al.
Published: (2026)
HERMES: High-Performance RISC-V Memory Hierarchy for ML Workloads
by: Suryadevara, Pranav
Published: (2025)
by: Suryadevara, Pranav
Published: (2025)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
A Spatio-Temporal Graph Neural Networks Approach for Predicting Silent Data Corruption inducing Circuit-Level Faults
by: Wei, Shaoqi, et al.
Published: (2025)
by: Wei, Shaoqi, et al.
Published: (2025)
Tekum: Balanced Ternary Tapered Precision Real Arithmetic
by: Hunhold, Laslo
Published: (2025)
by: Hunhold, Laslo
Published: (2025)
Fine Grain 3D Integration for Microarchitecture Design Through Cube Packing Exploration
by: Liu, Yongxiang, et al.
Published: (2025)
by: Liu, Yongxiang, et al.
Published: (2025)
OpenEye: A Scalable Open-Source Hardware Accelerator for DNNs
by: Lebold, Denis, et al.
Published: (2026)
by: Lebold, Denis, et al.
Published: (2026)
Computing with Clocks
by: Edwards, Jonathan, et al.
Published: (2024)
by: Edwards, Jonathan, et al.
Published: (2024)
A 5T-2MTJ STT-assisted Spin Orbit Torque based Ternary Content Addressable Memory for Hardware Accelerators
by: Narla, Siri, et al.
Published: (2024)
by: Narla, Siri, et al.
Published: (2024)
SISA: A Scale-In Systolic Array for GEMM Acceleration
by: Altamura, Luigi, et al.
Published: (2026)
by: Altamura, Luigi, et al.
Published: (2026)
Fast and Practical Strassen's Matrix Multiplication using FPGAs
by: Ahmad, Afzal, et al.
Published: (2024)
by: Ahmad, Afzal, et al.
Published: (2024)
Non-interfering On-line and In-field SoC Testing
by: Strauch, Tobias
Published: (2024)
by: Strauch, Tobias
Published: (2024)
Analog-to-Stochastic Converter Using Magnetic Tunnel Junction Devices for Vision Chips
by: Onizawa, Naoya, et al.
Published: (2026)
by: Onizawa, Naoya, et al.
Published: (2026)
CLSA-CIM: A Cross-Layer Scheduling Approach for Computing-in-Memory Architectures
by: Pelke, Rebecca, et al.
Published: (2024)
by: Pelke, Rebecca, et al.
Published: (2024)
OPIMA: Optical Processing-In-Memory for Convolutional Neural Network Acceleration
by: Sunny, Febin, et al.
Published: (2024)
by: Sunny, Febin, et al.
Published: (2024)
Sequencing on Silicon: AI SoC Design for Mobile Genomics at the Edge
by: Magierowski, Sebastian, et al.
Published: (2025)
by: Magierowski, Sebastian, et al.
Published: (2025)
Hardware Accelerators for Artificial Intelligence
by: Ahsan, S M Mojahidul, et al.
Published: (2024)
by: Ahsan, S M Mojahidul, et al.
Published: (2024)
AMC: Access to Miss Correlation Prefetcher for Evolving Graph Analytics
by: Singh, Abhishek, et al.
Published: (2024)
by: Singh, Abhishek, et al.
Published: (2024)
RACE-IT: A Reconfigurable Analog Computing Engine for In-Memory Transformer Acceleration
by: Zhao, Lei, et al.
Published: (2023)
by: Zhao, Lei, et al.
Published: (2023)
All-in-One Analog AI Hardware: On-Chip Training and Inference with Conductive-Metal-Oxide/HfOx ReRAM Devices
by: Falcone, Donato Francesco, et al.
Published: (2025)
by: Falcone, Donato Francesco, et al.
Published: (2025)
Similar Items
-
CMOS+X: Stacking Persistent Embedded Memories based on Oxide Transistors upon GPGPU Platforms
by: Waqar, Faaiq, et al.
Published: (2025) -
Monolithic 3D FPGAs Utilizing Back-End-of-Line Configuration Memories
by: Waqar, Faaiq, et al.
Published: (2025) -
Optimization and Benchmarking of Monolithically Stackable Gain Cell Memory for Last-Level Cache
by: Waqar, Faaiq, et al.
Published: (2025) -
Time Domain Near Memory Computing Engine
by: Antal, Sarthak, et al.
Published: (2026) -
A WASM-Subset Stack Architecture for Low-cost FPGAs using Open-Source EDA Flows
by: Chakrabarti, Aradhya
Published: (2025)