PaCKD: Pattern-Clustered Knowledge Distillation for Compressing Memory Access Prediction Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gupta, Neelesh, Zhang, Pengmiao, Kannan, Rajgopal, Prasanna, Viktor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attention, Distillation, and Tabularization: Towards Practical Neural Network-Based Prefetching
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023)
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023)
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
von: Gupta, Neelesh, et al.
Veröffentlicht: (2026)
von: Gupta, Neelesh, et al.
Veröffentlicht: (2026)
Enabling Long FFT Convolutions on Memory-Constrained FPGAs via Chunking
von: Wang, Peter, et al.
Veröffentlicht: (2025)
von: Wang, Peter, et al.
Veröffentlicht: (2025)
FAME: FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption
von: Xu, Zhihan, et al.
Veröffentlicht: (2025)
von: Xu, Zhihan, et al.
Veröffentlicht: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
Federated Knowledge Distillation for Multi-Model Architectures Lithography Hotspot Detection
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
Memory Access Characterization of Large Language Models in CPU Environment and its Potential Impacts
von: Banasik, Spencer
Veröffentlicht: (2025)
von: Banasik, Spencer
Veröffentlicht: (2025)
Dynamic Sparse Attention: Access Patterns and Architecture
von: Levy, Noam
Veröffentlicht: (2026)
von: Levy, Noam
Veröffentlicht: (2026)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
von: Wolters, Christopher, et al.
Veröffentlicht: (2024)
von: Wolters, Christopher, et al.
Veröffentlicht: (2024)
SLOFetch: Compressed-Hierarchical Instruction Prefetching for Cloud Microservices
von: Bao, Zerui, et al.
Veröffentlicht: (2025)
von: Bao, Zerui, et al.
Veröffentlicht: (2025)
Efficient Data Access Paths for Mixed Vector-Relational Search
von: Sanca, Viktor, et al.
Veröffentlicht: (2024)
von: Sanca, Viktor, et al.
Veröffentlicht: (2024)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
von: Xiang, Maoyang, et al.
Veröffentlicht: (2025)
von: Xiang, Maoyang, et al.
Veröffentlicht: (2025)
Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
von: Xu, Bin, et al.
Veröffentlicht: (2025)
von: Xu, Bin, et al.
Veröffentlicht: (2025)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
von: Kim, Wonung, et al.
Veröffentlicht: (2025)
von: Kim, Wonung, et al.
Veröffentlicht: (2025)
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
Autoformalizing Memory Specifications with Agents
von: Ernst, Jan Ole, et al.
Veröffentlicht: (2026)
von: Ernst, Jan Ole, et al.
Veröffentlicht: (2026)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
AnalogSAGE: Self-evolving Analog Design Multi-Agents with Stratified Memory and Grounded Experience
von: Wang, Zining, et al.
Veröffentlicht: (2025)
von: Wang, Zining, et al.
Veröffentlicht: (2025)
Kernel Approximation using Analog In-Memory Computing
von: Büchel, Julian, et al.
Veröffentlicht: (2024)
von: Büchel, Julian, et al.
Veröffentlicht: (2024)
CAMformer: Associative Memory is All You Need
von: Molom-Ochir, Tergel, et al.
Veröffentlicht: (2025)
von: Molom-Ochir, Tergel, et al.
Veröffentlicht: (2025)
Efficient VQ-QAT and Mixed Vector/Linear quantized Neural Networks
von: Gou, Terry, et al.
Veröffentlicht: (2026)
von: Gou, Terry, et al.
Veröffentlicht: (2026)
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
von: Jayanth, Rakshith, et al.
Veröffentlicht: (2026)
von: Jayanth, Rakshith, et al.
Veröffentlicht: (2026)
FRED: Flexible REduction-Distribution Interconnect and Communication Implementation for Wafer-Scale Distributed Training of DNN Models
von: Rashidi, Saeed, et al.
Veröffentlicht: (2024)
von: Rashidi, Saeed, et al.
Veröffentlicht: (2024)
PACiM: A Sparsity-Centric Hybrid Compute-in-Memory Architecture via Probabilistic Approximation
von: Zhang, Wenlun, et al.
Veröffentlicht: (2024)
von: Zhang, Wenlun, et al.
Veröffentlicht: (2024)
Reward-Weighted On-Policy Distillation with an Open Property-Equivalence Verifier for NL-to-SVA Generation
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2025)
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2025)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
von: Shin, Duckgyu, et al.
Veröffentlicht: (2026)
von: Shin, Duckgyu, et al.
Veröffentlicht: (2026)
DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
von: Shao, Kunming, et al.
Veröffentlicht: (2026)
von: Shao, Kunming, et al.
Veröffentlicht: (2026)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
Effective and Memory-Efficient Alternatives to ECC for Reliable Large-Scale DNNs
von: Ahmadilivani, Mohammad Hasan, et al.
Veröffentlicht: (2026)
von: Ahmadilivani, Mohammad Hasan, et al.
Veröffentlicht: (2026)
AnalogNAS-Bench: A NAS Benchmark for Analog In-Memory Computing
von: Bessalah, Aniss, et al.
Veröffentlicht: (2025)
von: Bessalah, Aniss, et al.
Veröffentlicht: (2025)
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
von: Zheng, Size, et al.
Veröffentlicht: (2024)
von: Zheng, Size, et al.
Veröffentlicht: (2024)
U-SWIM: Universal Selective Write-Verify for Computing-in-Memory Neural Accelerators
von: Yan, Zheyu, et al.
Veröffentlicht: (2023)
von: Yan, Zheyu, et al.
Veröffentlicht: (2023)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
H2PIPE: High throughput CNN Inference on FPGAs with High-Bandwidth Memory
von: Doumet, Mario, et al.
Veröffentlicht: (2024)
von: Doumet, Mario, et al.
Veröffentlicht: (2024)
ArtNet: Hierarchical Clustering-Based Artificial Netlist Generator for ML and DTCO Application
von: Kang, Andrew B. Kahng. Seokhyeong, et al.
Veröffentlicht: (2025)
von: Kang, Andrew B. Kahng. Seokhyeong, et al.
Veröffentlicht: (2025)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Attention, Distillation, and Tabularization: Towards Practical Neural Network-Based Prefetching
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023) -
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
von: Gupta, Neelesh, et al.
Veröffentlicht: (2026) -
Enabling Long FFT Convolutions on Memory-Constrained FPGAs via Chunking
von: Wang, Peter, et al.
Veröffentlicht: (2025) -
FAME: FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption
von: Xu, Zhihan, et al.
Veröffentlicht: (2025) -
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)