Mugi: Value Level Parallelism For Efficient LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Price, Daniel, Vellaisamy, Prabhu, Shen, John, Wu, Di |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
par: Nair, Harideep, et autres
Publié: (2024)
par: Nair, Harideep, et autres
Publié: (2024)
Catwalk: Unary Top-K for Efficient Ramp-No-Leak Neuron Design for Temporal Neural Networks
par: Lister, Devon, et autres
Publié: (2025)
par: Lister, Devon, et autres
Publié: (2025)
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
par: Vellaisamy, Prabhu, et autres
Publié: (2026)
par: Vellaisamy, Prabhu, et autres
Publié: (2026)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
par: Zhang, Zehuan, et autres
Publié: (2026)
par: Zhang, Zehuan, et autres
Publié: (2026)
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
par: de Lima, João Paulo Cardoso, et autres
Publié: (2025)
par: de Lima, João Paulo Cardoso, et autres
Publié: (2025)
Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
par: Nair, Harideep, et autres
Publié: (2024)
par: Nair, Harideep, et autres
Publié: (2024)
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
par: Leon, Vasileios, et autres
Publié: (2025)
par: Leon, Vasileios, et autres
Publié: (2025)
PICBench: Benchmarking LLMs for Photonic Integrated Circuits Design
par: Wu, Yuchao, et autres
Publié: (2025)
par: Wu, Yuchao, et autres
Publié: (2025)
TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
par: Shen, Chaoyao, et autres
Publié: (2026)
par: Shen, Chaoyao, et autres
Publié: (2026)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
par: Vellaisamy, Prabhu, et autres
Publié: (2024)
par: Vellaisamy, Prabhu, et autres
Publié: (2024)
Effective and Memory-Efficient Alternatives to ECC for Reliable Large-Scale DNNs
par: Ahmadilivani, Mohammad Hasan, et autres
Publié: (2026)
par: Ahmadilivani, Mohammad Hasan, et autres
Publié: (2026)
HLSFactory: A Framework Empowering High-Level Synthesis Datasets for Machine Learning and Beyond
par: Abi-Karam, Stefan, et autres
Publié: (2024)
par: Abi-Karam, Stefan, et autres
Publié: (2024)
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
par: Venkatachalam, Shanmuga, et autres
Publié: (2026)
par: Venkatachalam, Shanmuga, et autres
Publié: (2026)
RL-MUL 2.0: Multiplier Design Optimization with Parallel Deep Reinforcement Learning and Space Reduction
par: Zuo, Dongsheng, et autres
Publié: (2024)
par: Zuo, Dongsheng, et autres
Publié: (2024)
Hardware-Aware Data and Instruction Mapping for AI Tasks: Balancing Parallelism, I/O and Memory Tradeoffs
par: Chowdhury, Md Rownak Hossain, et autres
Publié: (2025)
par: Chowdhury, Md Rownak Hossain, et autres
Publié: (2025)
Deep Inverse Design for High-Level Synthesis
par: Chang, Ping, et autres
Publié: (2024)
par: Chang, Ping, et autres
Publié: (2024)
An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes
par: Wang, Miaoxin, et autres
Publié: (2024)
par: Wang, Miaoxin, et autres
Publié: (2024)
FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference
par: Hsieh, Fen-Yu, et autres
Publié: (2025)
par: Hsieh, Fen-Yu, et autres
Publié: (2025)
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
par: Wu, Qizhe, et autres
Publié: (2024)
par: Wu, Qizhe, et autres
Publié: (2024)
ACE-RTL: When Agentic Context Evolution Meets RTL-Specialized LLMs
par: Deng, Chenhui, et autres
Publié: (2026)
par: Deng, Chenhui, et autres
Publié: (2026)
EvolveGen: Algorithmic Level Hardware Model Checking Benchmark Generation through Reinforcement Learning
par: Hu, Guangyu, et autres
Publié: (2026)
par: Hu, Guangyu, et autres
Publié: (2026)
RTL-Repo: A Benchmark for Evaluating LLMs on Large-Scale RTL Design Projects
par: Allam, Ahmed, et autres
Publié: (2024)
par: Allam, Ahmed, et autres
Publié: (2024)
COMET: Towards Partical W4A4KV4 LLMs Serving
par: Liu, Lian, et autres
Publié: (2024)
par: Liu, Lian, et autres
Publié: (2024)
Efficient and Reliable Vector Similarity Search Using Asymmetric Encoding with NAND-Flash for Many-Class Few-Shot Learning
par: Chiang, Hao-Wei, et autres
Publié: (2024)
par: Chiang, Hao-Wei, et autres
Publié: (2024)
NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction
par: Wadle, Shayne, et autres
Publié: (2025)
par: Wadle, Shayne, et autres
Publié: (2025)
Active Imitation Learning for Thermal- and Kernel-Aware LFM Inference on 3D S-NUCA Many-Cores
par: Shen, Yixian, et autres
Publié: (2026)
par: Shen, Yixian, et autres
Publié: (2026)
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
par: Vellaisamy, Prabhu, et autres
Publié: (2024)
par: Vellaisamy, Prabhu, et autres
Publié: (2024)
Efficient Tabular Data Preprocessing of ML Pipelines
par: Zhu, Yu, et autres
Publié: (2024)
par: Zhu, Yu, et autres
Publié: (2024)
TransAxx: Efficient Transformers with Approximate Computing
par: Danopoulos, Dimitrios, et autres
Publié: (2024)
par: Danopoulos, Dimitrios, et autres
Publié: (2024)
Designing Efficient LLM Accelerators for Edge Devices
par: Haris, Jude, et autres
Publié: (2024)
par: Haris, Jude, et autres
Publié: (2024)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
par: Xie, Xilong, et autres
Publié: (2025)
par: Xie, Xilong, et autres
Publié: (2025)
A 65nm 8b-Activation 8b-Weight SRAM-Based Charge-Domain Computing-in-Memory Macro Using A Fully-Parallel Analog Adder Network and A Single-ADC Interface
par: Yin, Guodong, et autres
Publié: (2022)
par: Yin, Guodong, et autres
Publié: (2022)
Intelligent4DSE: Optimizing High-Level Synthesis Design Space Exploration with Graph Neural Networks and Large Language Models
par: Xu, Lei, et autres
Publié: (2025)
par: Xu, Lei, et autres
Publié: (2025)
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
par: Shin, Duckgyu, et autres
Publié: (2026)
par: Shin, Duckgyu, et autres
Publié: (2026)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
par: Wang, Chenyu, et autres
Publié: (2023)
par: Wang, Chenyu, et autres
Publié: (2023)
CircuitVAE: Efficient and Scalable Latent Circuit Optimization
par: Song, Jialin, et autres
Publié: (2024)
par: Song, Jialin, et autres
Publié: (2024)
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
par: Kang, Hao, et autres
Publié: (2024)
par: Kang, Hao, et autres
Publié: (2024)
TinyFormer: Efficient Transformer Design and Deployment on Tiny Devices
par: Yang, Jianlei, et autres
Publié: (2023)
par: Yang, Jianlei, et autres
Publié: (2023)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
par: İslamoğlu, Gamze, et autres
Publié: (2023)
par: İslamoğlu, Gamze, et autres
Publié: (2023)
Efficient and Mathematically Robust Operations for Certified Neural Networks Inference
par: Geyer, Fabien, et autres
Publié: (2024)
par: Geyer, Fabien, et autres
Publié: (2024)
Documents similaires
-
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
par: Nair, Harideep, et autres
Publié: (2024) -
Catwalk: Unary Top-K for Efficient Ramp-No-Leak Neuron Design for Temporal Neural Networks
par: Lister, Devon, et autres
Publié: (2025) -
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
par: Vellaisamy, Prabhu, et autres
Publié: (2026) -
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
par: Zhang, Zehuan, et autres
Publié: (2026) -
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
par: de Lima, João Paulo Cardoso, et autres
Publié: (2025)