LLM Inference Acceleration via Efficient Operation Fusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Salmani, Mahsa, Soloveychik, Ilya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accurate Block Quantization in LLMs with Outliers
von: Trukhanov, Nikita, et al.
Veröffentlicht: (2024)
von: Trukhanov, Nikita, et al.
Veröffentlicht: (2024)
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
von: Yang, Ze, et al.
Veröffentlicht: (2024)
von: Yang, Ze, et al.
Veröffentlicht: (2024)
Keyformer: KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023)
ProtocolLLM: RTL Benchmark for SystemVerilog Generation of Communication Protocols
von: Sheth, Arnav, et al.
Veröffentlicht: (2025)
von: Sheth, Arnav, et al.
Veröffentlicht: (2025)
LLM-FSM: Scaling Large Language Models for Finite-State Reasoning in RTL Code Generation
von: Wu, Yuheng, et al.
Veröffentlicht: (2026)
von: Wu, Yuheng, et al.
Veröffentlicht: (2026)
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
von: Gope, Dibakar, et al.
Veröffentlicht: (2024)
von: Gope, Dibakar, et al.
Veröffentlicht: (2024)
White-Box Reasoning: Synergizing LLM Strategy and gm/Id Data for Automated Analog Circuit Design
von: Chen, Jianqiu, et al.
Veröffentlicht: (2025)
von: Chen, Jianqiu, et al.
Veröffentlicht: (2025)
SLaNC: Static LayerNorm Calibration
von: Salmani, Mahsa, et al.
Veröffentlicht: (2024)
von: Salmani, Mahsa, et al.
Veröffentlicht: (2024)
Sangam: Chiplet-Based DRAM-PIM Accelerator with CXL Integration for LLM Inferencing
von: Kiyawat, Khyati, et al.
Veröffentlicht: (2025)
von: Kiyawat, Khyati, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
von: Fang, Yunhua, et al.
Veröffentlicht: (2025)
von: Fang, Yunhua, et al.
Veröffentlicht: (2025)
PRISM: Breaking the O(n) Memory Wall in Long-Context LLM Inference via O(1) Photonic Block Selection
von: Park, Hyoseok, et al.
Veröffentlicht: (2026)
von: Park, Hyoseok, et al.
Veröffentlicht: (2026)
HPU: High-Bandwidth Processing Unit for Scalable, Cost-effective LLM Inference via GPU Co-processing
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025)
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025)
HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
von: He, Zifan, et al.
Veröffentlicht: (2025)
von: He, Zifan, et al.
Veröffentlicht: (2025)
Chain-of-Descriptions: Improving Code LLMs for VHDL Code Generation and Summarization
von: Vijayaraghavan, Prashanth, et al.
Veröffentlicht: (2025)
von: Vijayaraghavan, Prashanth, et al.
Veröffentlicht: (2025)
FloorPlan-DeepSeek (FPDS): A multimodal approach to floorplan generation using vector-based next room prediction
von: Yin, Jun, et al.
Veröffentlicht: (2025)
von: Yin, Jun, et al.
Veröffentlicht: (2025)
Hardware Phi-1.5B: A Large Language Model Encodes Hardware Domain Specific Knowledge
von: Fu, Weimin, et al.
Veröffentlicht: (2024)
von: Fu, Weimin, et al.
Veröffentlicht: (2024)
Digital ASIC Design with Ongoing LLMs: Strategies and Prospects
von: Xiang, Maoyang, et al.
Veröffentlicht: (2024)
von: Xiang, Maoyang, et al.
Veröffentlicht: (2024)
From English to ASIC: Hardware Implementation with Large Language Model
von: Goh, Emil, et al.
Veröffentlicht: (2024)
von: Goh, Emil, et al.
Veröffentlicht: (2024)
InCoder-32B-Thinking: Industrial Code World Model for Thinking
von: Yang, Jian, et al.
Veröffentlicht: (2026)
von: Yang, Jian, et al.
Veröffentlicht: (2026)
DocEDA: Automated Extraction and Design of Analog Circuits from Documents with Large Language Model
von: Chen, Hong Cai, et al.
Veröffentlicht: (2024)
von: Chen, Hong Cai, et al.
Veröffentlicht: (2024)
EDA Corpus: A Large Language Model Dataset for Enhanced Interaction with OpenROAD
von: Wu, Bing-Yue, et al.
Veröffentlicht: (2024)
von: Wu, Bing-Yue, et al.
Veröffentlicht: (2024)
LLM-DSE: Searching Accelerator Parameters with LLM Agents
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models
von: Jiang, Wenqi, et al.
Veröffentlicht: (2023)
von: Jiang, Wenqi, et al.
Veröffentlicht: (2023)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
von: Zeng, Shulin, et al.
Veröffentlicht: (2024)
von: Zeng, Shulin, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
Accelerating Post-Quantum Cryptography via LLM-Driven Hardware-Software Co-Design
von: Liao, Yuchao, et al.
Veröffentlicht: (2026)
von: Liao, Yuchao, et al.
Veröffentlicht: (2026)
The Graph's Apprentice: Teaching an LLM Low Level Knowledge for Circuit Quality Estimation
von: Moravej, Reza, et al.
Veröffentlicht: (2024)
von: Moravej, Reza, et al.
Veröffentlicht: (2024)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
von: Ma, Songchen, et al.
Veröffentlicht: (2026)
von: Ma, Songchen, et al.
Veröffentlicht: (2026)
Pushing the Limits of BFP on Narrow Precision LLM Inference
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
DABench-LLM: Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators for LLMs
von: Hu, Ziyu, et al.
Veröffentlicht: (2025)
von: Hu, Ziyu, et al.
Veröffentlicht: (2025)
Automatically Improving LLM-based Verilog Generation using EDA Tool Feedback
von: Blocklove, Jason, et al.
Veröffentlicht: (2024)
von: Blocklove, Jason, et al.
Veröffentlicht: (2024)
RedFuser: An Automatic Operator Fusion Framework for Cascaded Reductions on AI Accelerators
von: Tang, Xinsheng, et al.
Veröffentlicht: (2026)
von: Tang, Xinsheng, et al.
Veröffentlicht: (2026)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
von: Wang, Hanrui, et al.
Veröffentlicht: (2020)
von: Wang, Hanrui, et al.
Veröffentlicht: (2020)
Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2025)
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2025)
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
von: Chen, Chun-Ting, et al.
Veröffentlicht: (2025)
von: Chen, Chun-Ting, et al.
Veröffentlicht: (2025)
BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge
von: Ji, Yuhao, et al.
Veröffentlicht: (2024)
von: Ji, Yuhao, et al.
Veröffentlicht: (2024)
Comparative Characterization of KV Cache Management Strategies for LLM Inference
von: Mamo, Oteo, et al.
Veröffentlicht: (2026)
von: Mamo, Oteo, et al.
Veröffentlicht: (2026)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Accurate Block Quantization in LLMs with Outliers
von: Trukhanov, Nikita, et al.
Veröffentlicht: (2024) -
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
von: Yang, Ze, et al.
Veröffentlicht: (2024) -
Keyformer: KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024) -
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023) -
ProtocolLLM: RTL Benchmark for SystemVerilog Generation of Communication Protocols
von: Sheth, Arnav, et al.
Veröffentlicht: (2025)