Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yubeaton, Patrick, Mahmoud, Tareq, Naga, Shehab, Taheri, Pooria, Xia, Tianhua, George, Arun, Khalil, Yasmein, Zhang, Sai Qian, Joshi, Siddharth, Hegde, Chinmay, Garg, Siddharth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring the Agentic Frontier of Verilog Code Generation
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2026)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2026)
VeriInteresting: An Empirical Study of Model Prompt Interactions in Verilog Code Generation
von: Collini, Luca, et al.
Veröffentlicht: (2026)
von: Collini, Luca, et al.
Veröffentlicht: (2026)
Rome was Not Built in a Single Step: Hierarchical Prompting for LLM-based Chip Design
von: Nakkab, Andre, et al.
Veröffentlicht: (2024)
von: Nakkab, Andre, et al.
Veröffentlicht: (2024)
PrefixLLM: LLM-aided Prefix Circuit Design
von: Xiao, Weihua, et al.
Veröffentlicht: (2024)
von: Xiao, Weihua, et al.
Veröffentlicht: (2024)
LLM-Aided Testbench Generation and Bug Detection for Finite-State Machines
von: Bhandari, Jitendra, et al.
Veröffentlicht: (2024)
von: Bhandari, Jitendra, et al.
Veröffentlicht: (2024)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
von: Wei, Jianyu, et al.
Veröffentlicht: (2025)
von: Wei, Jianyu, et al.
Veröffentlicht: (2025)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
von: Xia, Tianhua, et al.
Veröffentlicht: (2025)
von: Xia, Tianhua, et al.
Veröffentlicht: (2025)
Hyft: A Reconfigurable Softmax Accelerator with Hybrid Numeric Format for both Training and Inference
von: Xia, Tianhua, et al.
Veröffentlicht: (2023)
von: Xia, Tianhua, et al.
Veröffentlicht: (2023)
FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification
von: Wan, Gwok-Waa, et al.
Veröffentlicht: (2025)
von: Wan, Gwok-Waa, et al.
Veröffentlicht: (2025)
PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
C2HLSC: Can LLMs Bridge the Software-to-Hardware Design Gap?
von: Collini, Luca, et al.
Veröffentlicht: (2024)
von: Collini, Luca, et al.
Veröffentlicht: (2024)
SLDB: An End-To-End Heterogeneous System-on-Chip Benchmark Suite for LLM-Aided Design
von: Alvanaki, Elisavet Lydia, et al.
Veröffentlicht: (2025)
von: Alvanaki, Elisavet Lydia, et al.
Veröffentlicht: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
von: Yao, Jiayi, et al.
Veröffentlicht: (2026)
von: Yao, Jiayi, et al.
Veröffentlicht: (2026)
SZKP: A Scalable Accelerator Architecture for Zero-Knowledge Proofs
von: Daftardar, Alhad, et al.
Veröffentlicht: (2024)
von: Daftardar, Alhad, et al.
Veröffentlicht: (2024)
C2HLSC: Leveraging Large Language Models to Bridge the Software-to-Hardware Design Gap
von: Collini, Luca, et al.
Veröffentlicht: (2024)
von: Collini, Luca, et al.
Veröffentlicht: (2024)
TruncFormer: Private LLM Inference Using Only Truncations
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2024)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2024)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
Automatically Improving LLM-based Verilog Generation using EDA Tool Feedback
von: Blocklove, Jason, et al.
Veröffentlicht: (2024)
von: Blocklove, Jason, et al.
Veröffentlicht: (2024)
PowerFlow-DNN: Compiler-Directed Fine-Grained Power Orchestration for End-to-End Edge AI Inference
von: Chen, Paul, et al.
Veröffentlicht: (2026)
von: Chen, Paul, et al.
Veröffentlicht: (2026)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
von: DeBole, Michael V., et al.
Veröffentlicht: (2025)
von: DeBole, Michael V., et al.
Veröffentlicht: (2025)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
ArchXBench: A Complex Digital Systems Benchmark Suite for LLM Driven RTL Synthesis
von: Purini, Suresh, et al.
Veröffentlicht: (2025)
von: Purini, Suresh, et al.
Veröffentlicht: (2025)
TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
Masala-CHAI: A Large-Scale SPICE Netlist Dataset for Analog Circuits by Harnessing AI
von: Bhandari, Jitendra, et al.
Veröffentlicht: (2024)
von: Bhandari, Jitendra, et al.
Veröffentlicht: (2024)
Efficient yet Accurate End-to-End SC Accelerator Design
von: Li, Meng, et al.
Veröffentlicht: (2024)
von: Li, Meng, et al.
Veröffentlicht: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
von: Sun, Xiaotian, et al.
Veröffentlicht: (2024)
von: Sun, Xiaotian, et al.
Veröffentlicht: (2024)
Make Every Move Count: LLM-based High-Quality RTL Code Generation Using MCTS
von: DeLorenzo, Matthew, et al.
Veröffentlicht: (2024)
von: DeLorenzo, Matthew, et al.
Veröffentlicht: (2024)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
von: Guo, Yipin, et al.
Veröffentlicht: (2026)
von: Guo, Yipin, et al.
Veröffentlicht: (2026)
Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
von: Yoon, Dongho, et al.
Veröffentlicht: (2025)
von: Yoon, Dongho, et al.
Veröffentlicht: (2025)
Voyager: An End-to-End Framework for Design-Space Exploration and Generation of DNN Accelerators
von: Prabhu, Kartik, et al.
Veröffentlicht: (2025)
von: Prabhu, Kartik, et al.
Veröffentlicht: (2025)
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
HAAN: A Holistic Approach for Accelerating Normalization Operations in Large Language Models
von: Peng, Tianfan, et al.
Veröffentlicht: (2025)
von: Peng, Tianfan, et al.
Veröffentlicht: (2025)
AutoGNN: End-to-End Hardware-Driven Graph Preprocessing for Enhanced GNN Performance
von: Kang, Seungkwan, et al.
Veröffentlicht: (2026)
von: Kang, Seungkwan, et al.
Veröffentlicht: (2026)
Towards an End-To-End System for Real-Time Gesture Recognition from Surface Vibrations
von: Hettstedt, Florian, et al.
Veröffentlicht: (2026)
von: Hettstedt, Florian, et al.
Veröffentlicht: (2026)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
MemIntelli: A Generic End-to-End Simulation Framework for Memristive Intelligent Computing
von: Zhou, Houji, et al.
Veröffentlicht: (2025)
von: Zhou, Houji, et al.
Veröffentlicht: (2025)
FASE: FPGA-Assisted Syscall Emulation for Rapid End-to-End Processor Performance Validation
von: Meng, Chengzhen, et al.
Veröffentlicht: (2025)
von: Meng, Chengzhen, et al.
Veröffentlicht: (2025)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
von: Cheng, Jianyi, et al.
Veröffentlicht: (2023)
von: Cheng, Jianyi, et al.
Veröffentlicht: (2023)
MTU: The Multifunction Tree Unit for Accelerating Zero-Knowledge Proofs
von: Mo, Jianqiao, et al.
Veröffentlicht: (2025)
von: Mo, Jianqiao, et al.
Veröffentlicht: (2025)
HSCO-Bench: An Agent-Driven End-to-End Hardware-Software Co-design Benchmark for Systems-on-Chip
von: Tsai, Pei-Huan, et al.
Veröffentlicht: (2026)
von: Tsai, Pei-Huan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Exploring the Agentic Frontier of Verilog Code Generation
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2026) -
VeriInteresting: An Empirical Study of Model Prompt Interactions in Verilog Code Generation
von: Collini, Luca, et al.
Veröffentlicht: (2026) -
Rome was Not Built in a Single Step: Hierarchical Prompting for LLM-based Chip Design
von: Nakkab, Andre, et al.
Veröffentlicht: (2024) -
PrefixLLM: LLM-aided Prefix Circuit Design
von: Xiao, Weihua, et al.
Veröffentlicht: (2024) -
LLM-Aided Testbench Generation and Bug Detection for Finite-State Machines
von: Bhandari, Jitendra, et al.
Veröffentlicht: (2024)