COMET: Towards Partical W4A4KV4 LLMs Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Lian, Ren, Haimeng, Cheng, Long, Xu, Zhaohui, Pan, Yudong, Wang, Mengdi, Li, Xiaowei, Han, Yinhe, Wang, Ying |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025)
by: Liu, Lian, et al.
Published: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
by: Liu, Lian, et al.
Published: (2026)
by: Liu, Lian, et al.
Published: (2026)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
ChipGPT: How far are we from natural language hardware design
by: Chang, Kaiyan, et al.
Published: (2023)
by: Chang, Kaiyan, et al.
Published: (2023)
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
by: Zhao, Shixin, et al.
Published: (2025)
by: Zhao, Shixin, et al.
Published: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
by: Pan, Yudong, et al.
Published: (2026)
by: Pan, Yudong, et al.
Published: (2026)
ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators
by: Zou, Guoqiang, et al.
Published: (2025)
by: Zou, Guoqiang, et al.
Published: (2025)
CLASS: A Controller-Centric Layout Synthesizer for Dynamic Quantum Circuits
by: Chen, Yu, et al.
Published: (2025)
by: Chen, Yu, et al.
Published: (2025)
PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Data is all you need: Finetuning LLMs for Chip Design via an Automated design-data augmentation framework
by: Chang, Kaiyan, et al.
Published: (2024)
by: Chang, Kaiyan, et al.
Published: (2024)
A System Architecture for Low Latency Multiprogramming Quantum Computing
by: Zhao, Yilun, et al.
Published: (2026)
by: Zhao, Yilun, et al.
Published: (2026)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
by: Sun, Xiaotian, et al.
Published: (2024)
by: Sun, Xiaotian, et al.
Published: (2024)
PIMSYN: Synthesizing Processing-in-memory CNN Accelerators
by: Li, Wanqian, et al.
Published: (2024)
by: Li, Wanqian, et al.
Published: (2024)
Natural language is not enough: Benchmarking multi-modal generative AI for Verilog generation
by: Chang, Kaiyan, et al.
Published: (2024)
by: Chang, Kaiyan, et al.
Published: (2024)
Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL
by: Cai, Siyang, et al.
Published: (2026)
by: Cai, Siyang, et al.
Published: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
by: Meng, William, et al.
Published: (2025)
by: Meng, William, et al.
Published: (2025)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
by: li, Fei, et al.
Published: (2026)
by: li, Fei, et al.
Published: (2026)
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
by: Zou, Jiaxiang, et al.
Published: (2026)
by: Zou, Jiaxiang, et al.
Published: (2026)
ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement Learning
by: Chen, Zhirong, et al.
Published: (2025)
by: Chen, Zhirong, et al.
Published: (2025)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
by: Xia, Tianhua, et al.
Published: (2025)
by: Xia, Tianhua, et al.
Published: (2025)
TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
by: Bertuletti, Marco, et al.
Published: (2026)
by: Bertuletti, Marco, et al.
Published: (2026)
Towards Reliable Systems: A Scalable Approach to AXI4 Transaction Monitoring
by: Liang, Chaoqun, et al.
Published: (2025)
by: Liang, Chaoqun, et al.
Published: (2025)
Automated SVA Generation with LLMs
by: Fu, Lik Tung, et al.
Published: (2026)
by: Fu, Lik Tung, et al.
Published: (2026)
Lifecycle Cost-Effectiveness Modeling for Redundancy-Enhanced Multi-Chiplet Architectures
by: Liu, Zizhen, et al.
Published: (2026)
by: Liu, Zizhen, et al.
Published: (2026)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
by: Wang, Zhican, et al.
Published: (2025)
by: Wang, Zhican, et al.
Published: (2025)
Rethinking Compute Substrates for 3D-Stacked Near-Memory LLM Decoding: Microarchitecture-Scheduling Co-Design
by: Ai, Chenyang, et al.
Published: (2026)
by: Ai, Chenyang, et al.
Published: (2026)
HLSPilot: LLM-based High-Level Synthesis
by: Xiong, Chenwei, et al.
Published: (2024)
by: Xiong, Chenwei, et al.
Published: (2024)
UVLLM: An Automated Universal RTL Verification Framework using LLMs
by: Hu, Yuchen, et al.
Published: (2024)
by: Hu, Yuchen, et al.
Published: (2024)
Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems
by: Yamamoto, Yuji, et al.
Published: (2026)
by: Yamamoto, Yuji, et al.
Published: (2026)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
by: Li, Cong, et al.
Published: (2026)
by: Li, Cong, et al.
Published: (2026)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
by: Chen, Peilin, et al.
Published: (2025)
by: Chen, Peilin, et al.
Published: (2025)
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
by: Lin, Chenqi, et al.
Published: (2025)
by: Lin, Chenqi, et al.
Published: (2025)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
by: Xu, Weikai, et al.
Published: (2025)
by: Xu, Weikai, et al.
Published: (2025)
A Memory-Efficient Retrieval Architecture for RAG-Enabled Wearable Medical LLMs-Agents
by: Liao, Zhipeng, et al.
Published: (2025)
by: Liao, Zhipeng, et al.
Published: (2025)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
by: Zhang, Junming, et al.
Published: (2026)
by: Zhang, Junming, et al.
Published: (2026)
ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
by: Kim, Junsoo, et al.
Published: (2025)
by: Kim, Junsoo, et al.
Published: (2025)
Chiplet Cloud: Building AI Supercomputers for Serving Large Generative Language Models
by: Peng, Huwan, et al.
Published: (2023)
by: Peng, Huwan, et al.
Published: (2023)
MPM-LLM4DSE: Reaching the Pareto Frontier in HLS with Multimodal Learning and LLM-Driven Exploration
by: Xu, Lei, et al.
Published: (2026)
by: Xu, Lei, et al.
Published: (2026)
Similar Items
-
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025) -
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
by: Liu, Lian, et al.
Published: (2026) -
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026) -
ChipGPT: How far are we from natural language hardware design
by: Chang, Kaiyan, et al.
Published: (2023) -
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
by: Liu, Yiqi, et al.
Published: (2026)