ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Junsoo, Lee, Hunjong, Ko, Geonwoo, Choi, Gyubin, Ham, Seri, Hong, Seongmin, Kim, Joo-Young |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
di: Kim, Minsu, et al.
Pubblicazione: (2025)
di: Kim, Minsu, et al.
Pubblicazione: (2025)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
di: Moon, Seungjae, et al.
Pubblicazione: (2024)
di: Moon, Seungjae, et al.
Pubblicazione: (2024)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
di: Hong, Jeongmin, et al.
Pubblicazione: (2024)
di: Hong, Jeongmin, et al.
Pubblicazione: (2024)
MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization
di: Kim, Seeyeon, et al.
Pubblicazione: (2026)
di: Kim, Seeyeon, et al.
Pubblicazione: (2026)
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
di: Kim, Jae-Young, et al.
Pubblicazione: (2025)
di: Kim, Jae-Young, et al.
Pubblicazione: (2025)
HURRY: Highly Utilized, Reconfigurable ReRAM-based In-situ Accelerator with Multifunctionality
di: Shin, Hery, et al.
Pubblicazione: (2024)
di: Shin, Hery, et al.
Pubblicazione: (2024)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
di: Heo, Guseul, et al.
Pubblicazione: (2024)
di: Heo, Guseul, et al.
Pubblicazione: (2024)
Low-overhead General-purpose Near-Data Processing in CXL Memory Expanders
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024)
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
di: Yoon, Jieon, et al.
Pubblicazione: (2026)
di: Yoon, Jieon, et al.
Pubblicazione: (2026)
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
di: Han, Wontak, et al.
Pubblicazione: (2024)
di: Han, Wontak, et al.
Pubblicazione: (2024)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
di: Li, Boyu, et al.
Pubblicazione: (2025)
di: Li, Boyu, et al.
Pubblicazione: (2025)
System-Level Design Space Exploration for High-Level Synthesis under End-to-End Latency Constraints
di: Liao, Yuchao, et al.
Pubblicazione: (2024)
di: Liao, Yuchao, et al.
Pubblicazione: (2024)
ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration
di: Lee, Hangyeol, et al.
Pubblicazione: (2026)
di: Lee, Hangyeol, et al.
Pubblicazione: (2026)
SCRec: A Scalable Computational Storage System with Statistical Sharding and Tensor-train Decomposition for Recommendation Models
di: Yang, Jinho, et al.
Pubblicazione: (2025)
di: Yang, Jinho, et al.
Pubblicazione: (2025)
APINT: A Full-Stack Framework for Acceleration of Privacy-Preserving Inference of Transformers based on Garbled Circuits
di: Cho, Hyunjun, et al.
Pubblicazione: (2025)
di: Cho, Hyunjun, et al.
Pubblicazione: (2025)
SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
di: Lee, Minjae, et al.
Pubblicazione: (2023)
di: Lee, Minjae, et al.
Pubblicazione: (2023)
Securing DRAM at Scale: ARFM-Driven Row Hammer Defense with Unveiling the Threat of Short tRC Patterns
di: Joo, Nogeun, et al.
Pubblicazione: (2025)
di: Joo, Nogeun, et al.
Pubblicazione: (2025)
LIMCA: LLM for Automating Analog In-Memory Computing Architecture Design Exploration
di: Vungarala, Deepak, et al.
Pubblicazione: (2025)
di: Vungarala, Deepak, et al.
Pubblicazione: (2025)
RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNs
di: Ko, Hanum, et al.
Pubblicazione: (2026)
di: Ko, Hanum, et al.
Pubblicazione: (2026)
Polaris: Multi-Fidelity Design Space Exploration of Deep Learning Accelerators
di: Sakhuja, Chirag, et al.
Pubblicazione: (2024)
di: Sakhuja, Chirag, et al.
Pubblicazione: (2024)
SSR: Spatial Sequential Hybrid Architecture for Latency Throughput Tradeoff in Transformer Acceleration
di: Zhuang, Jinming, et al.
Pubblicazione: (2024)
di: Zhuang, Jinming, et al.
Pubblicazione: (2024)
Voyager: An End-to-End Framework for Design-Space Exploration and Generation of DNN Accelerators
di: Prabhu, Kartik, et al.
Pubblicazione: (2025)
di: Prabhu, Kartik, et al.
Pubblicazione: (2025)
STRAW: A Stress-Aware WL-Based Read Reclaim Technique for High-Density NAND Flash-Based SSDs
di: Chun, Myoungjun, et al.
Pubblicazione: (2025)
di: Chun, Myoungjun, et al.
Pubblicazione: (2025)
Cerberus: Cross-Layer ECC Co-Design for Robust and Efficient Memory Protection
di: Kim, Junhwan, et al.
Pubblicazione: (2026)
di: Kim, Junhwan, et al.
Pubblicazione: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
di: Meng, William, et al.
Pubblicazione: (2025)
di: Meng, William, et al.
Pubblicazione: (2025)
A 0.5V, 6.2$μ$W, 0.059mm$^{2}$ Sinusoidal Current Generator IC with 0.088% THD for Bio-Impedance Sensing
di: Kim, Kwantae, et al.
Pubblicazione: (2024)
di: Kim, Kwantae, et al.
Pubblicazione: (2024)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
di: Li, Cong, et al.
Pubblicazione: (2026)
di: Li, Cong, et al.
Pubblicazione: (2026)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
di: You, Dean, et al.
Pubblicazione: (2025)
di: You, Dean, et al.
Pubblicazione: (2025)
ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration
di: Lin, Wei-Fen, et al.
Pubblicazione: (2026)
di: Lin, Wei-Fen, et al.
Pubblicazione: (2026)
A Vertically Integrated Framework for Templatized Chip Design
di: Kim, Jeongeun, et al.
Pubblicazione: (2025)
di: Kim, Jeongeun, et al.
Pubblicazione: (2025)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
di: Hong, Junguk, et al.
Pubblicazione: (2026)
di: Hong, Junguk, et al.
Pubblicazione: (2026)
RFAmpDesigner: A Self-Evolving Multi-Agent LLM Framework for Automated Radio Frequency Amplifier Design
di: Lu, Hang, et al.
Pubblicazione: (2026)
di: Lu, Hang, et al.
Pubblicazione: (2026)
AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server Applications
di: Yahya, Jawad Haj, et al.
Pubblicazione: (2022)
di: Yahya, Jawad Haj, et al.
Pubblicazione: (2022)
RealProbe: An Automated and Lightweight Performance Profiler for In-FPGA Execution of High-Level Synthesis Designs
di: Kim, Jiho, et al.
Pubblicazione: (2025)
di: Kim, Jiho, et al.
Pubblicazione: (2025)
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024)
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024)
Lightweight Congruence Profiling for Early Design Exploration of Heterogeneous FPGAs
di: Boston, Allen, et al.
Pubblicazione: (2025)
di: Boston, Allen, et al.
Pubblicazione: (2025)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
di: Kwon, Hyucksung, et al.
Pubblicazione: (2024)
di: Kwon, Hyucksung, et al.
Pubblicazione: (2024)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
di: Seo, Minseok, et al.
Pubblicazione: (2024)
di: Seo, Minseok, et al.
Pubblicazione: (2024)
Cool-3D: An End-to-End Thermal-Aware Framework for Early-Phase Design Space Exploration of Microfluidic-Cooled 3DICs
di: Wang, Runxi, et al.
Pubblicazione: (2025)
di: Wang, Runxi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
di: Kim, Minsu, et al.
Pubblicazione: (2025) -
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
di: Moon, Seungjae, et al.
Pubblicazione: (2024) -
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
di: Hong, Jeongmin, et al.
Pubblicazione: (2024) -
MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization
di: Kim, Seeyeon, et al.
Pubblicazione: (2026) -
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
di: Kim, Jae-Young, et al.
Pubblicazione: (2025)