PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Cho, Eunyeong, Bang, Jehyeon, Hwang, Ranggi, Rhu, Minsoo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
vTrain: A Simulation Framework for Evaluating Cost-effective and Compute-optimal Large Language Model Training
di: Bang, Jehyeon, et al.
Pubblicazione: (2023)
di: Bang, Jehyeon, et al.
Pubblicazione: (2023)
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
di: Bang, Jehyeon, et al.
Pubblicazione: (2026)
di: Bang, Jehyeon, et al.
Pubblicazione: (2026)
The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
di: Kim, Jiin, et al.
Pubblicazione: (2025)
di: Kim, Jiin, et al.
Pubblicazione: (2025)
PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
Debunking the CUDA Myth Towards GPU-based AI Systems
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
di: Lee, Jungi, et al.
Pubblicazione: (2025)
di: Lee, Jungi, et al.
Pubblicazione: (2025)
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
di: Hwang, Ranggi, et al.
Pubblicazione: (2023)
di: Hwang, Ranggi, et al.
Pubblicazione: (2023)
Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
di: Kim, Wonung, et al.
Pubblicazione: (2025)
di: Kim, Wonung, et al.
Pubblicazione: (2025)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
di: Yeo, Gwangoo, et al.
Pubblicazione: (2024)
di: Yeo, Gwangoo, et al.
Pubblicazione: (2024)
Llumnix: Dynamic Scheduling for Large Language Model Serving
di: Sun, Biao, et al.
Pubblicazione: (2024)
di: Sun, Biao, et al.
Pubblicazione: (2024)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
di: Lee, Dongjae, et al.
Pubblicazione: (2025)
di: Lee, Dongjae, et al.
Pubblicazione: (2025)
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
di: Lee, Dongjae, et al.
Pubblicazione: (2024)
di: Lee, Dongjae, et al.
Pubblicazione: (2024)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
di: Hyun, Bongjoon, et al.
Pubblicazione: (2023)
di: Hyun, Bongjoon, et al.
Pubblicazione: (2023)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
di: Killian, Earl
Pubblicazione: (2026)
di: Killian, Earl
Pubblicazione: (2026)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
di: Xia, Haojun, et al.
Pubblicazione: (2024)
di: Xia, Haojun, et al.
Pubblicazione: (2024)
Unveiling Environmental Impacts of Large Language Model Serving: A Functional Unit View
di: Wu, Yanran, et al.
Pubblicazione: (2025)
di: Wu, Yanran, et al.
Pubblicazione: (2025)
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
di: Yun, Sungmin, et al.
Pubblicazione: (2024)
di: Yun, Sungmin, et al.
Pubblicazione: (2024)
A Data-Driven Approach to Dataflow-Aware Online Scheduling for Graph Neural Network Inference
di: Puigdemont, Pol, et al.
Pubblicazione: (2024)
di: Puigdemont, Pol, et al.
Pubblicazione: (2024)
LLM-USO: Large Language Model-based Universal Sizing Optimizer
di: S, Karthik Somayaji N., et al.
Pubblicazione: (2025)
di: S, Karthik Somayaji N., et al.
Pubblicazione: (2025)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
di: Yoon, Dongho, et al.
Pubblicazione: (2025)
di: Yoon, Dongho, et al.
Pubblicazione: (2025)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
di: Li, Jinhao, et al.
Pubblicazione: (2024)
di: Li, Jinhao, et al.
Pubblicazione: (2024)
Accelerating String-Key Learned Index Structures via Memoization-based Incremental Training
di: Kim, Minsu, et al.
Pubblicazione: (2024)
di: Kim, Minsu, et al.
Pubblicazione: (2024)
COMET: Towards Partical W4A4KV4 LLMs Serving
di: Liu, Lian, et al.
Pubblicazione: (2024)
di: Liu, Lian, et al.
Pubblicazione: (2024)
From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
di: Zhu, Tianhao, et al.
Pubblicazione: (2025)
di: Zhu, Tianhao, et al.
Pubblicazione: (2025)
GFormer: Accelerating Large Language Models with Optimized Transformers on Gaudi Processors
di: Zhang, Chengming, et al.
Pubblicazione: (2024)
di: Zhang, Chengming, et al.
Pubblicazione: (2024)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
di: Kim, Minsu, et al.
Pubblicazione: (2025)
di: Kim, Minsu, et al.
Pubblicazione: (2025)
SNIP: An Adaptive Mixed Precision Framework for Subbyte Large Language Model Training
di: Pan, Yunjie, et al.
Pubblicazione: (2026)
di: Pan, Yunjie, et al.
Pubblicazione: (2026)
AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
di: Sharma, Amit
Pubblicazione: (2025)
di: Sharma, Amit
Pubblicazione: (2025)
Accelerating Neural Networks for Large Language Models and Graph Processing with Silicon Photonics
di: Afifi, Salma, et al.
Pubblicazione: (2024)
di: Afifi, Salma, et al.
Pubblicazione: (2024)
Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization
di: Lee, Jungi, et al.
Pubblicazione: (2024)
di: Lee, Jungi, et al.
Pubblicazione: (2024)
GauS: Differentiable Scheduling Optimization via Gaussian Reparameterization
di: Cai, Yaohui, et al.
Pubblicazione: (2026)
di: Cai, Yaohui, et al.
Pubblicazione: (2026)
Memory Access Characterization of Large Language Models in CPU Environment and its Potential Impacts
di: Banasik, Spencer
Pubblicazione: (2025)
di: Banasik, Spencer
Pubblicazione: (2025)
LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation
di: Zhang, Zixi, et al.
Pubblicazione: (2023)
di: Zhang, Zixi, et al.
Pubblicazione: (2023)
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
di: Titopoulos, Vasileios, et al.
Pubblicazione: (2025)
di: Titopoulos, Vasileios, et al.
Pubblicazione: (2025)
TorR: Towards Brain-Inspired Task-Oriented Reasoning via Cache-Oriented Algorithm-Architecture Co-design
di: Oh, Hyunwoo, et al.
Pubblicazione: (2026)
di: Oh, Hyunwoo, et al.
Pubblicazione: (2026)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
di: Ding, Jianru, et al.
Pubblicazione: (2026)
di: Ding, Jianru, et al.
Pubblicazione: (2026)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
di: Wolters, Christopher, et al.
Pubblicazione: (2024)
di: Wolters, Christopher, et al.
Pubblicazione: (2024)
Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification
di: Pinckney, Nathaniel, et al.
Pubblicazione: (2025)
di: Pinckney, Nathaniel, et al.
Pubblicazione: (2025)
Energy Efficient Software Hardware CoDesign for Machine Learning: From TinyML to Large Language Models
di: Vahdatpour, Mohammad Saleh, et al.
Pubblicazione: (2026)
di: Vahdatpour, Mohammad Saleh, et al.
Pubblicazione: (2026)
Documenti analoghi
-
vTrain: A Simulation Framework for Evaluating Cost-effective and Compute-optimal Large Language Model Training
di: Bang, Jehyeon, et al.
Pubblicazione: (2023) -
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
di: Bang, Jehyeon, et al.
Pubblicazione: (2026) -
The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
di: Kim, Jiin, et al.
Pubblicazione: (2025) -
PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
di: Lee, Yunjae, et al.
Pubblicazione: (2024) -
Debunking the CUDA Myth Towards GPU-based AI Systems
di: Lee, Yunjae, et al.
Pubblicazione: (2024)