Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Joon Ha, Kim, Geon-Woo, Rachakonda, Anoop, Kim, Daehyeok |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Vulcan: Instance-Optimal Systems Heuristics Through LLM-Driven Search
par: Dwivedula, Rohit, et autres
Publié: (2025)
par: Dwivedula, Rohit, et autres
Publié: (2025)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
par: Cho, Jaehong, et autres
Publié: (2024)
par: Cho, Jaehong, et autres
Publié: (2024)
Large Language Models as Realistic Microservice Trace Generators
par: Kim, Donghyun, et autres
Publié: (2024)
par: Kim, Donghyun, et autres
Publié: (2024)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
par: Kim, Geon-Woo, et autres
Publié: (2025)
par: Kim, Geon-Woo, et autres
Publié: (2025)
Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
par: Sung, Mingyu, et autres
Publié: (2025)
par: Sung, Mingyu, et autres
Publié: (2025)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
par: Lyu, Hongtao, et autres
Publié: (2025)
par: Lyu, Hongtao, et autres
Publié: (2025)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
par: Stojkovic, Jovan, et autres
Publié: (2025)
par: Stojkovic, Jovan, et autres
Publié: (2025)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
par: Rhee, Myunghyun, et autres
Publié: (2025)
par: Rhee, Myunghyun, et autres
Publié: (2025)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
par: Argerich, Mauricio Fadel, et autres
Publié: (2026)
par: Argerich, Mauricio Fadel, et autres
Publié: (2026)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
par: Jeong, Bodon, et autres
Publié: (2026)
par: Jeong, Bodon, et autres
Publié: (2026)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
par: Qazi, Muhammad Azlan, et autres
Publié: (2026)
par: Qazi, Muhammad Azlan, et autres
Publié: (2026)
Accelerating LLM Inference with Precomputed Query Storage
par: Park, Jay H., et autres
Publié: (2025)
par: Park, Jay H., et autres
Publié: (2025)
LLM Inference Serving: Survey of Recent Advances and Opportunities
par: Li, Baolin, et autres
Publié: (2024)
par: Li, Baolin, et autres
Publié: (2024)
Decentralized AI: Permissionless LLM Inference on POKT Network
par: Olshansky, Daniel, et autres
Publié: (2024)
par: Olshansky, Daniel, et autres
Publié: (2024)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
par: Deng, Xiumei, et autres
Publié: (2025)
par: Deng, Xiumei, et autres
Publié: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
par: Li, Rongzhi, et autres
Publié: (2025)
par: Li, Rongzhi, et autres
Publié: (2025)
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
par: Yuan, Yichao, et autres
Publié: (2026)
par: Yuan, Yichao, et autres
Publié: (2026)
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
par: Chandrasekar, Ashok, et autres
Publié: (2026)
par: Chandrasekar, Ashok, et autres
Publié: (2026)
Seesaw: High-throughput LLM Inference via Model Re-sharding
par: Su, Qidong, et autres
Publié: (2025)
par: Su, Qidong, et autres
Publié: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
par: Wu, Linyu, et autres
Publié: (2025)
par: Wu, Linyu, et autres
Publié: (2025)
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
par: Feng, Yicheng, et autres
Publié: (2026)
par: Feng, Yicheng, et autres
Publié: (2026)
Frontier: Simulating the Next Generation of LLM Inference Systems
par: Feng, Yicheng, et autres
Publié: (2025)
par: Feng, Yicheng, et autres
Publié: (2025)
AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving
par: Xu, Tianhao, et autres
Publié: (2026)
par: Xu, Tianhao, et autres
Publié: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
par: Xie, Jincheng, et autres
Publié: (2026)
par: Xie, Jincheng, et autres
Publié: (2026)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
par: Song, Jingwei, et autres
Publié: (2025)
par: Song, Jingwei, et autres
Publié: (2025)
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
par: Wilkins, Grant, et autres
Publié: (2024)
par: Wilkins, Grant, et autres
Publié: (2024)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
par: Liu, Xing, et autres
Publié: (2025)
par: Liu, Xing, et autres
Publié: (2025)
LAPS: A Length-Aware-Prefill LLM Serving System
par: She, Jianshu, et autres
Publié: (2026)
par: She, Jianshu, et autres
Publié: (2026)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
par: Hankendi, Can, et autres
Publié: (2026)
par: Hankendi, Can, et autres
Publié: (2026)
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
par: Hu, Bodun, et autres
Publié: (2024)
par: Hu, Bodun, et autres
Publié: (2024)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
par: Wang, Shaoyu, et autres
Publié: (2025)
par: Wang, Shaoyu, et autres
Publié: (2025)
Why Do AI Agents Systematically Fail at Cloud Root Cause Analysis?
par: Kim, Taeyoon, et autres
Publié: (2026)
par: Kim, Taeyoon, et autres
Publié: (2026)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
par: Li, Wanqian, et autres
Publié: (2026)
par: Li, Wanqian, et autres
Publié: (2026)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
par: Zheng, Wanyi, et autres
Publié: (2025)
par: Zheng, Wanyi, et autres
Publié: (2025)
Multi-IaC-Eval: Benchmarking Cloud Infrastructure as Code Across Multiple Formats
par: Davidson, Sam, et autres
Publié: (2025)
par: Davidson, Sam, et autres
Publié: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
par: Liu, Dong, et autres
Publié: (2025)
par: Liu, Dong, et autres
Publié: (2025)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
par: Oh, Hyungjun, et autres
Publié: (2024)
par: Oh, Hyungjun, et autres
Publié: (2024)
Compare Where It Matters: Using Layer-Wise Regularization To Improve Federated Learning on Heterogeneous Data
par: Son, Ha Min, et autres
Publié: (2021)
par: Son, Ha Min, et autres
Publié: (2021)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
par: Ramachandran, Arun, et autres
Publié: (2025)
par: Ramachandran, Arun, et autres
Publié: (2025)
Documents similaires
-
Vulcan: Instance-Optimal Systems Heuristics Through LLM-Driven Search
par: Dwivedula, Rohit, et autres
Publié: (2025) -
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
par: Cho, Jaehong, et autres
Publié: (2024) -
Large Language Models as Realistic Microservice Trace Generators
par: Kim, Donghyun, et autres
Publié: (2024) -
OMEGA: A Low-Latency GNN Serving System for Large Graphs
par: Kim, Geon-Woo, et autres
Publié: (2025) -
Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
par: Sung, Mingyu, et autres
Publié: (2025)