Vidur: A Large-Scale Simulation Framework For LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Agrawal, Amey, Kedia, Nitin, Mohan, Jayashree, Panwar, Ashish, Kwatra, Nipun, Gulavani, Bhargav, Ramjee, Ramachandran, Tumanov, Alexey |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
On Evaluating Performance of LLM Inference Serving Systems
by: Agrawal, Amey, et al.
Published: (2025)
by: Agrawal, Amey, et al.
Published: (2025)
Niyama : Breaking the Silos of LLM Inference Serving
by: Goel, Kanishk, et al.
Published: (2025)
by: Goel, Kanishk, et al.
Published: (2025)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
by: Deshmukh, Dhruv, et al.
Published: (2025)
by: Deshmukh, Dhruv, et al.
Published: (2025)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
by: Gond, Raja, et al.
Published: (2025)
by: Gond, Raja, et al.
Published: (2025)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
by: Gond, Raja, et al.
Published: (2026)
by: Gond, Raja, et al.
Published: (2026)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
by: Kamath, Aditya K, et al.
Published: (2024)
by: Kamath, Aditya K, et al.
Published: (2024)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024)
by: Prabhu, Ramya, et al.
Published: (2024)
Accuracy is Not All You Need
by: Dutta, Abhinav, et al.
Published: (2024)
by: Dutta, Abhinav, et al.
Published: (2024)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models
by: Ramjee, Sharan
Published: (2026)
by: Ramjee, Sharan
Published: (2026)
Boosting Zero-Shot Crosslingual Performance using LLM-Based Augmentations with Effective Data Selection
by: Fazili, Barah, et al.
Published: (2024)
by: Fazili, Barah, et al.
Published: (2024)
ASTRA: Accurate and Scalable ANNS-based Training of Extreme Classifiers
by: Mehta, Sonu, et al.
Published: (2024)
by: Mehta, Sonu, et al.
Published: (2024)
Boosting the Capabilities of Compact Models in Low-Data Contexts with Large Language Models and Retrieval-Augmented Generation
by: Shandilya, Bhargav, et al.
Published: (2024)
by: Shandilya, Bhargav, et al.
Published: (2024)
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference
by: Yang, Mengtian, et al.
Published: (2026)
by: Yang, Mengtian, et al.
Published: (2026)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
by: Behnam, Payman, et al.
Published: (2025)
by: Behnam, Payman, et al.
Published: (2025)
Initialization using Update Approximation is a Silver Bullet for Extremely Efficient Low-Rank Fine-Tuning
by: Ponkshe, Kaustubh, et al.
Published: (2024)
by: Ponkshe, Kaustubh, et al.
Published: (2024)
PLUM: Improving Inference Efficiency By Leveraging Repetition-Sparsity Trade-Off
by: Kuhar, Sachit, et al.
Published: (2023)
by: Kuhar, Sachit, et al.
Published: (2023)
Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
by: Agrawal, Vishakha, et al.
Published: (2025)
by: Agrawal, Vishakha, et al.
Published: (2025)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
by: Biswas, Anish, et al.
Published: (2026)
by: Biswas, Anish, et al.
Published: (2026)
Scaling LLM Inference with Optimized Sample Compute Allocation
by: Zhang, Kexun, et al.
Published: (2024)
by: Zhang, Kexun, et al.
Published: (2024)
Building pre-train LLM Dataset for the INDIC Languages: a case study on Hindi
by: Parida, Shantipriya, et al.
Published: (2024)
by: Parida, Shantipriya, et al.
Published: (2024)
ConSCompF: Consistency-focused Similarity Comparison Framework for Generative Large Language Models
by: Karev, Alexey, et al.
Published: (2025)
by: Karev, Alexey, et al.
Published: (2025)
CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation
by: Bougie, Nicolas, et al.
Published: (2025)
by: Bougie, Nicolas, et al.
Published: (2025)
"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models
by: Tao, Yufei, et al.
Published: (2025)
by: Tao, Yufei, et al.
Published: (2025)
Simul-LLM: A Framework for Exploring High-Quality Simultaneous Translation with Large Language Models
by: Agostinelli, Victor, et al.
Published: (2023)
by: Agostinelli, Victor, et al.
Published: (2023)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
Query-Efficient Planning with Language Models
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2024)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2024)
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
by: Ouyang, Haojie, et al.
Published: (2025)
by: Ouyang, Haojie, et al.
Published: (2025)
An Investigation of Linguistic Biases in LLM-Based Recommendations
by: Venkateswaran, Nitin, et al.
Published: (2026)
by: Venkateswaran, Nitin, et al.
Published: (2026)
A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
by: Ke, Zixuan, et al.
Published: (2025)
by: Ke, Zixuan, et al.
Published: (2025)
Addressing LLM Diversity by Infusing Random Concepts
by: Agrawal, Pulin, et al.
Published: (2026)
by: Agrawal, Pulin, et al.
Published: (2026)
Spoken Grammar Assessment Using LLM
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
Bullying the Machine: How Personas Increase LLM Vulnerability
by: Xu, Ziwei, et al.
Published: (2025)
by: Xu, Ziwei, et al.
Published: (2025)
Window-based Membership Inference Attacks Against Fine-tuned Large Language Models
by: Chen, Yuetian, et al.
Published: (2026)
by: Chen, Yuetian, et al.
Published: (2026)
Construction of Hyper-Relational Knowledge Graphs Using Pre-Trained Large Language Models
by: Datta, Preetha, et al.
Published: (2024)
by: Datta, Preetha, et al.
Published: (2024)
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
by: Şenol, Ali, et al.
Published: (2026)
by: Şenol, Ali, et al.
Published: (2026)
The Impact of Model Scaling on Seen and Unseen Language Performance
by: Pokharel, Rhitabrat, et al.
Published: (2025)
by: Pokharel, Rhitabrat, et al.
Published: (2025)
An Efficient Inference Framework for Early-exit Large Language Models
by: Miao, Ruijie, et al.
Published: (2024)
by: Miao, Ruijie, et al.
Published: (2024)
Similar Items
-
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
by: Agrawal, Amey, et al.
Published: (2024) -
Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
by: Agrawal, Amey, et al.
Published: (2024) -
On Evaluating Performance of LLM Inference Serving Systems
by: Agrawal, Amey, et al.
Published: (2025) -
Niyama : Breaking the Silos of LLM Inference Serving
by: Goel, Kanishk, et al.
Published: (2025) -
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
by: Deshmukh, Dhruv, et al.
Published: (2025)