Plato: Plan to Efficiently Decode for Large Language Model Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Jin, Shuowei, Liu, Xueshen, Wu, Yongji, Zheng, Haizhong, Zhang, Qingzhao, Prakash, Atul, Lentz, Matthew, Zhuo, Danyang, Qian, Feng, Mao, Z. Morley |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
por: Wu, Yongji, et al.
Publicado: (2025)
por: Wu, Yongji, et al.
Publicado: (2025)
Compute Or Load KV Cache? Why Not Both?
por: Jin, Shuowei, et al.
Publicado: (2024)
por: Jin, Shuowei, et al.
Publicado: (2024)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
por: Zheng, Haizhong, et al.
Publicado: (2024)
por: Zheng, Haizhong, et al.
Publicado: (2024)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
por: Liu, Xueshen, et al.
Publicado: (2026)
por: Liu, Xueshen, et al.
Publicado: (2026)
Eagle: Efficient Training-Free Router for Multi-LLM Inference
por: Zhao, Zesen, et al.
Publicado: (2024)
por: Zhao, Zesen, et al.
Publicado: (2024)
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs
por: Zheng, Haizhong, et al.
Publicado: (2026)
por: Zheng, Haizhong, et al.
Publicado: (2026)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
por: Wu, Yongji, et al.
Publicado: (2025)
por: Wu, Yongji, et al.
Publicado: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
por: Wu, Yongji, et al.
Publicado: (2024)
por: Wu, Yongji, et al.
Publicado: (2024)
From Stealthy Data Fabrication to Unsafe Driving: Realistic Scenario Attacks on Collaborative Perception
por: Zhang, Qingzhao, et al.
Publicado: (2026)
por: Zhang, Qingzhao, et al.
Publicado: (2026)
LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks
por: Zhang, Qingzhao, et al.
Publicado: (2024)
por: Zhang, Qingzhao, et al.
Publicado: (2024)
Class-Proportional Coreset Selection for Difficulty-Separable Data
por: Tsai, Elisa, et al.
Publicado: (2025)
por: Tsai, Elisa, et al.
Publicado: (2025)
Hydra: Efficient, Correct Code Generation via Checkpoint-and-Rollback Support
por: Du, Alexander, et al.
Publicado: (2026)
por: Du, Alexander, et al.
Publicado: (2026)
VcLLM: Video Codecs are Secretly Tensor Codecs
por: Xu, Ceyu, et al.
Publicado: (2024)
por: Xu, Ceyu, et al.
Publicado: (2024)
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
por: Hu, Pingbang, et al.
Publicado: (2026)
por: Hu, Pingbang, et al.
Publicado: (2026)
Leveraging Hierarchical Feature Sharing for Efficient Dataset Condensation
por: Zheng, Haizhong, et al.
Publicado: (2023)
por: Zheng, Haizhong, et al.
Publicado: (2023)
QUIC is not Quick Enough over Fast Internet
por: Zhang, Xumiao, et al.
Publicado: (2023)
por: Zhang, Xumiao, et al.
Publicado: (2023)
Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale
por: Tsai, Elisa, et al.
Publicado: (2025)
por: Tsai, Elisa, et al.
Publicado: (2025)
Curator: Efficient Indexing for Multi-Tenant Vector Databases
por: Jin, Yicheng, et al.
Publicado: (2024)
por: Jin, Yicheng, et al.
Publicado: (2024)
AutoSpec: Automated Generation of Neural Network Specifications
por: Jin, Shuowei, et al.
Publicado: (2024)
por: Jin, Shuowei, et al.
Publicado: (2024)
What Would Trojans Do? Exploiting Partial-Information Vulnerabilities in Autonomous Vehicle Sensing
por: Hallyburton, R. Spencer, et al.
Publicado: (2023)
por: Hallyburton, R. Spencer, et al.
Publicado: (2023)
SoK: How Sensor Attacks Disrupt Autonomous Vehicles: An End-to-end Analysis, Challenges, and Missed Threats
por: Zhang, Qingzhao, et al.
Publicado: (2025)
por: Zhang, Qingzhao, et al.
Publicado: (2025)
Curator: Efficient Vector Search with Low-Selectivity Filters
por: Jin, Yicheng, et al.
Publicado: (2026)
por: Jin, Yicheng, et al.
Publicado: (2026)
CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving
por: Zhu, Ruiyang, et al.
Publicado: (2026)
por: Zhu, Ruiyang, et al.
Publicado: (2026)
Automatic Teller Machines for Offline E-cash
por: Chakraborti, Anrin, et al.
Publicado: (2026)
por: Chakraborti, Anrin, et al.
Publicado: (2026)
Nonparametric Inference for Extreme CoVaR and CoES
por: Zhong, Qingzhao, et al.
Publicado: (2025)
por: Zhong, Qingzhao, et al.
Publicado: (2025)
MARS: Harmonizing Multimodal Convergence via Adaptive Rank Search
por: Cho, Minkyoung, et al.
Publicado: (2026)
por: Cho, Minkyoung, et al.
Publicado: (2026)
SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass
por: Qian, Chen, et al.
Publicado: (2026)
por: Qian, Chen, et al.
Publicado: (2026)
ELFS: Label-Free Coreset Selection with Proxy Training Dynamics
por: Zheng, Haizhong, et al.
Publicado: (2024)
por: Zheng, Haizhong, et al.
Publicado: (2024)
A Two-Stage Proactive Dialogue Generator for Efficient Clinical Information Collection Using Large Language Model
por: Li, Xueshen, et al.
Publicado: (2024)
por: Li, Xueshen, et al.
Publicado: (2024)
CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
por: Liu, Dong, et al.
Publicado: (2025)
por: Liu, Dong, et al.
Publicado: (2025)
Estimations of Extreme CoVaR and CoES under Asymptotic Independence
por: Zhong, Qingzhao
Publicado: (2026)
por: Zhong, Qingzhao
Publicado: (2026)
Cocoon: Robust Multi-Modal Perception with Uncertainty-Aware Sensor Fusion
por: Cho, Minkyoung, et al.
Publicado: (2024)
por: Cho, Minkyoung, et al.
Publicado: (2024)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
por: Zheng, Wenhao, et al.
Publicado: (2025)
por: Zheng, Wenhao, et al.
Publicado: (2025)
FlashDecoding++: Faster Large Language Model Inference on GPUs
por: Hong, Ke, et al.
Publicado: (2023)
por: Hong, Ke, et al.
Publicado: (2023)
How Large Language Models Balance Internal Knowledge with User and Document Assertions
por: Li, Shuowei, et al.
Publicado: (2026)
por: Li, Shuowei, et al.
Publicado: (2026)
DecodeX: Exploring and Benchmarking of LDPC Decoding across CPU, GPU, and ASIC Platforms
por: Qi, Zhenzhou, et al.
Publicado: (2025)
por: Qi, Zhenzhou, et al.
Publicado: (2025)
Language and dialogue in Plato
por: SAMUEL SCOLNICOV
Publicado: (2006)
por: SAMUEL SCOLNICOV
Publicado: (2006)
Prompt Inversion Attack against Collaborative Inference of Large Language Models
por: Qu, Wenjie, et al.
Publicado: (2025)
por: Qu, Wenjie, et al.
Publicado: (2025)
Flat Rotation Curves from Vacuum Pressure in a Self-Gravitating Nonlinear Scalar Field
por: An, Haizhong
Publicado: (2026)
por: An, Haizhong
Publicado: (2026)
Structural Conjectures for Particle Topology and Black Hole Morphology in a Discrete Field-Quantum Network
por: An, Haizhong
Publicado: (2026)
por: An, Haizhong
Publicado: (2026)
Ejemplares similares
-
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
por: Wu, Yongji, et al.
Publicado: (2025) -
Compute Or Load KV Cache? Why Not Both?
por: Jin, Shuowei, et al.
Publicado: (2024) -
Learn To be Efficient: Build Structured Sparsity in Large Language Models
por: Zheng, Haizhong, et al.
Publicado: (2024) -
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
por: Liu, Xueshen, et al.
Publicado: (2026) -
Eagle: Efficient Training-Free Router for Multi-LLM Inference
por: Zhao, Zesen, et al.
Publicado: (2024)