Fast Inference for Augmented Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shahout, Rana, Liang, Cong, Xin, Shiji, Lao, Qianru, Cui, Yong, Yu, Minlan, Mitzenmacher, Michael |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
par: Shahout, Rana, et autres
Publié: (2025)
par: Shahout, Rana, et autres
Publié: (2025)
Orla: A Library for Serving LLM-Based Multi-Agent Systems
par: Shahout, Rana, et autres
Publié: (2026)
par: Shahout, Rana, et autres
Publié: (2026)
Queueing, Predictions, and LLMs: Challenges and Open Problems
par: Mitzenmacher, Michael, et autres
Publié: (2025)
par: Mitzenmacher, Michael, et autres
Publié: (2025)
Intra-request branch orchestration for efficient LLM reasoning
par: Jiang, Weifan, et autres
Publié: (2025)
par: Jiang, Weifan, et autres
Publié: (2025)
Learning-Based Heavy Hitters and Flow Frequency Estimation in Streams
par: Shahout, Rana, et autres
Publié: (2024)
par: Shahout, Rana, et autres
Publié: (2024)
Learning-Augmented Frequency Estimation in Sliding Windows
par: Shahout, Rana, et autres
Publié: (2024)
par: Shahout, Rana, et autres
Publié: (2024)
SkipPredict: When to Invest in Predictions for Scheduling
par: Shahout, Rana, et autres
Publié: (2024)
par: Shahout, Rana, et autres
Publié: (2024)
Don't Stop Me Now: Embedding Based Scheduling for LLMs
par: Shahout, Rana, et autres
Publié: (2024)
par: Shahout, Rana, et autres
Publié: (2024)
THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
par: Li, Minghao, et autres
Publié: (2023)
par: Li, Minghao, et autres
Publié: (2023)
Predictive Scheduling for Efficient Inference-Time Reasoning in Large Language Models
par: Brown, Katrina, et autres
Publié: (2026)
par: Brown, Katrina, et autres
Publié: (2026)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
par: Hankendi, Can, et autres
Publié: (2026)
par: Hankendi, Can, et autres
Publié: (2026)
EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
par: Lao, ChonLam, et autres
Publié: (2024)
par: Lao, ChonLam, et autres
Publié: (2024)
Learning Multimodal Energy-Based Model with Multimodal Variational Auto-Encoder via MCMC Revision
par: Cui, Jiali, et autres
Publié: (2026)
par: Cui, Jiali, et autres
Publié: (2026)
DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
par: Yue, Yang, et autres
Publié: (2024)
par: Yue, Yang, et autres
Publié: (2024)
Conda: Column-Normalized Adam for Training Large Language Models Faster
par: Wang, Junjie, et autres
Publié: (2025)
par: Wang, Junjie, et autres
Publié: (2025)
FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
par: Bajpai, Divya Jyoti, et autres
Publié: (2025)
par: Bajpai, Divya Jyoti, et autres
Publié: (2025)
SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL Workloads
par: Lao, Jiale, et autres
Publié: (2025)
par: Lao, Jiale, et autres
Publié: (2025)
Model-Distributed Inference for Large Language Models at the Edge
par: Macario, Davide, et autres
Publié: (2025)
par: Macario, Davide, et autres
Publié: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
par: Jiang, Xuanlin, et autres
Publié: (2024)
par: Jiang, Xuanlin, et autres
Publié: (2024)
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
par: Chen, Sihan, et autres
Publié: (2025)
par: Chen, Sihan, et autres
Publié: (2025)
An LLM-based Agentic Framework for Accessible Network Control
par: Lin, Samuel, et autres
Publié: (2025)
par: Lin, Samuel, et autres
Publié: (2025)
Over-Searching in Search-Augmented Large Language Models
par: Xie, Roy, et autres
Publié: (2026)
par: Xie, Roy, et autres
Publié: (2026)
Federated Learning Clients Clustering with Adaptation to Data Drifts
par: Li, Minghao, et autres
Publié: (2024)
par: Li, Minghao, et autres
Publié: (2024)
Fast Large Language Model Collaborative Decoding via Speculation
par: Fu, Jiale, et autres
Publié: (2025)
par: Fu, Jiale, et autres
Publié: (2025)
Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees
par: Chen, Sijia, et autres
Publié: (2024)
par: Chen, Sijia, et autres
Publié: (2024)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
par: Lin, Huawei, et autres
Publié: (2025)
par: Lin, Huawei, et autres
Publié: (2025)
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
par: Schmied, Thomas, et autres
Publié: (2024)
par: Schmied, Thomas, et autres
Publié: (2024)
Democratizing Large Language Model-Based Graph Data Augmentation via Latent Knowledge Graphs
par: Feng, Yushi, et autres
Publié: (2025)
par: Feng, Yushi, et autres
Publié: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
par: Erdogan, Mete, et autres
Publié: (2025)
par: Erdogan, Mete, et autres
Publié: (2025)
Automatic Calibration for Membership Inference Attack on Large Language Models
par: Zade, Saleh Zare, et autres
Publié: (2025)
par: Zade, Saleh Zare, et autres
Publié: (2025)
Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate
par: Fang, Zhiyuan, et autres
Publié: (2025)
par: Fang, Zhiyuan, et autres
Publié: (2025)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
par: Ho, Namgyu, et autres
Publié: (2024)
par: Ho, Namgyu, et autres
Publié: (2024)
SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation
par: Ren, Yanwei, et autres
Publié: (2025)
par: Ren, Yanwei, et autres
Publié: (2025)
The Shape of Wisdom: Decision Trajectories in Language Models
par: Rana, Shailesh
Publié: (2026)
par: Rana, Shailesh
Publié: (2026)
Consistency Models for Scalable and Fast Simulation-Based Inference
par: Schmitt, Marvin, et autres
Publié: (2023)
par: Schmitt, Marvin, et autres
Publié: (2023)
TimeCAP: Learning to Contextualize, Augment, and Predict Time Series Events with Large Language Model Agents
par: Lee, Geon, et autres
Publié: (2025)
par: Lee, Geon, et autres
Publié: (2025)
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
par: Cai, Yuxuan, et autres
Publié: (2025)
par: Cai, Yuxuan, et autres
Publié: (2025)
Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
par: Sun, Yifan, et autres
Publié: (2025)
par: Sun, Yifan, et autres
Publié: (2025)
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
par: Li, Junsong, et autres
Publié: (2025)
par: Li, Junsong, et autres
Publié: (2025)
Do Language Models Have Bayesian Brains? Distinguishing Stochastic and Deterministic Decision Patterns within Large Language Models
par: Cui, Andrea Yaoyun, et autres
Publié: (2025)
par: Cui, Andrea Yaoyun, et autres
Publié: (2025)
Documents similaires
-
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
par: Shahout, Rana, et autres
Publié: (2025) -
Orla: A Library for Serving LLM-Based Multi-Agent Systems
par: Shahout, Rana, et autres
Publié: (2026) -
Queueing, Predictions, and LLMs: Challenges and Open Problems
par: Mitzenmacher, Michael, et autres
Publié: (2025) -
Intra-request branch orchestration for efficient LLM reasoning
par: Jiang, Weifan, et autres
Publié: (2025) -
Learning-Based Heavy Hitters and Flow Frequency Estimation in Streams
par: Shahout, Rana, et autres
Publié: (2024)