Fast Inference for Augmented Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shahout, Rana, Liang, Cong, Xin, Shiji, Lao, Qianru, Cui, Yong, Yu, Minlan, Mitzenmacher, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
Orla: A Library for Serving LLM-Based Multi-Agent Systems
von: Shahout, Rana, et al.
Veröffentlicht: (2026)
von: Shahout, Rana, et al.
Veröffentlicht: (2026)
Queueing, Predictions, and LLMs: Challenges and Open Problems
von: Mitzenmacher, Michael, et al.
Veröffentlicht: (2025)
von: Mitzenmacher, Michael, et al.
Veröffentlicht: (2025)
Intra-request branch orchestration for efficient LLM reasoning
von: Jiang, Weifan, et al.
Veröffentlicht: (2025)
von: Jiang, Weifan, et al.
Veröffentlicht: (2025)
Learning-Based Heavy Hitters and Flow Frequency Estimation in Streams
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
Learning-Augmented Frequency Estimation in Sliding Windows
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
SkipPredict: When to Invest in Predictions for Scheduling
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
Don't Stop Me Now: Embedding Based Scheduling for LLMs
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
von: Li, Minghao, et al.
Veröffentlicht: (2023)
von: Li, Minghao, et al.
Veröffentlicht: (2023)
Predictive Scheduling for Efficient Inference-Time Reasoning in Large Language Models
von: Brown, Katrina, et al.
Veröffentlicht: (2026)
von: Brown, Katrina, et al.
Veröffentlicht: (2026)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
von: Lao, ChonLam, et al.
Veröffentlicht: (2024)
von: Lao, ChonLam, et al.
Veröffentlicht: (2024)
Learning Multimodal Energy-Based Model with Multimodal Variational Auto-Encoder via MCMC Revision
von: Cui, Jiali, et al.
Veröffentlicht: (2026)
von: Cui, Jiali, et al.
Veröffentlicht: (2026)
DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
von: Yue, Yang, et al.
Veröffentlicht: (2024)
von: Yue, Yang, et al.
Veröffentlicht: (2024)
Conda: Column-Normalized Adam for Training Large Language Models Faster
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL Workloads
von: Lao, Jiale, et al.
Veröffentlicht: (2025)
von: Lao, Jiale, et al.
Veröffentlicht: (2025)
Model-Distributed Inference for Large Language Models at the Edge
von: Macario, Davide, et al.
Veröffentlicht: (2025)
von: Macario, Davide, et al.
Veröffentlicht: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
von: Chen, Sihan, et al.
Veröffentlicht: (2025)
von: Chen, Sihan, et al.
Veröffentlicht: (2025)
An LLM-based Agentic Framework for Accessible Network Control
von: Lin, Samuel, et al.
Veröffentlicht: (2025)
von: Lin, Samuel, et al.
Veröffentlicht: (2025)
Over-Searching in Search-Augmented Large Language Models
von: Xie, Roy, et al.
Veröffentlicht: (2026)
von: Xie, Roy, et al.
Veröffentlicht: (2026)
Federated Learning Clients Clustering with Adaptation to Data Drifts
von: Li, Minghao, et al.
Veröffentlicht: (2024)
von: Li, Minghao, et al.
Veröffentlicht: (2024)
Fast Large Language Model Collaborative Decoding via Speculation
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees
von: Chen, Sijia, et al.
Veröffentlicht: (2024)
von: Chen, Sijia, et al.
Veröffentlicht: (2024)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
von: Schmied, Thomas, et al.
Veröffentlicht: (2024)
von: Schmied, Thomas, et al.
Veröffentlicht: (2024)
Democratizing Large Language Model-Based Graph Data Augmentation via Latent Knowledge Graphs
von: Feng, Yushi, et al.
Veröffentlicht: (2025)
von: Feng, Yushi, et al.
Veröffentlicht: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
Automatic Calibration for Membership Inference Attack on Large Language Models
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2025)
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2025)
Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate
von: Fang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Fang, Zhiyuan, et al.
Veröffentlicht: (2025)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
von: Ho, Namgyu, et al.
Veröffentlicht: (2024)
von: Ho, Namgyu, et al.
Veröffentlicht: (2024)
SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation
von: Ren, Yanwei, et al.
Veröffentlicht: (2025)
von: Ren, Yanwei, et al.
Veröffentlicht: (2025)
The Shape of Wisdom: Decision Trajectories in Language Models
von: Rana, Shailesh
Veröffentlicht: (2026)
von: Rana, Shailesh
Veröffentlicht: (2026)
Consistency Models for Scalable and Fast Simulation-Based Inference
von: Schmitt, Marvin, et al.
Veröffentlicht: (2023)
von: Schmitt, Marvin, et al.
Veröffentlicht: (2023)
TimeCAP: Learning to Contextualize, Augment, and Predict Time Series Events with Large Language Model Agents
von: Lee, Geon, et al.
Veröffentlicht: (2025)
von: Lee, Geon, et al.
Veröffentlicht: (2025)
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
von: Li, Junsong, et al.
Veröffentlicht: (2025)
von: Li, Junsong, et al.
Veröffentlicht: (2025)
Do Language Models Have Bayesian Brains? Distinguishing Stochastic and Deterministic Decision Patterns within Large Language Models
von: Cui, Andrea Yaoyun, et al.
Veröffentlicht: (2025)
von: Cui, Andrea Yaoyun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
von: Shahout, Rana, et al.
Veröffentlicht: (2025) -
Orla: A Library for Serving LLM-Based Multi-Agent Systems
von: Shahout, Rana, et al.
Veröffentlicht: (2026) -
Queueing, Predictions, and LLMs: Challenges and Open Problems
von: Mitzenmacher, Michael, et al.
Veröffentlicht: (2025) -
Intra-request branch orchestration for efficient LLM reasoning
von: Jiang, Weifan, et al.
Veröffentlicht: (2025) -
Learning-Based Heavy Hitters and Flow Frequency Estimation in Streams
von: Shahout, Rana, et al.
Veröffentlicht: (2024)