Don't Stop Me Now: Embedding Based Scheduling for LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Shahout, Rana, Malach, Eran, Liu, Chunwei, Jiang, Weifan, Yu, Minlan, Mitzenmacher, Michael |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Intra-request branch orchestration for efficient LLM reasoning
di: Jiang, Weifan, et al.
Pubblicazione: (2025)
di: Jiang, Weifan, et al.
Pubblicazione: (2025)
SkipPredict: When to Invest in Predictions for Scheduling
di: Shahout, Rana, et al.
Pubblicazione: (2024)
di: Shahout, Rana, et al.
Pubblicazione: (2024)
Learning-Based Heavy Hitters and Flow Frequency Estimation in Streams
di: Shahout, Rana, et al.
Pubblicazione: (2024)
di: Shahout, Rana, et al.
Pubblicazione: (2024)
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
di: Shahout, Rana, et al.
Pubblicazione: (2025)
di: Shahout, Rana, et al.
Pubblicazione: (2025)
Learning-Augmented Frequency Estimation in Sliding Windows
di: Shahout, Rana, et al.
Pubblicazione: (2024)
di: Shahout, Rana, et al.
Pubblicazione: (2024)
Fast Inference for Augmented Large Language Models
di: Shahout, Rana, et al.
Pubblicazione: (2024)
di: Shahout, Rana, et al.
Pubblicazione: (2024)
Orla: A Library for Serving LLM-Based Multi-Agent Systems
di: Shahout, Rana, et al.
Pubblicazione: (2026)
di: Shahout, Rana, et al.
Pubblicazione: (2026)
Queueing, Predictions, and LLMs: Challenges and Open Problems
di: Mitzenmacher, Michael, et al.
Pubblicazione: (2025)
di: Mitzenmacher, Michael, et al.
Pubblicazione: (2025)
Auto-Regressive Next-Token Predictors are Universal Learners
di: Malach, Eran
Pubblicazione: (2023)
di: Malach, Eran
Pubblicazione: (2023)
Don't Stop Me Yet: Sampling Loss Minima via Dissipative Riemannian Mechanics
di: Jacobsen, Albert Kjøller, et al.
Pubblicazione: (2026)
di: Jacobsen, Albert Kjøller, et al.
Pubblicazione: (2026)
Federated Learning Clients Clustering with Adaptation to Data Drifts
di: Li, Minghao, et al.
Pubblicazione: (2024)
di: Li, Minghao, et al.
Pubblicazione: (2024)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
di: Karchmer, Ari, et al.
Pubblicazione: (2025)
di: Karchmer, Ari, et al.
Pubblicazione: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
di: Jelassi, Samy, et al.
Pubblicazione: (2024)
di: Jelassi, Samy, et al.
Pubblicazione: (2024)
Don't Waste Your Time: Early Stopping Cross-Validation
di: Bergman, Edward, et al.
Pubblicazione: (2024)
di: Bergman, Edward, et al.
Pubblicazione: (2024)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
di: Mirtaheri, Parsa, et al.
Pubblicazione: (2025)
di: Mirtaheri, Parsa, et al.
Pubblicazione: (2025)
Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning
di: Ravie, Navin Sriram, et al.
Pubblicazione: (2026)
di: Ravie, Navin Sriram, et al.
Pubblicazione: (2026)
LLM Priors for ERM over Programs
di: Singhal, Shivam, et al.
Pubblicazione: (2025)
di: Singhal, Shivam, et al.
Pubblicazione: (2025)
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
di: Tsilivis, Nikolaos, et al.
Pubblicazione: (2025)
di: Tsilivis, Nikolaos, et al.
Pubblicazione: (2025)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
Don't Stop Me Now: Investigating the Information Interactions Involved in Overcoming Creative Blocks
di: Patricia Sanchez, et al.
Pubblicazione: (2025)
di: Patricia Sanchez, et al.
Pubblicazione: (2025)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
di: Qin, Tian, et al.
Pubblicazione: (2025)
di: Qin, Tian, et al.
Pubblicazione: (2025)
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
di: Johnson, Daniel D., et al.
Pubblicazione: (2024)
di: Johnson, Daniel D., et al.
Pubblicazione: (2024)
Show Me What You Don't Know: Efficient Sampling from Invariant Sets for Model Validation
di: Rousselot, Armand, et al.
Pubblicazione: (2026)
di: Rousselot, Armand, et al.
Pubblicazione: (2026)
Universal Length Generalization with Turing Programs
di: Hou, Kaiying, et al.
Pubblicazione: (2024)
di: Hou, Kaiying, et al.
Pubblicazione: (2024)
Whitespaces Don't Lie: Feature-Driven and Embedding-Based Approaches for Detecting Machine-Generated Code
di: Nirob, Syed Mehedi Hasan, et al.
Pubblicazione: (2026)
di: Nirob, Syed Mehedi Hasan, et al.
Pubblicazione: (2026)
Quantizing With Randomized Hadamard Transforms: Efficient Heuristic Now Proven
di: Ben-Basat, Ran, et al.
Pubblicazione: (2026)
di: Ben-Basat, Ran, et al.
Pubblicazione: (2026)
Ignore Me But Don't Replace Me: Utilizing Non-Linguistic Elements for Pretraining on the Cybersecurity Domain
di: Jang, Eugene, et al.
Pubblicazione: (2024)
di: Jang, Eugene, et al.
Pubblicazione: (2024)
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
di: Edelman, Benjamin L., et al.
Pubblicazione: (2024)
di: Edelman, Benjamin L., et al.
Pubblicazione: (2024)
Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
di: Roux, Christophe, et al.
Pubblicazione: (2025)
di: Roux, Christophe, et al.
Pubblicazione: (2025)
Don't Transform the Code, Code the Transforms: Towards Precise Code Rewriting using LLMs
di: Cummins, Chris, et al.
Pubblicazione: (2024)
di: Cummins, Chris, et al.
Pubblicazione: (2024)
THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
di: Li, Minghao, et al.
Pubblicazione: (2023)
di: Li, Minghao, et al.
Pubblicazione: (2023)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
di: Hernandez, Adriano
Pubblicazione: (2024)
di: Hernandez, Adriano
Pubblicazione: (2024)
Don't Get Me Wrong: How to Apply Deep Visual Interpretations to Time Series
di: Loeffler, Christoffer, et al.
Pubblicazione: (2022)
di: Loeffler, Christoffer, et al.
Pubblicazione: (2022)
Bayesian Mixture-of-Experts: Towards Making LLMs Know What They Don't Know
di: Li, Albus Yizhuo
Pubblicazione: (2025)
di: Li, Albus Yizhuo
Pubblicazione: (2025)
Don't Mesh with Me: Generating Constructive Solid Geometry Instead of Meshes by Fine-Tuning a Code-Generation LLM
di: Mews, Maximilian, et al.
Pubblicazione: (2024)
di: Mews, Maximilian, et al.
Pubblicazione: (2024)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
di: Hankendi, Can, et al.
Pubblicazione: (2026)
di: Hankendi, Can, et al.
Pubblicazione: (2026)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
di: Zhang, Jiefu, et al.
Pubblicazione: (2026)
di: Zhang, Jiefu, et al.
Pubblicazione: (2026)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
What LLMs Think When You Don't Tell Them What to Think About?
di: Kwon, Yongchan, et al.
Pubblicazione: (2026)
di: Kwon, Yongchan, et al.
Pubblicazione: (2026)
Position: Don't be Afraid of Over-Smoothing And Over-Squashing
di: Kormann, Niklas, et al.
Pubblicazione: (2026)
di: Kormann, Niklas, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Intra-request branch orchestration for efficient LLM reasoning
di: Jiang, Weifan, et al.
Pubblicazione: (2025) -
SkipPredict: When to Invest in Predictions for Scheduling
di: Shahout, Rana, et al.
Pubblicazione: (2024) -
Learning-Based Heavy Hitters and Flow Frequency Estimation in Streams
di: Shahout, Rana, et al.
Pubblicazione: (2024) -
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
di: Shahout, Rana, et al.
Pubblicazione: (2025) -
Learning-Augmented Frequency Estimation in Sliding Windows
di: Shahout, Rana, et al.
Pubblicazione: (2024)