Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Behdin, Kayhan, Fatahibaarzi, Ata, Song, Qingquan, Dai, Yun, Gupta, Aman, Wang, Zhipeng, Tang, Shao, Sang, Hejian, Dexter, Gregory, Zhu, Sirou, Zhu, Siyu, Dharamsi, Tejas, Kothapalli, Vignesh, Fu, Zhoutong, Cao, Yihan, Hsu, Pin-Lun, Borisyuk, Fedor, Pillai, Natesh, Simon, Luke, Mazumder, Rahul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLM Query Scheduling with Prefix Reuse and Latency Constraints
por: Dexter, Gregory, et al.
Publicado: (2025)
por: Dexter, Gregory, et al.
Publicado: (2025)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
por: Chen, Wei-Rui, et al.
Publicado: (2025)
por: Chen, Wei-Rui, et al.
Publicado: (2025)
Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
por: Behdin, Kayhan, et al.
Publicado: (2025)
por: Behdin, Kayhan, et al.
Publicado: (2025)
Sparse NMF with Archetypal Regularization: Computational and Robustness Properties
por: Behdin, Kayhan, et al.
Publicado: (2021)
por: Behdin, Kayhan, et al.
Publicado: (2021)
Sparse PCA: A New Scalable Estimator Based On Integer Programming
por: Behdin, Kayhan, et al.
Publicado: (2021)
por: Behdin, Kayhan, et al.
Publicado: (2021)
BP-Seg: A graphical model approach to unsupervised and non-contiguous text segmentation using belief propagation
por: Li, Fengyi, et al.
Publicado: (2025)
por: Li, Fengyi, et al.
Publicado: (2025)
Sparse Gaussian Graphical Models with Discrete Optimization: Computational and Statistical Perspectives
por: Behdin, Kayhan, et al.
Publicado: (2023)
por: Behdin, Kayhan, et al.
Publicado: (2023)
End-to-end Feature Selection Approach for Learning Skinny Trees
por: Ibrahim, Shibal, et al.
Publicado: (2023)
por: Ibrahim, Shibal, et al.
Publicado: (2023)
Differentially Private High-dimensional Variable Selection via Integer Programming
por: Prastakos, Petros, et al.
Publicado: (2025)
por: Prastakos, Petros, et al.
Publicado: (2025)
Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
por: Lucas, Ryan, et al.
Publicado: (2025)
por: Lucas, Ryan, et al.
Publicado: (2025)
Efficient user history modeling with amortized inference for deep learning recommendation models
por: Hertel, Lars, et al.
Publicado: (2024)
por: Hertel, Lars, et al.
Publicado: (2024)
Modeling with Categorical Features via Exact Fusion and Sparsity Regularisation
por: Behdin, Kayhan, et al.
Publicado: (2026)
por: Behdin, Kayhan, et al.
Publicado: (2026)
ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language Models
por: Meng, Xiang, et al.
Publicado: (2024)
por: Meng, Xiang, et al.
Publicado: (2024)
Microergodicity implies orthogonality of Matérn fields on bounded domains in $\mathbb{R}^4$
por: Pillai, Natesh S.
Publicado: (2026)
por: Pillai, Natesh S.
Publicado: (2026)
Optimal Scaling for the Proximal Langevin Algorithm in High Dimensions
por: Pillai, Natesh S.
Publicado: (2022)
por: Pillai, Natesh S.
Publicado: (2022)
AlphaPO: Reward Shape Matters for LLM Alignment
por: Gupta, Aman, et al.
Publicado: (2025)
por: Gupta, Aman, et al.
Publicado: (2025)
HASSLE-free: A unified Framework for Sparse plus Low-Rank Matrix Decomposition for LLMs
por: Makni, Mehdi, et al.
Publicado: (2025)
por: Makni, Mehdi, et al.
Publicado: (2025)
Multi-Task Learning for Sparsity Pattern Heterogeneity: Statistical and Computational Perspectives
por: Behdin, Kayhan, et al.
Publicado: (2022)
por: Behdin, Kayhan, et al.
Publicado: (2022)
Liger Kernel: Efficient Triton Kernels for LLM Training
por: Hsu, Pin-Lun, et al.
Publicado: (2024)
por: Hsu, Pin-Lun, et al.
Publicado: (2024)
Policy Gradients for Optimal Parallel Tempering MCMC
por: Zhao, Daniel, et al.
Publicado: (2024)
por: Zhao, Daniel, et al.
Publicado: (2024)
Mixing on $k$ Columns of the Transvection Walk
por: Pillai, Natesh, et al.
Publicado: (2026)
por: Pillai, Natesh, et al.
Publicado: (2026)
Kac's walk on rotation matrices mixes in $n^2 \log n$ steps
por: Pillai, Natesh S., et al.
Publicado: (2026)
por: Pillai, Natesh S., et al.
Publicado: (2026)
OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial Optimization
por: Meng, Xiang, et al.
Publicado: (2024)
por: Meng, Xiang, et al.
Publicado: (2024)
Can Kernel Methods Explain How the Data Affects Neural Collapse?
por: Kothapalli, Vignesh, et al.
Publicado: (2024)
por: Kothapalli, Vignesh, et al.
Publicado: (2024)
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
por: Markovic-Voronov, Jelena, et al.
Publicado: (2026)
por: Markovic-Voronov, Jelena, et al.
Publicado: (2026)
No Free Lunch for Stochastic Gradient Langevin Dynamics
por: Pillai, Natesh S., et al.
Publicado: (2024)
por: Pillai, Natesh S., et al.
Publicado: (2024)
Enhancing Stability for Large Language Models Training in Constrained Bandwidth Networks
por: Dai, Yun, et al.
Publicado: (2024)
por: Dai, Yun, et al.
Publicado: (2024)
An Optimization Framework for Differentially Private Sparse Fine-Tuning
por: Makni, Mehdi, et al.
Publicado: (2025)
por: Makni, Mehdi, et al.
Publicado: (2025)
CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context Demonstrations
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
No Free Lunch for Approximate MCMC
por: Johndrow, James E., et al.
Publicado: (2020)
por: Johndrow, James E., et al.
Publicado: (2020)
A Heavily Right Strategy for Statistical Inference with Dependent Studies in Any Dimension
por: Liu, Tianle, et al.
Publicado: (2025)
por: Liu, Tianle, et al.
Publicado: (2025)
Support Tokens, Stability Margins, and a New Foundation for Robust LLMs
por: Agarwal, Deepak, et al.
Publicado: (2026)
por: Agarwal, Deepak, et al.
Publicado: (2026)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
por: Lin, Xiaofeng, et al.
Publicado: (2026)
por: Lin, Xiaofeng, et al.
Publicado: (2026)
MultiSlot ReRanker: A Generic Model-based Re-Ranking Framework in Recommendation Systems
por: Xiao, Qiang Charles, et al.
Publicado: (2024)
por: Xiao, Qiang Charles, et al.
Publicado: (2024)
Adaptive Divide and Conquer with Two Rounds of Communication
por: Kal, Niladri, et al.
Publicado: (2025)
por: Kal, Niladri, et al.
Publicado: (2025)
MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
por: Li, Guoyao, et al.
Publicado: (2025)
por: Li, Guoyao, et al.
Publicado: (2025)
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
por: Chen, Daiwei, et al.
Publicado: (2026)
por: Chen, Daiwei, et al.
Publicado: (2026)
Sampling for Quality: Training-Free Reward-Guided LLM Decoding via Sequential Monte Carlo
por: Markovic-Voronov, Jelena, et al.
Publicado: (2026)
por: Markovic-Voronov, Jelena, et al.
Publicado: (2026)
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
por: Kothapalli, Vignesh, et al.
Publicado: (2024)
por: Kothapalli, Vignesh, et al.
Publicado: (2024)
Ejemplares similares
-
LLM Query Scheduling with Prefix Reuse and Latency Constraints
por: Dexter, Gregory, et al.
Publicado: (2025) -
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
por: Kothapalli, Vignesh, et al.
Publicado: (2025) -
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
por: Chen, Wei-Rui, et al.
Publicado: (2025) -
Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
por: Behdin, Kayhan, et al.
Publicado: (2025) -
Sparse NMF with Archetypal Regularization: Computational and Robustness Properties
por: Behdin, Kayhan, et al.
Publicado: (2021)