Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Behdin, Kayhan, Fatahibaarzi, Ata, Song, Qingquan, Dai, Yun, Gupta, Aman, Wang, Zhipeng, Tang, Shao, Sang, Hejian, Dexter, Gregory, Zhu, Sirou, Zhu, Siyu, Dharamsi, Tejas, Kothapalli, Vignesh, Fu, Zhoutong, Cao, Yihan, Hsu, Pin-Lun, Borisyuk, Fedor, Pillai, Natesh, Simon, Luke, Mazumder, Rahul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM Query Scheduling with Prefix Reuse and Latency Constraints
von: Dexter, Gregory, et al.
Veröffentlicht: (2025)
von: Dexter, Gregory, et al.
Veröffentlicht: (2025)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2025)
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2025)
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
von: Chen, Wei-Rui, et al.
Veröffentlicht: (2025)
von: Chen, Wei-Rui, et al.
Veröffentlicht: (2025)
Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
von: Behdin, Kayhan, et al.
Veröffentlicht: (2025)
von: Behdin, Kayhan, et al.
Veröffentlicht: (2025)
Sparse NMF with Archetypal Regularization: Computational and Robustness Properties
von: Behdin, Kayhan, et al.
Veröffentlicht: (2021)
von: Behdin, Kayhan, et al.
Veröffentlicht: (2021)
Sparse PCA: A New Scalable Estimator Based On Integer Programming
von: Behdin, Kayhan, et al.
Veröffentlicht: (2021)
von: Behdin, Kayhan, et al.
Veröffentlicht: (2021)
BP-Seg: A graphical model approach to unsupervised and non-contiguous text segmentation using belief propagation
von: Li, Fengyi, et al.
Veröffentlicht: (2025)
von: Li, Fengyi, et al.
Veröffentlicht: (2025)
Sparse Gaussian Graphical Models with Discrete Optimization: Computational and Statistical Perspectives
von: Behdin, Kayhan, et al.
Veröffentlicht: (2023)
von: Behdin, Kayhan, et al.
Veröffentlicht: (2023)
End-to-end Feature Selection Approach for Learning Skinny Trees
von: Ibrahim, Shibal, et al.
Veröffentlicht: (2023)
von: Ibrahim, Shibal, et al.
Veröffentlicht: (2023)
Differentially Private High-dimensional Variable Selection via Integer Programming
von: Prastakos, Petros, et al.
Veröffentlicht: (2025)
von: Prastakos, Petros, et al.
Veröffentlicht: (2025)
Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
von: Lucas, Ryan, et al.
Veröffentlicht: (2025)
von: Lucas, Ryan, et al.
Veröffentlicht: (2025)
Efficient user history modeling with amortized inference for deep learning recommendation models
von: Hertel, Lars, et al.
Veröffentlicht: (2024)
von: Hertel, Lars, et al.
Veröffentlicht: (2024)
Modeling with Categorical Features via Exact Fusion and Sparsity Regularisation
von: Behdin, Kayhan, et al.
Veröffentlicht: (2026)
von: Behdin, Kayhan, et al.
Veröffentlicht: (2026)
ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language Models
von: Meng, Xiang, et al.
Veröffentlicht: (2024)
von: Meng, Xiang, et al.
Veröffentlicht: (2024)
Microergodicity implies orthogonality of Matérn fields on bounded domains in $\mathbb{R}^4$
von: Pillai, Natesh S.
Veröffentlicht: (2026)
von: Pillai, Natesh S.
Veröffentlicht: (2026)
Optimal Scaling for the Proximal Langevin Algorithm in High Dimensions
von: Pillai, Natesh S.
Veröffentlicht: (2022)
von: Pillai, Natesh S.
Veröffentlicht: (2022)
AlphaPO: Reward Shape Matters for LLM Alignment
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
HASSLE-free: A unified Framework for Sparse plus Low-Rank Matrix Decomposition for LLMs
von: Makni, Mehdi, et al.
Veröffentlicht: (2025)
von: Makni, Mehdi, et al.
Veröffentlicht: (2025)
Multi-Task Learning for Sparsity Pattern Heterogeneity: Statistical and Computational Perspectives
von: Behdin, Kayhan, et al.
Veröffentlicht: (2022)
von: Behdin, Kayhan, et al.
Veröffentlicht: (2022)
Liger Kernel: Efficient Triton Kernels for LLM Training
von: Hsu, Pin-Lun, et al.
Veröffentlicht: (2024)
von: Hsu, Pin-Lun, et al.
Veröffentlicht: (2024)
Policy Gradients for Optimal Parallel Tempering MCMC
von: Zhao, Daniel, et al.
Veröffentlicht: (2024)
von: Zhao, Daniel, et al.
Veröffentlicht: (2024)
Mixing on $k$ Columns of the Transvection Walk
von: Pillai, Natesh, et al.
Veröffentlicht: (2026)
von: Pillai, Natesh, et al.
Veröffentlicht: (2026)
Kac's walk on rotation matrices mixes in $n^2 \log n$ steps
von: Pillai, Natesh S., et al.
Veröffentlicht: (2026)
von: Pillai, Natesh S., et al.
Veröffentlicht: (2026)
OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial Optimization
von: Meng, Xiang, et al.
Veröffentlicht: (2024)
von: Meng, Xiang, et al.
Veröffentlicht: (2024)
Can Kernel Methods Explain How the Data Affects Neural Collapse?
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2024)
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2024)
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
von: Markovic-Voronov, Jelena, et al.
Veröffentlicht: (2026)
von: Markovic-Voronov, Jelena, et al.
Veröffentlicht: (2026)
No Free Lunch for Stochastic Gradient Langevin Dynamics
von: Pillai, Natesh S., et al.
Veröffentlicht: (2024)
von: Pillai, Natesh S., et al.
Veröffentlicht: (2024)
Enhancing Stability for Large Language Models Training in Constrained Bandwidth Networks
von: Dai, Yun, et al.
Veröffentlicht: (2024)
von: Dai, Yun, et al.
Veröffentlicht: (2024)
An Optimization Framework for Differentially Private Sparse Fine-Tuning
von: Makni, Mehdi, et al.
Veröffentlicht: (2025)
von: Makni, Mehdi, et al.
Veröffentlicht: (2025)
CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context Demonstrations
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2025)
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2025)
No Free Lunch for Approximate MCMC
von: Johndrow, James E., et al.
Veröffentlicht: (2020)
von: Johndrow, James E., et al.
Veröffentlicht: (2020)
A Heavily Right Strategy for Statistical Inference with Dependent Studies in Any Dimension
von: Liu, Tianle, et al.
Veröffentlicht: (2025)
von: Liu, Tianle, et al.
Veröffentlicht: (2025)
Support Tokens, Stability Margins, and a New Foundation for Robust LLMs
von: Agarwal, Deepak, et al.
Veröffentlicht: (2026)
von: Agarwal, Deepak, et al.
Veröffentlicht: (2026)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
MultiSlot ReRanker: A Generic Model-based Re-Ranking Framework in Recommendation Systems
von: Xiao, Qiang Charles, et al.
Veröffentlicht: (2024)
von: Xiao, Qiang Charles, et al.
Veröffentlicht: (2024)
Adaptive Divide and Conquer with Two Rounds of Communication
von: Kal, Niladri, et al.
Veröffentlicht: (2025)
von: Kal, Niladri, et al.
Veröffentlicht: (2025)
MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
von: Li, Guoyao, et al.
Veröffentlicht: (2025)
von: Li, Guoyao, et al.
Veröffentlicht: (2025)
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
von: Chen, Daiwei, et al.
Veröffentlicht: (2026)
von: Chen, Daiwei, et al.
Veröffentlicht: (2026)
Sampling for Quality: Training-Free Reward-Guided LLM Decoding via Sequential Monte Carlo
von: Markovic-Voronov, Jelena, et al.
Veröffentlicht: (2026)
von: Markovic-Voronov, Jelena, et al.
Veröffentlicht: (2026)
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2024)
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLM Query Scheduling with Prefix Reuse and Latency Constraints
von: Dexter, Gregory, et al.
Veröffentlicht: (2025) -
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2025) -
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
von: Chen, Wei-Rui, et al.
Veröffentlicht: (2025) -
Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
von: Behdin, Kayhan, et al.
Veröffentlicht: (2025) -
Sparse NMF with Archetypal Regularization: Computational and Robustness Properties
von: Behdin, Kayhan, et al.
Veröffentlicht: (2021)