Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
Fuente:
arXiv
Guardado en:
| Autores principales: | Behdin, Kayhan, Song, Qingquan, Vasudevan, Sriram, Sheng, Jian, Ma, Xiaojing, Zhou, Z, Zhu, Chuanrui, Li, Guoyao, Nguyen, Chanh, Ghosh, Sayan, Sang, Hejian, Baarzi, Ata Fatahi, Ramachandran, Sundara Raman, Wang, Xiaoqing, Lan, Qing, S, Vinay Y, Guo, Qi, Johnson, Caleb, Wang, Zhipeng, Borisyuk, Fedor |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
por: Li, Guoyao, et al.
Publicado: (2025)
por: Li, Guoyao, et al.
Publicado: (2025)
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
por: Behdin, Kayhan, et al.
Publicado: (2025)
por: Behdin, Kayhan, et al.
Publicado: (2025)
LLM Query Scheduling with Prefix Reuse and Latency Constraints
por: Dexter, Gregory, et al.
Publicado: (2025)
por: Dexter, Gregory, et al.
Publicado: (2025)
Sparse NMF with Archetypal Regularization: Computational and Robustness Properties
por: Behdin, Kayhan, et al.
Publicado: (2021)
por: Behdin, Kayhan, et al.
Publicado: (2021)
Sparse PCA: A New Scalable Estimator Based On Integer Programming
por: Behdin, Kayhan, et al.
Publicado: (2021)
por: Behdin, Kayhan, et al.
Publicado: (2021)
Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
por: Lucas, Ryan, et al.
Publicado: (2025)
por: Lucas, Ryan, et al.
Publicado: (2025)
Emerging memory technologies at room/cryogenic temperature
por: Raman, Siddhartha Raman Sundara
Publicado: (2026)
por: Raman, Siddhartha Raman Sundara
Publicado: (2026)
Sparse Gaussian Graphical Models with Discrete Optimization: Computational and Statistical Perspectives
por: Behdin, Kayhan, et al.
Publicado: (2023)
por: Behdin, Kayhan, et al.
Publicado: (2023)
End-to-end Feature Selection Approach for Learning Skinny Trees
por: Ibrahim, Shibal, et al.
Publicado: (2023)
por: Ibrahim, Shibal, et al.
Publicado: (2023)
Differentially Private High-dimensional Variable Selection via Integer Programming
por: Prastakos, Petros, et al.
Publicado: (2025)
por: Prastakos, Petros, et al.
Publicado: (2025)
Learning to Retrieve for Job Matching
por: Shen, Jianqiang, et al.
Publicado: (2024)
por: Shen, Jianqiang, et al.
Publicado: (2024)
Weak Supervision for Improved Precision in Search Systems
por: Vasudevan, Sriram
Publicado: (2025)
por: Vasudevan, Sriram
Publicado: (2025)
Modeling with Categorical Features via Exact Fusion and Sparsity Regularisation
por: Behdin, Kayhan, et al.
Publicado: (2026)
por: Behdin, Kayhan, et al.
Publicado: (2026)
ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language Models
por: Meng, Xiang, et al.
Publicado: (2024)
por: Meng, Xiang, et al.
Publicado: (2024)
ABI: A tightly integrated, unified, sparsity-aware, reconfigurable, compute near-register file/cache GPU architecture with light-weight softmax for deep learning, linear algebra, and Ising compute
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
HASSLE-free: A unified Framework for Sparse plus Low-Rank Matrix Decomposition for LLMs
por: Makni, Mehdi, et al.
Publicado: (2025)
por: Makni, Mehdi, et al.
Publicado: (2025)
Multi-Task Learning for Sparsity Pattern Heterogeneity: Statistical and Computational Perspectives
por: Behdin, Kayhan, et al.
Publicado: (2022)
por: Behdin, Kayhan, et al.
Publicado: (2022)
A comparative study on power delivery aspects of compute-in/near-memory approaches using DRAM
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
A complete discussion on fully reconfigurable, digital, scalable, graph and sparsity-aware near-memory accelerator for graph neural networks
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
por: Chen, Wei-Rui, et al.
Publicado: (2025)
por: Chen, Wei-Rui, et al.
Publicado: (2025)
A comprehensive study on ILP acceleration accounting for sparsity, area, energy, data movement using near-memory architecture
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
Complexity at Scale: A Quantitative Analysis of an Alibaba Microservice Deployment
por: Winchester, Giles, et al.
Publicado: (2025)
por: Winchester, Giles, et al.
Publicado: (2025)
BP-Seg: A graphical model approach to unsupervised and non-contiguous text segmentation using belief propagation
por: Li, Fengyi, et al.
Publicado: (2025)
por: Li, Fengyi, et al.
Publicado: (2025)
OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial Optimization
por: Meng, Xiang, et al.
Publicado: (2024)
por: Meng, Xiang, et al.
Publicado: (2024)
Few-Shot Deployment of Pretrained MRI Transformers in Brain Imaging Tasks
por: Li, Mengyu, et al.
Publicado: (2025)
por: Li, Mengyu, et al.
Publicado: (2025)
Efficient user history modeling with amortized inference for deep learning recommendation models
por: Hertel, Lars, et al.
Publicado: (2024)
por: Hertel, Lars, et al.
Publicado: (2024)
Predictability of Global AI Weather Models
por: Kieu, Chanh
Publicado: (2024)
por: Kieu, Chanh
Publicado: (2024)
Semantic Search At LinkedIn
por: Borisyuk, Fedor, et al.
Publicado: (2026)
por: Borisyuk, Fedor, et al.
Publicado: (2026)
Optimal Neutron Spectrum Database for In‐reactor 238 Pu Production
por: Qingquan Pan, et al.
Publicado: (2025)
por: Qingquan Pan, et al.
Publicado: (2025)
Serving the Job Hunter
por: Gerhan, David
Publicado: (1973)
por: Gerhan, David
Publicado: (1973)
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
por: Markovic-Voronov, Jelena, et al.
Publicado: (2026)
por: Markovic-Voronov, Jelena, et al.
Publicado: (2026)
Tiered Reward: Designing Rewards for Specification and Fast Learning of Desired Behavior
por: Zhou, Zhiyuan, et al.
Publicado: (2022)
por: Zhou, Zhiyuan, et al.
Publicado: (2022)
A Role of Environmental Complexity on Representation Learning in Deep Reinforcement Learning Agents
por: Liu, Andrew, et al.
Publicado: (2024)
por: Liu, Andrew, et al.
Publicado: (2024)
A Flexible Job Shop Scheduling Problem Involving Reconfigurable Machine Tools Under Industry 5.0
por: Bakhshi-Khaniki, Hessam, et al.
Publicado: (2024)
por: Bakhshi-Khaniki, Hessam, et al.
Publicado: (2024)
RLEEGNet: Integrating Brain-Computer Interfaces with Adaptive AI for Intuitive Responsiveness and High-Accuracy Motor Imagery Classification
por: Nallani, Sriram V. C., et al.
Publicado: (2024)
por: Nallani, Sriram V. C., et al.
Publicado: (2024)
Serving Compound Inference Systems on Datacenter GPUs
por: Devata, Sriram, et al.
Publicado: (2026)
por: Devata, Sriram, et al.
Publicado: (2026)
Causal relationship between immune cells and the risk of myeloperoxidase antineutrophil cytoplasmic antibody‐associated vasculitis: A Mendelian randomization study
por: Xiaojing Cai, et al.
Publicado: (2024)
por: Xiaojing Cai, et al.
Publicado: (2024)
LinkSAGE: Optimizing Job Matching Using Graph Neural Networks
por: Liu, Ping, et al.
Publicado: (2024)
por: Liu, Ping, et al.
Publicado: (2024)
Fingerprinting Deep Packet Inspection Devices by Their Ambiguities
por: Xue, Diwen, et al.
Publicado: (2025)
por: Xue, Diwen, et al.
Publicado: (2025)
An Optimization Framework for Differentially Private Sparse Fine-Tuning
por: Makni, Mehdi, et al.
Publicado: (2025)
por: Makni, Mehdi, et al.
Publicado: (2025)
Ejemplares similares
-
MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
por: Li, Guoyao, et al.
Publicado: (2025) -
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
por: Behdin, Kayhan, et al.
Publicado: (2025) -
LLM Query Scheduling with Prefix Reuse and Latency Constraints
por: Dexter, Gregory, et al.
Publicado: (2025) -
Sparse NMF with Archetypal Regularization: Computational and Robustness Properties
por: Behdin, Kayhan, et al.
Publicado: (2021) -
Sparse PCA: A New Scalable Estimator Based On Integer Programming
por: Behdin, Kayhan, et al.
Publicado: (2021)