MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Guoyao, He, Ran, Jing, Shusen, Behdin, Kayhan, Wang, Yubo, Ramachandran, Sundara Raman, Nguyen, Chanh, Sheng, Jian, Ma, Xiaojing, Zhu, Chuanrui, Vasudevan, Sriram, Wu, Muchen, Ghosh, Sayan, Su, Lin, Song, Qingquan, Wang, Xiaoqing, Wang, Zhipeng, Lan, Qing, Chen, Yanning, Wu, Jingwei, Simon, Luke, Zhang, Wenjing, Guo, Qi, Borisyuk, Fedor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
by: Behdin, Kayhan, et al.
Published: (2025)
by: Behdin, Kayhan, et al.
Published: (2025)
Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
by: Lucas, Ryan, et al.
Published: (2025)
by: Lucas, Ryan, et al.
Published: (2025)
Sparse NMF with Archetypal Regularization: Computational and Robustness Properties
by: Behdin, Kayhan, et al.
Published: (2021)
by: Behdin, Kayhan, et al.
Published: (2021)
Sparse PCA: A New Scalable Estimator Based On Integer Programming
by: Behdin, Kayhan, et al.
Published: (2021)
by: Behdin, Kayhan, et al.
Published: (2021)
ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language Models
by: Meng, Xiang, et al.
Published: (2024)
by: Meng, Xiang, et al.
Published: (2024)
Emerging memory technologies at room/cryogenic temperature
by: Raman, Siddhartha Raman Sundara
Published: (2026)
by: Raman, Siddhartha Raman Sundara
Published: (2026)
Sparse Gaussian Graphical Models with Discrete Optimization: Computational and Statistical Perspectives
by: Behdin, Kayhan, et al.
Published: (2023)
by: Behdin, Kayhan, et al.
Published: (2023)
End-to-end Feature Selection Approach for Learning Skinny Trees
by: Ibrahim, Shibal, et al.
Published: (2023)
by: Ibrahim, Shibal, et al.
Published: (2023)
Differentially Private High-dimensional Variable Selection via Integer Programming
by: Prastakos, Petros, et al.
Published: (2025)
by: Prastakos, Petros, et al.
Published: (2025)
Weak Supervision for Improved Precision in Search Systems
by: Vasudevan, Sriram
Published: (2025)
by: Vasudevan, Sriram
Published: (2025)
High Fidelity Textual User Representation over Heterogeneous Sources via Reinforcement Learning
by: Arora, Rajat, et al.
Published: (2026)
by: Arora, Rajat, et al.
Published: (2026)
Modeling with Categorical Features via Exact Fusion and Sparsity Regularisation
by: Behdin, Kayhan, et al.
Published: (2026)
by: Behdin, Kayhan, et al.
Published: (2026)
BP-Seg: A graphical model approach to unsupervised and non-contiguous text segmentation using belief propagation
by: Li, Fengyi, et al.
Published: (2025)
by: Li, Fengyi, et al.
Published: (2025)
Inverse Mixed Strategy Games with Generative Trajectory Models
by: Sun, Max Muchen, et al.
Published: (2025)
by: Sun, Max Muchen, et al.
Published: (2025)
ABI: A tightly integrated, unified, sparsity-aware, reconfigurable, compute near-register file/cache GPU architecture with light-weight softmax for deep learning, linear algebra, and Ising compute
by: Raman, Siddhartha Raman Sundara, et al.
Published: (2026)
by: Raman, Siddhartha Raman Sundara, et al.
Published: (2026)
Semantic Search At LinkedIn
by: Borisyuk, Fedor, et al.
Published: (2026)
by: Borisyuk, Fedor, et al.
Published: (2026)
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
Mixed Strategy Nash Equilibrium for Crowd Navigation
by: Sun, Max Muchen, et al.
Published: (2024)
by: Sun, Max Muchen, et al.
Published: (2024)
HASSLE-free: A unified Framework for Sparse plus Low-Rank Matrix Decomposition for LLMs
by: Makni, Mehdi, et al.
Published: (2025)
by: Makni, Mehdi, et al.
Published: (2025)
Multi-Task Learning for Sparsity Pattern Heterogeneity: Statistical and Computational Perspectives
by: Behdin, Kayhan, et al.
Published: (2022)
by: Behdin, Kayhan, et al.
Published: (2022)
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
by: Behdin, Kayhan, et al.
Published: (2025)
by: Behdin, Kayhan, et al.
Published: (2025)
Learning to Retrieve for Job Matching
by: Shen, Jianqiang, et al.
Published: (2024)
by: Shen, Jianqiang, et al.
Published: (2024)
A comparative study on power delivery aspects of compute-in/near-memory approaches using DRAM
by: Raman, Siddhartha Raman Sundara, et al.
Published: (2026)
by: Raman, Siddhartha Raman Sundara, et al.
Published: (2026)
A complete discussion on fully reconfigurable, digital, scalable, graph and sparsity-aware near-memory accelerator for graph neural networks
by: Raman, Siddhartha Raman Sundara, et al.
Published: (2026)
by: Raman, Siddhartha Raman Sundara, et al.
Published: (2026)
Intrinsic Mixed-state Topological Order
by: Wang, Zijian, et al.
Published: (2023)
by: Wang, Zijian, et al.
Published: (2023)
Causal relationship between immune cells and the risk of myeloperoxidase antineutrophil cytoplasmic antibody‐associated vasculitis: A Mendelian randomization study
by: Xiaojing Cai, et al.
Published: (2024)
by: Xiaojing Cai, et al.
Published: (2024)
A comprehensive study on ILP acceleration accounting for sparsity, area, energy, data movement using near-memory architecture
by: Raman, Siddhartha Raman Sundara, et al.
Published: (2026)
by: Raman, Siddhartha Raman Sundara, et al.
Published: (2026)
Harnessing Machine Learning to Understand and Design Disordered Solids
by: Muchen Wang, et al.
Published: (2026)
by: Muchen Wang, et al.
Published: (2026)
Stability Evaluation of Landfill Slopes Using Quantitative Risk Assessment
by: Wang Xiong, et al.
Published: (2025)
by: Wang Xiong, et al.
Published: (2025)
Fingerprinting Deep Packet Inspection Devices by Their Ambiguities
by: Xue, Diwen, et al.
Published: (2025)
by: Xue, Diwen, et al.
Published: (2025)
OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial Optimization
by: Meng, Xiang, et al.
Published: (2024)
by: Meng, Xiang, et al.
Published: (2024)
Mixing Configurations for Downstream Prediction
by: Wang, Juntang, et al.
Published: (2025)
by: Wang, Juntang, et al.
Published: (2025)
scDFM: Distributional Flow Matching Model for Robust Single-Cell Perturbation Prediction
by: Yu, Chenglei, et al.
Published: (2026)
by: Yu, Chenglei, et al.
Published: (2026)
PACE: Geometry-Aware Bridge Transport for Single-Cell Trajectory Inference
by: Yu, Chenglei, et al.
Published: (2026)
by: Yu, Chenglei, et al.
Published: (2026)
Efficient user history modeling with amortized inference for deep learning recommendation models
by: Hertel, Lars, et al.
Published: (2024)
by: Hertel, Lars, et al.
Published: (2024)
Olmix: A Framework for Data Mixing Throughout LM Development
by: Chen, Mayee F., et al.
Published: (2026)
by: Chen, Mayee F., et al.
Published: (2026)
Predictability of Global AI Weather Models
by: Kieu, Chanh
Published: (2024)
by: Kieu, Chanh
Published: (2024)
Optimal Neutron Spectrum Database for In‐reactor 238 Pu Production
by: Qingquan Pan, et al.
Published: (2025)
by: Qingquan Pan, et al.
Published: (2025)
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
by: Dong, Yanhao, et al.
Published: (2025)
by: Dong, Yanhao, et al.
Published: (2025)
Mixed State Entanglement Entropy in CFT
by: Jiang, Xin, et al.
Published: (2025)
by: Jiang, Xin, et al.
Published: (2025)
Similar Items
-
Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
by: Behdin, Kayhan, et al.
Published: (2025) -
Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
by: Lucas, Ryan, et al.
Published: (2025) -
Sparse NMF with Archetypal Regularization: Computational and Robustness Properties
by: Behdin, Kayhan, et al.
Published: (2021) -
Sparse PCA: A New Scalable Estimator Based On Integer Programming
by: Behdin, Kayhan, et al.
Published: (2021) -
ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language Models
by: Meng, Xiang, et al.
Published: (2024)