An Efficient Long-Context Ranking Architecture With Calibrated LLM Distillation: Application to Person-Job Fit

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jouanneau, Warren, Jouffroy, Emma, Palyart, Marc
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908770481733632
author Jouanneau, Warren
Jouffroy, Emma
Palyart, Marc
author_facet Jouanneau, Warren
Jouffroy, Emma
Palyart, Marc
contents Finding the most relevant person for a job proposal in real time is challenging, especially when resumes are long, structured, and multilingual. In this paper, we propose a re-ranking model based on a new generation of late cross-attention architecture, that decomposes both resumes and project briefs to efficiently handle long-context inputs with minimal computational overhead. To mitigate historical data biases, we use a generative large language model (LLM) as a teacher, generating fine-grained, semantically grounded supervision. This signal is distilled into our student model via an enriched distillation loss function. The resulting model produces skill-fit scores that enable consistent and interpretable person-job matching. Experiments on relevance, ranking, and calibration metrics demonstrate that our approach outperforms state-of-the-art baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10321
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle An Efficient Long-Context Ranking Architecture With Calibrated LLM Distillation: Application to Person-Job Fit
Jouanneau, Warren
Jouffroy, Emma
Palyart, Marc
Computation and Language
Information Retrieval
Machine Learning
Social and Information Networks
Finding the most relevant person for a job proposal in real time is challenging, especially when resumes are long, structured, and multilingual. In this paper, we propose a re-ranking model based on a new generation of late cross-attention architecture, that decomposes both resumes and project briefs to efficiently handle long-context inputs with minimal computational overhead. To mitigate historical data biases, we use a generative large language model (LLM) as a teacher, generating fine-grained, semantically grounded supervision. This signal is distilled into our student model via an enriched distillation loss function. The resulting model produces skill-fit scores that enable consistent and interpretable person-job matching. Experiments on relevance, ranking, and calibration metrics demonstrate that our approach outperforms state-of-the-art baselines.
title An Efficient Long-Context Ranking Architecture With Calibrated LLM Distillation: Application to Person-Job Fit
topic Computation and Language
Information Retrieval
Machine Learning
Social and Information Networks
url https://arxiv.org/abs/2601.10321