CIR at the NTCIR-17 ULTRE-2 Task

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yu, Lulu, Bi, Keping, Guo, Jiafeng, Cheng, Xueqi
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910919617937408
author Yu, Lulu
Bi, Keping
Guo, Jiafeng
Cheng, Xueqi
author_facet Yu, Lulu
Bi, Keping
Guo, Jiafeng
Cheng, Xueqi
contents The Chinese academy of sciences Information Retrieval team (CIR) has participated in the NTCIR-17 ULTRE-2 task. This paper describes our approaches and reports our results on the ULTRE-2 task. We recognize the issue of false negatives in the Baidu search data in this competition is very severe, much more severe than position bias. Hence, we adopt the Dual Learning Algorithm (DLA) to address the position bias and use it as an auxiliary model to study how to alleviate the false negative issue. We approach the problem from two perspectives: 1) correcting the labels for non-clicked items by a relevance judgment model trained from DLA, and learn a new ranker that is initialized from DLA; 2) including random documents as true negatives and documents that have partial matching as hard negatives. Both methods can enhance the model performance and our best method has achieved nDCG@10 of 0.5355, which is 2.66% better than the best score from the organizer.
format Preprint
id arxiv_https___arxiv_org_abs_2310_11852
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CIR at the NTCIR-17 ULTRE-2 Task
Yu, Lulu
Bi, Keping
Guo, Jiafeng
Cheng, Xueqi
Information Retrieval
The Chinese academy of sciences Information Retrieval team (CIR) has participated in the NTCIR-17 ULTRE-2 task. This paper describes our approaches and reports our results on the ULTRE-2 task. We recognize the issue of false negatives in the Baidu search data in this competition is very severe, much more severe than position bias. Hence, we adopt the Dual Learning Algorithm (DLA) to address the position bias and use it as an auxiliary model to study how to alleviate the false negative issue. We approach the problem from two perspectives: 1) correcting the labels for non-clicked items by a relevance judgment model trained from DLA, and learn a new ranker that is initialized from DLA; 2) including random documents as true negatives and documents that have partial matching as hard negatives. Both methods can enhance the model performance and our best method has achieved nDCG@10 of 0.5355, which is 2.66% better than the best score from the organizer.
title CIR at the NTCIR-17 ULTRE-2 Task
topic Information Retrieval
url https://arxiv.org/abs/2310.11852