LLMs can be Fooled into Labelling a Document as Relevant (best café near me; this paper is perfectly relevant)
Fuente:
arXiv
Saved in:
| Main Authors: | Alaofi, Marwah, Thomas, Paul, Scholer, Falk, Sanderson, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Generative LLMs Create Query Variants for Test Collections? An Exploratory Study
by: Alaofi, Marwah, et al.
Published: (2025)
by: Alaofi, Marwah, et al.
Published: (2025)
Demographically-Inspired Query Variants Using an LLM
by: Alaofi, Marwah, et al.
Published: (2025)
by: Alaofi, Marwah, et al.
Published: (2025)
Generative Information Retrieval Evaluation
by: Alaofi, Marwah, et al.
Published: (2024)
by: Alaofi, Marwah, et al.
Published: (2024)
Online and Offline Evaluation in Search Clarification
by: Tavakoli, Leila, et al.
Published: (2024)
by: Tavakoli, Leila, et al.
Published: (2024)
Personalisation of Generic Library Search Results Using Student Enrolment Information
by: Alaofi, Marwah, et al.
Published: (2015)
by: Alaofi, Marwah, et al.
Published: (2015)
Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment
by: Schnabel, Julian A., et al.
Published: (2025)
by: Schnabel, Julian A., et al.
Published: (2025)
RMIT-ADM+S at the MMU-RAG NeurIPS 2025 Competition
by: Ran, Kun, et al.
Published: (2026)
by: Ran, Kun, et al.
Published: (2026)
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Towards Investigating Biases in Spoken Conversational Search
by: Cherumanal, Sachin Pathiyan, et al.
Published: (2024)
by: Cherumanal, Sachin Pathiyan, et al.
Published: (2024)
Control Search Rankings, Control the World: What is a Good Search Engine?
by: Coghlan, Simon, et al.
Published: (2025)
by: Coghlan, Simon, et al.
Published: (2025)
LLM-based relevance assessment still can't replace human relevance assessment
by: Clarke, Charles L. A., et al.
Published: (2024)
by: Clarke, Charles L. A., et al.
Published: (2024)
The Effects of Demographic Instructions on LLM Personas
by: de Paula, Angel Felipe Magnossão, et al.
Published: (2025)
by: de Paula, Angel Felipe Magnossão, et al.
Published: (2025)
Metamorphic Evaluation of ChatGPT as a Recommender System
by: Khirbat, Madhurima, et al.
Published: (2024)
by: Khirbat, Madhurima, et al.
Published: (2024)
Validating LLM-Generated Relevance Labels for Educational Resource Search
by: Sebastian, Ratan J., et al.
Published: (2025)
by: Sebastian, Ratan J., et al.
Published: (2025)
Don't Use LLMs to Make Relevance Judgments
by: Soboroff, Ian
Published: (2024)
by: Soboroff, Ian
Published: (2024)
Diversity-Augmented Negative Sampling for Implicit Collaborative Filtering
by: Xuan, Yueqing, et al.
Published: (2025)
by: Xuan, Yueqing, et al.
Published: (2025)
Evaluating and Addressing Fairness Across User Groups in Negative Sampling for Recommender Systems
by: Xuan, Yueqing, et al.
Published: (2023)
by: Xuan, Yueqing, et al.
Published: (2023)
Hybrid Pooling with LLMs via Relevance Context Learning
by: Otero, David, et al.
Published: (2026)
by: Otero, David, et al.
Published: (2026)
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
Contextual Relevance and Adaptive Sampling for LLM-Based Document Reranking
by: Huang, Jerry, et al.
Published: (2025)
by: Huang, Jerry, et al.
Published: (2025)
Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
by: Gienapp, Lukas, et al.
Published: (2025)
by: Gienapp, Lukas, et al.
Published: (2025)
Stairway to Fairness: Connecting Group and Individual Fairness
by: Rampisela, Theresia Veronika, et al.
Published: (2025)
by: Rampisela, Theresia Veronika, et al.
Published: (2025)
Walert: Putting Conversational Search Knowledge into Action by Building and Evaluating a Large Language Model-Powered Chatbot
by: Cherumanal, Sachin Pathiyan, et al.
Published: (2024)
by: Cherumanal, Sachin Pathiyan, et al.
Published: (2024)
REALM: Recursive Relevance Modeling for LLM-based Document Re-Ranking
by: Wang, Pinhuan, et al.
Published: (2025)
by: Wang, Pinhuan, et al.
Published: (2025)
Beyond Yes and No: Improving Zero-Shot LLM Rankers via Scoring Fine-Grained Relevance Labels
by: Zhuang, Honglei, et al.
Published: (2023)
by: Zhuang, Honglei, et al.
Published: (2023)
Domain Adaptation for Dense Retrieval and Conversational Dense Retrieval through Self-Supervision by Meticulous Pseudo-Relevance Labeling
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
Towards Detecting and Mitigating Cognitive Bias in Spoken Conversational Search
by: Ji, Kaixin, et al.
Published: (2024)
by: Ji, Kaixin, et al.
Published: (2024)
Judging the Judges: A Collection of LLM-Generated Relevance Judgements
by: Rahmani, Hossein A., et al.
Published: (2025)
by: Rahmani, Hossein A., et al.
Published: (2025)
A Systematic Study of Pseudo-Relevance Feedback with LLMs
by: Jedidi, Nour, et al.
Published: (2026)
by: Jedidi, Nour, et al.
Published: (2026)
Detecting Cryptographically Relevant Software Packages with Collaborative LLMs
by: Hirsch, Eduard, et al.
Published: (2026)
by: Hirsch, Eduard, et al.
Published: (2026)
Leveraging LLMs to Evaluate Usefulness of Document
by: Wang, Xingzhu, et al.
Published: (2025)
by: Wang, Xingzhu, et al.
Published: (2025)
The Relevance of Item-Co-Exposure For Exposure Bias Mitigation
by: Krause, Thorsten, et al.
Published: (2024)
by: Krause, Thorsten, et al.
Published: (2024)
AutoMIR: Effective Zero-Shot Medical Information Retrieval without Relevance Labels
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
Migrating a Job Search Relevance Function
by: Mountain, Bennett, et al.
Published: (2025)
by: Mountain, Bennett, et al.
Published: (2025)
Addressing Labelled Data Scarcity: Taxonomy-Agnostic Annotation of PII Values in HTTP Traffic using LLMs
by: Cory, Thomas, et al.
Published: (2026)
by: Cory, Thomas, et al.
Published: (2026)
Generalized Pseudo-Relevance Feedback
by: Tu, Yiteng, et al.
Published: (2025)
by: Tu, Yiteng, et al.
Published: (2025)
Generative Pseudo-Labeling for Pre-Ranking with LLMs
by: Bi, Junyu, et al.
Published: (2026)
by: Bi, Junyu, et al.
Published: (2026)
How Relevance Emerges: Interpreting LoRA Fine-Tuning in Reranking LLMs
by: Nijasure, Atharva, et al.
Published: (2025)
by: Nijasure, Atharva, et al.
Published: (2025)
LLM-Assisted Pseudo-Relevance Feedback
by: Otero, David, et al.
Published: (2026)
by: Otero, David, et al.
Published: (2026)
LUCid: Redefining Relevance For Lifelong Personalization
by: Okite, Chimaobi, et al.
Published: (2026)
by: Okite, Chimaobi, et al.
Published: (2026)
Similar Items
-
Can Generative LLMs Create Query Variants for Test Collections? An Exploratory Study
by: Alaofi, Marwah, et al.
Published: (2025) -
Demographically-Inspired Query Variants Using an LLM
by: Alaofi, Marwah, et al.
Published: (2025) -
Generative Information Retrieval Evaluation
by: Alaofi, Marwah, et al.
Published: (2024) -
Online and Offline Evaluation in Search Clarification
by: Tavakoli, Leila, et al.
Published: (2024) -
Personalisation of Generic Library Search Results Using Student Enrolment Information
by: Alaofi, Marwah, et al.
Published: (2015)