A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
Fuente:
arXiv
Saved in:
| Main Authors: | Arabzadeh, Negar, Clarke, Charles L. A . |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
Adversarial Attacks against Neural Ranking Models via In-Context Learning
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
A Reproducibility Study of LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
ReFormeR: Learning and Applying Explicit Query Reformulation Patterns
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
Benchmarking Prompt Sensitivity in Large Language Models
by: Razavi, Amirhossein, et al.
Published: (2025)
by: Razavi, Amirhossein, et al.
Published: (2025)
A Comparison of Methods for Evaluating Generative IR
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
The Effect of Document Summarization on LLM-Based Relevance Judgments
by: Mohtadi, Samaneh, et al.
Published: (2025)
by: Mohtadi, Samaneh, et al.
Published: (2025)
Adapting Standard Retrieval Benchmarks to Evaluate Generated Answers
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis
by: Mohtadi, Samaneh, et al.
Published: (2026)
by: Mohtadi, Samaneh, et al.
Published: (2026)
exHarmony: Authorship and Citations for Benchmarking the Reviewer Assignment Problem
by: Ebrahimi, Sajad, et al.
Published: (2025)
by: Ebrahimi, Sajad, et al.
Published: (2025)
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
Query-driven Relevant Paragraph Extraction from Legal Judgments
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
EMPRA: Embedding Perturbation Rank Attack against Neural Ranking Models
by: Bigdeli, Amin, et al.
Published: (2024)
by: Bigdeli, Amin, et al.
Published: (2024)
Generative Information Retrieval Evaluation
by: Alaofi, Marwah, et al.
Published: (2024)
by: Alaofi, Marwah, et al.
Published: (2024)
Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
Mitigating the Threshold Priming Effect in Large Language Model-Based Relevance Judgments via Personality Infusing
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
Can We Use Large Language Models to Fill Relevance Judgment Holes?
by: Abbasiantaeb, Zahra, et al.
Published: (2024)
by: Abbasiantaeb, Zahra, et al.
Published: (2024)
Peerispect: Claim Verification in Scientific Peer Reviews
by: Ghorbanpour, Ali, et al.
Published: (2026)
by: Ghorbanpour, Ali, et al.
Published: (2026)
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Offline Evaluation of Set-Based Text-to-Image Generation
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
A Comparative Analysis of Faithfulness Metrics and Humans in Citation Evaluation
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
Judging with Personality and Confidence: A Study on Personality-Conditioned LLM Relevance Assessment
by: Chen, Nuo, et al.
Published: (2026)
by: Chen, Nuo, et al.
Published: (2026)
Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
by: Amirshahi, Shakiba, et al.
Published: (2025)
by: Amirshahi, Shakiba, et al.
Published: (2025)
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs
by: Zhang, Hengran, et al.
Published: (2024)
by: Zhang, Hengran, et al.
Published: (2024)
Improving Topic Relevance Model by Mix-structured Summarization and LLM-based Data Augmentation
by: Liu, Yizhu, et al.
Published: (2024)
by: Liu, Yizhu, et al.
Published: (2024)
AR-Med: Automated Relevance Enhancement in Medical Search via LLM-Driven Information Augmentation
by: Wang, Chuyue, et al.
Published: (2025)
by: Wang, Chuyue, et al.
Published: (2025)
Robust Interaction-Based Relevance Modeling for Online e-Commerce Search
by: Chen, Ben, et al.
Published: (2024)
by: Chen, Ben, et al.
Published: (2024)
CSRM-LLM: Embracing Multilingual LLMs for Cold-Start Relevance Matching in Emerging E-commerce Markets
by: Wang, Yujing, et al.
Published: (2025)
by: Wang, Yujing, et al.
Published: (2025)
An Investigation of Prompt Variations for Zero-shot LLM-based Rankers
by: Sun, Shuoqi, et al.
Published: (2024)
by: Sun, Shuoqi, et al.
Published: (2024)
On the Robustness of LLM-Based Dense Retrievers: A Systematic Analysis of Generalizability and Stability
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
Led to Mislead: Adversarial Content Injection for Attacks on Neural Ranking Models
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
by: Yang, Jheng-Hong, et al.
Published: (2024)
by: Yang, Jheng-Hong, et al.
Published: (2024)
Ranked List Truncation for Large Language Model-based Re-Ranking
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
A Systematic Study of Pseudo-Relevance Feedback with LLMs
by: Jedidi, Nour, et al.
Published: (2026)
by: Jedidi, Nour, et al.
Published: (2026)
LLM, Reporting In! Medical Information Extraction Across Prompting, Fine-tuning and Post-correction
by: Belmadani, Ikram, et al.
Published: (2025)
by: Belmadani, Ikram, et al.
Published: (2025)
Similar Items
-
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025) -
QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2025) -
Adversarial Attacks against Neural Ranking Models via In-Context Learning
by: Bigdeli, Amin, et al.
Published: (2025) -
A Reproducibility Study of LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2026) -
ReFormeR: Learning and Applying Explicit Query Reformulation Patterns
by: Bigdeli, Amin, et al.
Published: (2026)