Online and Offline Evaluation in Search Clarification
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tavakoli, Leila, Trippas, Johanne R., Zamani, Hamed, Scholer, Falk, Sanderson, Mark |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
par: Tavakoli, Leila, et autres
Publié: (2025)
par: Tavakoli, Leila, et autres
Publié: (2025)
Towards Investigating Biases in Spoken Conversational Search
par: Cherumanal, Sachin Pathiyan, et autres
Publié: (2024)
par: Cherumanal, Sachin Pathiyan, et autres
Publié: (2024)
Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment
par: Schnabel, Julian A., et autres
Publié: (2025)
par: Schnabel, Julian A., et autres
Publié: (2025)
LLMs can be Fooled into Labelling a Document as Relevant (best café near me; this paper is perfectly relevant)
par: Alaofi, Marwah, et autres
Publié: (2025)
par: Alaofi, Marwah, et autres
Publié: (2025)
Towards Detecting and Mitigating Cognitive Bias in Spoken Conversational Search
par: Ji, Kaixin, et autres
Publié: (2024)
par: Ji, Kaixin, et autres
Publié: (2024)
The Effects of Demographic Instructions on LLM Personas
par: de Paula, Angel Felipe Magnossão, et autres
Publié: (2025)
par: de Paula, Angel Felipe Magnossão, et autres
Publié: (2025)
Understanding Modality Preferences in Search Clarification
par: Tavakoli, Leila, et autres
Publié: (2024)
par: Tavakoli, Leila, et autres
Publié: (2024)
Walert: Putting Conversational Search Knowledge into Action by Building and Evaluating a Large Language Model-Powered Chatbot
par: Cherumanal, Sachin Pathiyan, et autres
Publié: (2024)
par: Cherumanal, Sachin Pathiyan, et autres
Publié: (2024)
Can Generative LLMs Create Query Variants for Test Collections? An Exploratory Study
par: Alaofi, Marwah, et autres
Publié: (2025)
par: Alaofi, Marwah, et autres
Publié: (2025)
Demographically-Inspired Query Variants Using an LLM
par: Alaofi, Marwah, et autres
Publié: (2025)
par: Alaofi, Marwah, et autres
Publié: (2025)
Can Users Detect Biases or Factual Errors in Generated Responses in Conversational Information-Seeking?
par: Łajewska, Weronika, et autres
Publié: (2024)
par: Łajewska, Weronika, et autres
Publié: (2024)
Can We Hide Machines in the Crowd? Quantifying Equivalence in LLM-in-the-loop Annotation Tasks
par: He, Jiaman, et autres
Publié: (2025)
par: He, Jiaman, et autres
Publié: (2025)
Characterising Topic Familiarity and Query Specificity Using Eye-Tracking Data
par: He, Jiaman, et autres
Publié: (2025)
par: He, Jiaman, et autres
Publié: (2025)
Explainability for Transparent Conversational Information-Seeking
par: Łajewska, Weronika, et autres
Publié: (2024)
par: Łajewska, Weronika, et autres
Publié: (2024)
Characterizing Personality from Eye-Tracking: The Role of Gaze and Its Absence in Interactive Search Environments
par: He, Jiaman, et autres
Publié: (2026)
par: He, Jiaman, et autres
Publié: (2026)
Towards a Search Engine for Machines: Unified Ranking for Multiple Retrieval-Augmented Large Language Models
par: Salemi, Alireza, et autres
Publié: (2024)
par: Salemi, Alireza, et autres
Publié: (2024)
Control Search Rankings, Control the World: What is a Good Search Engine?
par: Coghlan, Simon, et autres
Publié: (2025)
par: Coghlan, Simon, et autres
Publié: (2025)
Evaluating Retrieval Quality in Retrieval-Augmented Generation
par: Salemi, Alireza, et autres
Publié: (2024)
par: Salemi, Alireza, et autres
Publié: (2024)
ProCIS: A Benchmark for Proactive Retrieval in Conversations
par: Samarinas, Chris, et autres
Publié: (2024)
par: Samarinas, Chris, et autres
Publié: (2024)
Analyzing Coherency in Facet-based Clarification Prompt Generation for Search
par: Litvinov, Oleg, et autres
Publié: (2024)
par: Litvinov, Oleg, et autres
Publié: (2024)
Interactions with Generative Information Retrieval Systems
par: Aliannejadi, Mohammad, et autres
Publié: (2024)
par: Aliannejadi, Mohammad, et autres
Publié: (2024)
Open-Ended and Knowledge-Intensive Video Question Answering
par: Alam, Md Zarif Ul, et autres
Publié: (2025)
par: Alam, Md Zarif Ul, et autres
Publié: (2025)
Uncertainty Quantification for Retrieval-Augmented Reasoning
par: Soudani, Heydar, et autres
Publié: (2025)
par: Soudani, Heydar, et autres
Publié: (2025)
Scaling Sparse and Dense Retrieval in Decoder-Only LLMs
par: Zeng, Hansi, et autres
Publié: (2025)
par: Zeng, Hansi, et autres
Publié: (2025)
Learning to Rank for Multiple Retrieval-Augmented Models through Iterative Utility Maximization
par: Salemi, Alireza, et autres
Publié: (2024)
par: Salemi, Alireza, et autres
Publié: (2024)
Distillation and Refinement of Reasoning in Small Language Models for Document Re-ranking
par: Samarinas, Chris, et autres
Publié: (2025)
par: Samarinas, Chris, et autres
Publié: (2025)
Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedback
par: Alam, Md Zarif Ul, et autres
Publié: (2026)
par: Alam, Md Zarif Ul, et autres
Publié: (2026)
Metamorphic Evaluation of ChatGPT as a Recommender System
par: Khirbat, Madhurima, et autres
Publié: (2024)
par: Khirbat, Madhurima, et autres
Publié: (2024)
Evaluation of Agents under Simulated AI Marketplace Dynamics
par: Kim, To Eun, et autres
Publié: (2026)
par: Kim, To Eun, et autres
Publié: (2026)
Evaluating and Addressing Fairness Across User Groups in Negative Sampling for Recommender Systems
par: Xuan, Yueqing, et autres
Publié: (2023)
par: Xuan, Yueqing, et autres
Publié: (2023)
RAC: Retrieval-Augmented Clarification for Faithful Conversational Search
par: Kebir, Ahmed Rayane, et autres
Publié: (2026)
par: Kebir, Ahmed Rayane, et autres
Publié: (2026)
Generative Information Retrieval Evaluation
par: Alaofi, Marwah, et autres
Publié: (2024)
par: Alaofi, Marwah, et autres
Publié: (2024)
Total Recall QA: A Verifiable Evaluation Suite for Deep Research Agents
par: Rafiee, Mahta, et autres
Publié: (2026)
par: Rafiee, Mahta, et autres
Publié: (2026)
ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation
par: Salemi, Alireza, et autres
Publié: (2025)
par: Salemi, Alireza, et autres
Publié: (2025)
Scaling Laws for Cross-Encoder Reranking
par: Seetharaman, Rahul, et autres
Publié: (2026)
par: Seetharaman, Rahul, et autres
Publié: (2026)
CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search
par: Zeng, Hansi, et autres
Publié: (2026)
par: Zeng, Hansi, et autres
Publié: (2026)
Bridging Personalization and Control in Scientific Personalized Search
par: Mysore, Sheshera, et autres
Publié: (2024)
par: Mysore, Sheshera, et autres
Publié: (2024)
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
par: Tavakoli, Mohammad, et autres
Publié: (2025)
par: Tavakoli, Mohammad, et autres
Publié: (2025)
Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization
par: Zamani, Hamed, et autres
Publié: (2024)
par: Zamani, Hamed, et autres
Publié: (2024)
Learning from Natural Language Feedback for Personalized Question Answering
par: Salemi, Alireza, et autres
Publié: (2025)
par: Salemi, Alireza, et autres
Publié: (2025)
Documents similaires
-
Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
par: Tavakoli, Leila, et autres
Publié: (2025) -
Towards Investigating Biases in Spoken Conversational Search
par: Cherumanal, Sachin Pathiyan, et autres
Publié: (2024) -
Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment
par: Schnabel, Julian A., et autres
Publié: (2025) -
LLMs can be Fooled into Labelling a Document as Relevant (best café near me; this paper is perfectly relevant)
par: Alaofi, Marwah, et autres
Publié: (2025) -
Towards Detecting and Mitigating Cognitive Bias in Spoken Conversational Search
par: Ji, Kaixin, et autres
Publié: (2024)