Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.I
Fuente:
arXiv
Guardado en:
| Autores principales: | Oosterhuis, Harrie, Jagerman, Rolf, Qin, Zhen, Wang, Xuanhui, Bendersky, Michael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Optimizing Compound Retrieval Systems
por: Oosterhuis, Harrie, et al.
Publicado: (2025)
por: Oosterhuis, Harrie, et al.
Publicado: (2025)
Consolidating Ranking and Relevance Predictions of Large Language Models through Post-Processing
por: Yan, Le, et al.
Publicado: (2024)
por: Yan, Le, et al.
Publicado: (2024)
Learning to Rank with Variable Result Presentation Lengths
por: Knyazev, Norman, et al.
Publicado: (2025)
por: Knyazev, Norman, et al.
Publicado: (2025)
Practical and Robust Safety Guarantees for Advanced Counterfactual Learning to Rank
por: Gupta, Shashank, et al.
Publicado: (2024)
por: Gupta, Shashank, et al.
Publicado: (2024)
Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank
por: Gupta, Shashank, et al.
Publicado: (2024)
por: Gupta, Shashank, et al.
Publicado: (2024)
Estimating the Hessian Matrix of Ranking Objectives for Stochastic Learning to Rank with Gradient Boosted Trees
por: Kang, Jingwei, et al.
Publicado: (2024)
por: Kang, Jingwei, et al.
Publicado: (2024)
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting
por: Qin, Zhen, et al.
Publicado: (2023)
por: Qin, Zhen, et al.
Publicado: (2023)
Can Query Expansion Improve Generalization of Strong Cross-Encoder Rankers?
por: Li, Minghan, et al.
Publicado: (2023)
por: Li, Minghan, et al.
Publicado: (2023)
Optimal Baseline Corrections for Off-Policy Contextual Bandits
por: Gupta, Shashank, et al.
Publicado: (2024)
por: Gupta, Shashank, et al.
Publicado: (2024)
A Non-Parametric Choice Model That Learns How Users Choose Between Recommended Options
por: Krause, Thorsten, et al.
Publicado: (2025)
por: Krause, Thorsten, et al.
Publicado: (2025)
Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization
por: Zamani, Hamed, et al.
Publicado: (2024)
por: Zamani, Hamed, et al.
Publicado: (2024)
A First Look at Selection Bias in Preference Elicitation for Recommendation
por: Gupta, Shashank, et al.
Publicado: (2024)
por: Gupta, Shashank, et al.
Publicado: (2024)
Harnessing Pairwise Ranking Prompting Through Sample-Efficient Ranking Distillation
por: Wu, Junru, et al.
Publicado: (2025)
por: Wu, Junru, et al.
Publicado: (2025)
Searching Personal Collections
por: Bendersky, Michael, et al.
Publicado: (2024)
por: Bendersky, Michael, et al.
Publicado: (2024)
Is Interpretable Machine Learning Effective at Feature Selection for Neural Learning-to-Rank?
por: Lyu, Lijun, et al.
Publicado: (2024)
por: Lyu, Lijun, et al.
Publicado: (2024)
Adaptive Orchestration of Modular Generative Information Access Systems
por: Hoveyda, Mohanna, et al.
Publicado: (2025)
por: Hoveyda, Mohanna, et al.
Publicado: (2025)
Beyond Yes and No: Improving Zero-Shot LLM Rankers via Scoring Fine-Grained Relevance Labels
por: Zhuang, Honglei, et al.
Publicado: (2023)
por: Zhuang, Honglei, et al.
Publicado: (2023)
Following the Eye-Tracking Evidence: Established Web-Search Assumptions Fail in Carousel Interfaces
por: Kang, Jingwei, et al.
Publicado: (2026)
por: Kang, Jingwei, et al.
Publicado: (2026)
Beyond Static Evaluation: Rethinking the Assessment of Personalized Agent Adaptability in Information Retrieval
por: Kaur, Kirandeep, et al.
Publicado: (2025)
por: Kaur, Kirandeep, et al.
Publicado: (2025)
On the Reliability of Sampling Strategies in Offline Recommender Evaluation
por: Pereira, Bruno L., et al.
Publicado: (2025)
por: Pereira, Bruno L., et al.
Publicado: (2025)
Rethinking Click Models in Light of Carousel Interfaces: Theory-Based Categorization and Design of Click Models
por: Kang, Jingwei, et al.
Publicado: (2025)
por: Kang, Jingwei, et al.
Publicado: (2025)
RETLLM: Training and Data-Free MLLMs for Multimodal Information Retrieval
por: Su, Dawei, et al.
Publicado: (2026)
por: Su, Dawei, et al.
Publicado: (2026)
Hypencoder: Hypernetworks for Information Retrieval
por: Killingback, Julian, et al.
Publicado: (2025)
por: Killingback, Julian, et al.
Publicado: (2025)
Evaluating Generative AI Tools for Personalized Offline Recommendations: A Comparative Study
por: Salinas-Buestan, Rafael, et al.
Publicado: (2025)
por: Salinas-Buestan, Rafael, et al.
Publicado: (2025)
UniPinRec: Unifying Generative Retrieval and Ranking at Pinterest Scale
por: Li, Hanyu, et al.
Publicado: (2026)
por: Li, Hanyu, et al.
Publicado: (2026)
Leveraging Retrieval-Augmented Generation for Persian University Knowledge Retrieval
por: Hemmat, Arshia, et al.
Publicado: (2024)
por: Hemmat, Arshia, et al.
Publicado: (2024)
Going Beyond Popularity and Positivity Bias: Correcting for Multifactorial Bias in Recommender Systems
por: Huang, Jin, et al.
Publicado: (2024)
por: Huang, Jin, et al.
Publicado: (2024)
Probing Ranking LLMs: A Mechanistic Analysis for Information Retrieval
por: Chowdhury, Tanya, et al.
Publicado: (2024)
por: Chowdhury, Tanya, et al.
Publicado: (2024)
A Survey of Controllable Learning: Methods and Applications in Information Retrieval
por: Shen, Chenglei, et al.
Publicado: (2024)
por: Shen, Chenglei, et al.
Publicado: (2024)
Scaling Laws for Embedding Dimension in Information Retrieval
por: Killingback, Julian, et al.
Publicado: (2026)
por: Killingback, Julian, et al.
Publicado: (2026)
MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation
por: Yu, Qinhan, et al.
Publicado: (2025)
por: Yu, Qinhan, et al.
Publicado: (2025)
Context Awareness Gate For Retrieval Augmented Generation
por: Heydari, Mohammad Hassan, et al.
Publicado: (2024)
por: Heydari, Mohammad Hassan, et al.
Publicado: (2024)
Not Just What, But When: Integrating Irregular Intervals to LLM for Sequential Recommendation
por: Du, Wei-Wei, et al.
Publicado: (2025)
por: Du, Wei-Wei, et al.
Publicado: (2025)
MST-R: Multi-Stage Tuning for Retrieval Systems and Metric Evaluation
por: Malviya, Yash, et al.
Publicado: (2024)
por: Malviya, Yash, et al.
Publicado: (2024)
Contradictions in Context: Challenges for Retrieval-Augmented Generation in Healthcare
por: Javadi, Saeedeh, et al.
Publicado: (2025)
por: Javadi, Saeedeh, et al.
Publicado: (2025)
RELIANCE: Reliable Ensemble Learning for Information and News Credibility Evaluation
por: Ramezani, Majid, et al.
Publicado: (2024)
por: Ramezani, Majid, et al.
Publicado: (2024)
Evaluation of LLM-based Strategies for the Extraction of Food Product Information from Online Shops
por: Brosch, Christoph, et al.
Publicado: (2025)
por: Brosch, Christoph, et al.
Publicado: (2025)
Re3: Learning to Balance Relevance & Recency for Temporal Information Retrieval
por: Cao, Jiawei, et al.
Publicado: (2025)
por: Cao, Jiawei, et al.
Publicado: (2025)
MFBE: Leveraging Multi-Field Information of FAQs for Efficient Dense Retrieval
por: Banerjee, Debopriyo, et al.
Publicado: (2023)
por: Banerjee, Debopriyo, et al.
Publicado: (2023)
Retrieval-Augmented Generation for Predicting Cellular Responses to Gene Perturbation
por: Di Francesco, Andrea Giuseppe, et al.
Publicado: (2026)
por: Di Francesco, Andrea Giuseppe, et al.
Publicado: (2026)
Ejemplares similares
-
Optimizing Compound Retrieval Systems
por: Oosterhuis, Harrie, et al.
Publicado: (2025) -
Consolidating Ranking and Relevance Predictions of Large Language Models through Post-Processing
por: Yan, Le, et al.
Publicado: (2024) -
Learning to Rank with Variable Result Presentation Lengths
por: Knyazev, Norman, et al.
Publicado: (2025) -
Practical and Robust Safety Guarantees for Advanced Counterfactual Learning to Rank
por: Gupta, Shashank, et al.
Publicado: (2024) -
Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank
por: Gupta, Shashank, et al.
Publicado: (2024)