HORIZON: A Benchmark for In-the-wild User Behaviour Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Goel, Arnav, Chitale, Pranjal A, Paliwal, Bhawna, Santra, Bishal, Sharma, Amit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
von: Chitale, Pranjal A., et al.
Veröffentlicht: (2025)
von: Chitale, Pranjal A., et al.
Veröffentlicht: (2025)
Quantifying Positional Biases in Text Embedding Models
von: Lee, Reagan J., et al.
Veröffentlicht: (2024)
von: Lee, Reagan J., et al.
Veröffentlicht: (2024)
HIRO: Hierarchical Information Retrieval Optimization
von: Goel, Krish, et al.
Veröffentlicht: (2024)
von: Goel, Krish, et al.
Veröffentlicht: (2024)
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
von: Opsahl-Ong, Krista, et al.
Veröffentlicht: (2026)
von: Opsahl-Ong, Krista, et al.
Veröffentlicht: (2026)
Attribution in Scientific Literature: New Benchmark and Methods
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
Understanding the Role of User Profile in the Personalization of Large Language Models
von: Wu, Bin, et al.
Veröffentlicht: (2024)
von: Wu, Bin, et al.
Veröffentlicht: (2024)
HARNESS-LM: A Three-Phase Training Recipe for Harnessing SLMs in Sponsored Search Retrieval
von: Gupta, Vipul, et al.
Veröffentlicht: (2026)
von: Gupta, Vipul, et al.
Veröffentlicht: (2026)
SymTax: Symbiotic Relationship and Taxonomy Fusion for Effective Citation Recommendation
von: Goyal, Karan, et al.
Veröffentlicht: (2024)
von: Goyal, Karan, et al.
Veröffentlicht: (2024)
MMTEB: Massive Multilingual Text Embedding Benchmark
von: Enevoldsen, Kenneth, et al.
Veröffentlicht: (2025)
von: Enevoldsen, Kenneth, et al.
Veröffentlicht: (2025)
LangProBe: a Language Programs Benchmark
von: Tan, Shangyin, et al.
Veröffentlicht: (2025)
von: Tan, Shangyin, et al.
Veröffentlicht: (2025)
Faux Polyglot: A Study on Information Disparity in Multilingual Large Language Models
von: Sharma, Nikhil, et al.
Veröffentlicht: (2024)
von: Sharma, Nikhil, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation
von: Yoon, Se-eun, et al.
Veröffentlicht: (2024)
von: Yoon, Se-eun, et al.
Veröffentlicht: (2024)
Benchmarking Prompt Sensitivity in Large Language Models
von: Razavi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Razavi, Amirhossein, et al.
Veröffentlicht: (2025)
Deep Learning Based Named Entity Recognition Models for Recipes
von: Goel, Mansi, et al.
Veröffentlicht: (2024)
von: Goel, Mansi, et al.
Veröffentlicht: (2024)
Using Large Language Models to Generate, Validate, and Apply User Intent Taxonomies
von: Shah, Chirag, et al.
Veröffentlicht: (2023)
von: Shah, Chirag, et al.
Veröffentlicht: (2023)
URAG: A Benchmark for Uncertainty Quantification in Retrieval-Augmented Large Language Models
von: Nguyen, Vinh, et al.
Veröffentlicht: (2026)
von: Nguyen, Vinh, et al.
Veröffentlicht: (2026)
Benchmarking Information Retrieval Models on Complex Retrieval Tasks
von: Killingback, Julian, et al.
Veröffentlicht: (2025)
von: Killingback, Julian, et al.
Veröffentlicht: (2025)
Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models
von: Shi, Zhengliang, et al.
Veröffentlicht: (2025)
von: Shi, Zhengliang, et al.
Veröffentlicht: (2025)
Memory-Based vs. Context-Only Conditioning Produces Distinct Behavioral Patterns in Stateful Personalization
von: Park, Junsoo, et al.
Veröffentlicht: (2026)
von: Park, Junsoo, et al.
Veröffentlicht: (2026)
Domain-Partitioned Hybrid RAG for Legal Reasoning: Toward Modular and Explainable Legal AI for India
von: Goel, Rakshita, et al.
Veröffentlicht: (2025)
von: Goel, Rakshita, et al.
Veröffentlicht: (2025)
OAEI-LLM: A Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching
von: Qiang, Zhangcheng, et al.
Veröffentlicht: (2024)
von: Qiang, Zhangcheng, et al.
Veröffentlicht: (2024)
TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
von: Ezerceli, Özay, et al.
Veröffentlicht: (2025)
von: Ezerceli, Özay, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models on Reference Extraction and Parsing in the Social Sciences and Humanities
von: Zhu, Yurui, et al.
Veröffentlicht: (2026)
von: Zhu, Yurui, et al.
Veröffentlicht: (2026)
MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)
von: Bahaj, Adil, et al.
Veröffentlicht: (2025)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
von: Jin, Zhuoran, et al.
Veröffentlicht: (2024)
von: Jin, Zhuoran, et al.
Veröffentlicht: (2024)
Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability
von: Aggarwal, Shashank, et al.
Veröffentlicht: (2026)
von: Aggarwal, Shashank, et al.
Veröffentlicht: (2026)
Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study
von: Sui, Yuan, et al.
Veröffentlicht: (2023)
von: Sui, Yuan, et al.
Veröffentlicht: (2023)
GISA: A Benchmark for General Information-Seeking Assistant
von: Zhu, Yutao, et al.
Veröffentlicht: (2026)
von: Zhu, Yutao, et al.
Veröffentlicht: (2026)
ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
von: Su, Hongjin, et al.
Veröffentlicht: (2024)
von: Su, Hongjin, et al.
Veröffentlicht: (2024)
FinRetrieval: A Benchmark for Financial Data Retrieval by AI Agents
von: Kim, Eric Y., et al.
Veröffentlicht: (2026)
von: Kim, Eric Y., et al.
Veröffentlicht: (2026)
SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
von: Su, Weihang, et al.
Veröffentlicht: (2025)
von: Su, Weihang, et al.
Veröffentlicht: (2025)
The Massive Legal Embedding Benchmark (MLEB)
von: Butler, Umar, et al.
Veröffentlicht: (2025)
von: Butler, Umar, et al.
Veröffentlicht: (2025)
Benchmarking Retrieval-Augmented Generation for Chemistry
von: Zhong, Xianrui, et al.
Veröffentlicht: (2025)
von: Zhong, Xianrui, et al.
Veröffentlicht: (2025)
Exploring the Escalation of Source Bias in User, Data, and Recommender System Feedback Loop
von: Zhou, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2024)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
Towards Personalized Deep Research: Benchmarks and Evaluations
von: Liang, Yuan, et al.
Veröffentlicht: (2025)
von: Liang, Yuan, et al.
Veröffentlicht: (2025)
ALARB: An Arabic Legal Argument Reasoning Benchmark
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025)
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025)
Reasoning over User Preferences: Knowledge Graph-Augmented LLMs for Explainable Conversational Recommendations
von: Qiu, Zhangchi, et al.
Veröffentlicht: (2024)
von: Qiu, Zhangchi, et al.
Veröffentlicht: (2024)
Benchmarking Advanced Text Anonymisation Methods: A Comparative Study on Novel and Traditional Approaches
von: Asimopoulos, Dimitris, et al.
Veröffentlicht: (2024)
von: Asimopoulos, Dimitris, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
von: Chitale, Pranjal A., et al.
Veröffentlicht: (2025) -
Quantifying Positional Biases in Text Embedding Models
von: Lee, Reagan J., et al.
Veröffentlicht: (2024) -
HIRO: Hierarchical Information Retrieval Optimization
von: Goel, Krish, et al.
Veröffentlicht: (2024) -
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
von: Opsahl-Ong, Krista, et al.
Veröffentlicht: (2026) -
Attribution in Scientific Literature: New Benchmark and Methods
von: Saxena, Yash, et al.
Veröffentlicht: (2024)