Navigating Tomorrow: Reliably Assessing Large Language Models Performance on Future Event Prediction
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Nako, Petraq, Jatowt, Adam |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Wisdom of the Crowds in Forecasting: Forecast Summarization for Supporting Future Event Prediction
par: Saha, Anisha, et autres
Publié: (2025)
par: Saha, Anisha, et autres
Publié: (2025)
Analyzing the Role of Context in Forecasting with Large Language Models
par: Mutschlechner, Gerrit, et autres
Publié: (2025)
par: Mutschlechner, Gerrit, et autres
Publié: (2025)
Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring
par: Mozafari, Jamshid, et autres
Publié: (2026)
par: Mozafari, Jamshid, et autres
Publié: (2026)
Detecting Future-related Contexts of Entity Mentions
par: Prashar, Puneet, et autres
Publié: (2025)
par: Prashar, Puneet, et autres
Publié: (2025)
Wrong Answers Can Also Be Useful: PlausibleQA -- A Large-Scale QA Dataset with Answer Plausibility Scores
par: Mozafari, Jamshid, et autres
Publié: (2025)
par: Mozafari, Jamshid, et autres
Publié: (2025)
Evaluating Answer Reranking Strategies in Time-sensitive Question Answering
par: Kardan, Mehmet, et autres
Publié: (2025)
par: Kardan, Mehmet, et autres
Publié: (2025)
Context Convergence Improves Answering Inferential Questions
par: Mozafari, Jamshid, et autres
Publié: (2026)
par: Mozafari, Jamshid, et autres
Publié: (2026)
WikiHint: A Human-Annotated Dataset for Hint Ranking and Generation
par: Mozafari, Jamshid, et autres
Publié: (2024)
par: Mozafari, Jamshid, et autres
Publié: (2024)
PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian
par: Mozafari, Jamshid, et autres
Publié: (2026)
par: Mozafari, Jamshid, et autres
Publié: (2026)
How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillation
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions
par: Mozafari, Jamshid, et autres
Publié: (2025)
par: Mozafari, Jamshid, et autres
Publié: (2025)
Inferential Question Answering
par: Mozafari, Jamshid, et autres
Publié: (2026)
par: Mozafari, Jamshid, et autres
Publié: (2026)
Exploring Hint Generation Approaches in Open-Domain Question Answering
par: Mozafari, Jamshid, et autres
Publié: (2024)
par: Mozafari, Jamshid, et autres
Publié: (2024)
Multi-hop Question Answering
par: Mavi, Vaibhav, et autres
Publié: (2022)
par: Mavi, Vaibhav, et autres
Publié: (2022)
It's High Time: A Survey of Temporal Question Answering
par: Piryani, Bhawna, et autres
Publié: (2025)
par: Piryani, Bhawna, et autres
Publié: (2025)
TempRetriever: Fusion-based Temporal Dense Passage Retrieval for Time-Sensitive Questions
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
A Comprehensive Evaluation of Large Language Models on Temporal Event Forecasting
par: Chang, He, et autres
Publié: (2024)
par: Chang, He, et autres
Publié: (2024)
ArabicaQA: A Comprehensive Dataset for Arabic Question Answering
par: Abdallah, Abdelrahman, et autres
Publié: (2024)
par: Abdallah, Abdelrahman, et autres
Publié: (2024)
Assessing the Performance Gap Between Lexical and Semantic Models for Information Retrieval With Formulaic Legal Language
par: Mori, Larissa, et autres
Publié: (2025)
par: Mori, Larissa, et autres
Publié: (2025)
LLMTemporalComparator: A Tool for Analysing Differences in Temporal Adaptations of Large Language Models
par: Fritsch, Reinhard Friedrich, et autres
Publié: (2024)
par: Fritsch, Reinhard Friedrich, et autres
Publié: (2024)
Large Language Model-Powered Query-Driven Event Timeline Summarization in Industrial Search
par: Wang, Mingyue, et autres
Publié: (2026)
par: Wang, Mingyue, et autres
Publié: (2026)
GDLLM: A Global Distance-aware Modeling Approach Based on Large Language Models for Event Temporal Relation Extraction
par: Zhao, Jie, et autres
Publié: (2025)
par: Zhao, Jie, et autres
Publié: (2025)
RAG-Enhanced Large Language Models for Dynamic Content Expiration Prediction in Web Search
par: Chen, Tingyu, et autres
Publié: (2026)
par: Chen, Tingyu, et autres
Publié: (2026)
Link Prediction for Event Logs in the Process Industry
par: Zhukova, Anastasia, et autres
Publié: (2025)
par: Zhukova, Anastasia, et autres
Publié: (2025)
Assessing SPARQL capabilities of Large Language Models
par: Meyer, Lars-Peter, et autres
Publié: (2024)
par: Meyer, Lars-Peter, et autres
Publié: (2024)
M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions
par: Wang, Zheng, et autres
Publié: (2024)
par: Wang, Zheng, et autres
Publié: (2024)
Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search
par: DeVerna, Matthew R., et autres
Publié: (2025)
par: DeVerna, Matthew R., et autres
Publié: (2025)
Graph Fusion Across Languages using Large Language Models
par: Kyaw, Kaung Myat, et autres
Publié: (2026)
par: Kyaw, Kaung Myat, et autres
Publié: (2026)
BracketRank: Large Language Model Document Ranking via Reasoning-based Competitive Elimination
par: Abdallah, Abdelrahman, et autres
Publié: (2026)
par: Abdallah, Abdelrahman, et autres
Publié: (2026)
Hallucination Detection and Evaluation of Large Language Model
par: Zhang, Chenggong, et autres
Publié: (2025)
par: Zhang, Chenggong, et autres
Publié: (2025)
Metacognitive Retrieval-Augmented Large Language Models
par: Zhou, Yujia, et autres
Publié: (2024)
par: Zhou, Yujia, et autres
Publié: (2024)
Large Language Models as Evaluators for Recommendation Explanations
par: Zhang, Xiaoyu, et autres
Publié: (2024)
par: Zhang, Xiaoyu, et autres
Publié: (2024)
Enhance Graph Alignment for Large Language Models
par: Luo, Haitong, et autres
Publié: (2024)
par: Luo, Haitong, et autres
Publié: (2024)
Improving Text Embeddings with Large Language Models
par: Wang, Liang, et autres
Publié: (2023)
par: Wang, Liang, et autres
Publié: (2023)
A Study into Investigating Temporal Robustness of LLMs
par: Wallat, Jonas, et autres
Publié: (2025)
par: Wallat, Jonas, et autres
Publié: (2025)
Evaluating Large Language Models for Cross-Lingual Retrieval
par: Zuo, Longfei, et autres
Publié: (2025)
par: Zuo, Longfei, et autres
Publié: (2025)
Making Large Language Models Efficient Dense Retrievers
par: Lei, Yibin, et autres
Publié: (2025)
par: Lei, Yibin, et autres
Publié: (2025)
ConExion: Concept Extraction with Large Language Models
par: Norouzi, Ebrahim, et autres
Publié: (2025)
par: Norouzi, Ebrahim, et autres
Publié: (2025)
Documents similaires
-
Wisdom of the Crowds in Forecasting: Forecast Summarization for Supporting Future Event Prediction
par: Saha, Anisha, et autres
Publié: (2025) -
Analyzing the Role of Context in Forecasting with Large Language Models
par: Mutschlechner, Gerrit, et autres
Publié: (2025) -
Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring
par: Mozafari, Jamshid, et autres
Publié: (2026) -
Detecting Future-related Contexts of Entity Mentions
par: Prashar, Puneet, et autres
Publié: (2025) -
Wrong Answers Can Also Be Useful: PlausibleQA -- A Large-Scale QA Dataset with Answer Plausibility Scores
par: Mozafari, Jamshid, et autres
Publié: (2025)