Evaluating LLMs on Real-World Forecasting Against Expert Forecasters
Fuente:
arXiv
Guardado en:
| Autor principal: | Lu, Janna |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LeMoLE: LLM-Enhanced Mixture of Linear Experts for Time Series Forecasting
por: Zhang, Lingzheng, et al.
Publicado: (2024)
por: Zhang, Lingzheng, et al.
Publicado: (2024)
Forecasting Time Series with LLMs via Patch-Based Prompting and Decomposition
por: Bumb, Mayank, et al.
Publicado: (2025)
por: Bumb, Mayank, et al.
Publicado: (2025)
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
por: Karger, Ezra, et al.
Publicado: (2024)
por: Karger, Ezra, et al.
Publicado: (2024)
Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs
por: Zhao, Taibiao, et al.
Publicado: (2025)
por: Zhao, Taibiao, et al.
Publicado: (2025)
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
por: Dai, Hui, et al.
Publicado: (2026)
por: Dai, Hui, et al.
Publicado: (2026)
Rethinking Time Series Forecasting with LLMs via Nearest Neighbor Contrastive Learning
por: Bogahawatte, Jayanie, et al.
Publicado: (2024)
por: Bogahawatte, Jayanie, et al.
Publicado: (2024)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
por: Hou, Yufang, et al.
Publicado: (2024)
por: Hou, Yufang, et al.
Publicado: (2024)
Integrating Time Series into LLMs via Multi-layer Steerable Embedding Fusion for Enhanced Forecasting
por: Chen, Zhuomin, et al.
Publicado: (2025)
por: Chen, Zhuomin, et al.
Publicado: (2025)
Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation
por: Mule, Srujan P, et al.
Publicado: (2026)
por: Mule, Srujan P, et al.
Publicado: (2026)
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
por: Qian, Cheng, et al.
Publicado: (2025)
por: Qian, Cheng, et al.
Publicado: (2025)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
por: Zhuang, Haomin, et al.
Publicado: (2024)
por: Zhuang, Haomin, et al.
Publicado: (2024)
Consistency Checks for Language Model Forecasters
por: Paleka, Daniel, et al.
Publicado: (2024)
por: Paleka, Daniel, et al.
Publicado: (2024)
Structured Extraction of Real World Medical Knowledge using LLMs for Summarization and Search
por: Kim, Edward, et al.
Publicado: (2024)
por: Kim, Edward, et al.
Publicado: (2024)
Nexus : An Agentic Framework for Time Series Forecasting
por: Das, Sarkar Snigdha Sarathi, et al.
Publicado: (2026)
por: Das, Sarkar Snigdha Sarathi, et al.
Publicado: (2026)
Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
por: Han, Yixuan, et al.
Publicado: (2025)
por: Han, Yixuan, et al.
Publicado: (2025)
Are LLMs Ready for Real-World Materials Discovery?
por: Miret, Santiago, et al.
Publicado: (2024)
por: Miret, Santiago, et al.
Publicado: (2024)
Efficient Safety Retrofitting Against Jailbreaking for LLMs
por: Garcia-Gasulla, Dario, et al.
Publicado: (2025)
por: Garcia-Gasulla, Dario, et al.
Publicado: (2025)
Bench to the Future: A Pastcasting Benchmark for Forecasting Agents
por: FutureSearch, et al.
Publicado: (2025)
por: FutureSearch, et al.
Publicado: (2025)
Large Language Models for Medical Forecasting -- Foresight 2
por: Kraljevic, Zeljko, et al.
Publicado: (2024)
por: Kraljevic, Zeljko, et al.
Publicado: (2024)
ThinkTank-ME: A Multi-Expert Framework for Middle East Event Forecasting
por: Li, Haoxuan, et al.
Publicado: (2026)
por: Li, Haoxuan, et al.
Publicado: (2026)
When Does Multimodality Lead to Better Time Series Forecasting?
por: Zhang, Xiyuan, et al.
Publicado: (2025)
por: Zhang, Xiyuan, et al.
Publicado: (2025)
AQuA -- Combining Experts' and Non-Experts' Views To Assess Deliberation Quality in Online Discussions Using LLMs
por: Behrendt, Maike, et al.
Publicado: (2024)
por: Behrendt, Maike, et al.
Publicado: (2024)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
por: Guo, Zikang, et al.
Publicado: (2025)
por: Guo, Zikang, et al.
Publicado: (2025)
Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective
por: Franzen, Daniel, et al.
Publicado: (2025)
por: Franzen, Daniel, et al.
Publicado: (2025)
Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
por: Bandarkar, Lucas, et al.
Publicado: (2026)
por: Bandarkar, Lucas, et al.
Publicado: (2026)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
por: Lu, Xudong, et al.
Publicado: (2024)
por: Lu, Xudong, et al.
Publicado: (2024)
STELLA: Guiding Large Language Models for Time Series Forecasting with Semantic Abstractions
por: Fan, Junjie, et al.
Publicado: (2025)
por: Fan, Junjie, et al.
Publicado: (2025)
Intervention-Aware Forecasting: Breaking Historical Limits from a System Perspective
por: Xu, Zhijian, et al.
Publicado: (2024)
por: Xu, Zhijian, et al.
Publicado: (2024)
GenTKG: Generative Forecasting on Temporal Knowledge Graph with Large Language Models
por: Liao, Ruotong, et al.
Publicado: (2023)
por: Liao, Ruotong, et al.
Publicado: (2023)
REAM: Merging Improves Pruning of Experts in LLMs
por: Jha, Saurav, et al.
Publicado: (2026)
por: Jha, Saurav, et al.
Publicado: (2026)
On the Performance of LLMs for Real Estate Appraisal
por: Geerts, Margot, et al.
Publicado: (2025)
por: Geerts, Margot, et al.
Publicado: (2025)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
por: Sicilia, Anthony, et al.
Publicado: (2024)
por: Sicilia, Anthony, et al.
Publicado: (2024)
Reasoning and Tools for Human-Level Forecasting
por: Hsieh, Elvis, et al.
Publicado: (2024)
por: Hsieh, Elvis, et al.
Publicado: (2024)
Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
por: Ding, Zifeng, et al.
Publicado: (2025)
por: Ding, Zifeng, et al.
Publicado: (2025)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
por: Riachi, Roland, et al.
Publicado: (2025)
por: Riachi, Roland, et al.
Publicado: (2025)
RePST: Language Model Empowered Spatio-Temporal Forecasting via Semantic-Oriented Reprogramming
por: Wang, Hao, et al.
Publicado: (2024)
por: Wang, Hao, et al.
Publicado: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts
por: Mujahid, Zain Muhammad, et al.
Publicado: (2025)
por: Mujahid, Zain Muhammad, et al.
Publicado: (2025)
Agentic Reinforcement Learning for Real-World Code Repair
por: Zhu, Siyu, et al.
Publicado: (2025)
por: Zhu, Siyu, et al.
Publicado: (2025)
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
por: Marius, Dumitran Adrian, et al.
Publicado: (2025)
por: Marius, Dumitran Adrian, et al.
Publicado: (2025)
Ejemplares similares
-
LeMoLE: LLM-Enhanced Mixture of Linear Experts for Time Series Forecasting
por: Zhang, Lingzheng, et al.
Publicado: (2024) -
Forecasting Time Series with LLMs via Patch-Based Prompting and Decomposition
por: Bumb, Mayank, et al.
Publicado: (2025) -
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
por: Karger, Ezra, et al.
Publicado: (2024) -
Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs
por: Zhao, Taibiao, et al.
Publicado: (2025) -
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
por: Dai, Hui, et al.
Publicado: (2026)