ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Karger, Ezra, Bastani, Houtan, Yueh-Han, Chen, Jacobs, Zachary, Halawi, Danny, Zhang, Fred, Tetlock, Philip E. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Approaching Human-Level Forecasting with Language Models
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024)
by: Schoenegger, Philipp, et al.
Published: (2024)
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024)
by: Schoenegger, Philipp, et al.
Published: (2024)
Prompt Engineering Large Language Models' Forecasting Capabilities
by: Schoenegger, Philipp, et al.
Published: (2025)
by: Schoenegger, Philipp, et al.
Published: (2025)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
by: Halawi, Danny, et al.
Published: (2023)
by: Halawi, Danny, et al.
Published: (2023)
Bench to the Future: A Pastcasting Benchmark for Forecasting Agents
by: FutureSearch, et al.
Published: (2025)
by: FutureSearch, et al.
Published: (2025)
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025)
by: Pimpale, Govind, et al.
Published: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
Are AI Capabilities Increasing Exponentially? A Competing Hypothesis
by: Ge, Haosen, et al.
Published: (2026)
by: Ge, Haosen, et al.
Published: (2026)
What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
by: Jang, Jinkwan, et al.
Published: (2026)
by: Jang, Jinkwan, et al.
Published: (2026)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
by: Zhu, Hengchuan, et al.
Published: (2025)
by: Zhu, Hengchuan, et al.
Published: (2025)
PreScience: A Benchmark for Forecasting Scientific Contributions
by: Ajith, Anirudh, et al.
Published: (2026)
by: Ajith, Anirudh, et al.
Published: (2026)
Dominion: A New Frontier for AI Research
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
EpiCastBench: Datasets and Benchmarks for Multivariate Epidemic Forecasting
by: Panja, Madhurima, et al.
Published: (2026)
by: Panja, Madhurima, et al.
Published: (2026)
AI Idea Bench 2025: AI Research Idea Generation Benchmark
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
by: Ezra, Elon, et al.
Published: (2025)
by: Ezra, Elon, et al.
Published: (2025)
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
by: Han, Wenhan, et al.
Published: (2025)
by: Han, Wenhan, et al.
Published: (2025)
PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data
by: Cheng, Pu, et al.
Published: (2026)
by: Cheng, Pu, et al.
Published: (2026)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
by: Bragg, Jonathan, et al.
Published: (2025)
by: Bragg, Jonathan, et al.
Published: (2025)
CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
by: Sun, Guangzhi, et al.
Published: (2025)
by: Sun, Guangzhi, et al.
Published: (2025)
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
by: Macina, Jakub, et al.
Published: (2025)
by: Macina, Jakub, et al.
Published: (2025)
BizFinBench.v2: A Unified Dual-Mode Bilingual Benchmark for Expert-Level Financial Capability Alignment
by: Guo, Xin, et al.
Published: (2026)
by: Guo, Xin, et al.
Published: (2026)
UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
by: Fouda, Aya E., et al.
Published: (2025)
by: Fouda, Aya E., et al.
Published: (2025)
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
by: Prandi, Matteo, et al.
Published: (2025)
by: Prandi, Matteo, et al.
Published: (2025)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
by: He, Muyu, et al.
Published: (2026)
by: He, Muyu, et al.
Published: (2026)
CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
by: Deng, Hexuan, et al.
Published: (2026)
by: Deng, Hexuan, et al.
Published: (2026)
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
by: Wan, Luanbo, et al.
Published: (2025)
by: Wan, Luanbo, et al.
Published: (2025)
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
by: Siegel, Zachary S., et al.
Published: (2024)
by: Siegel, Zachary S., et al.
Published: (2024)
Evaluating LLMs on Real-World Forecasting Against Expert Forecasters
by: Lu, Janna
Published: (2025)
by: Lu, Janna
Published: (2025)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Forecasting Clinical Risk from Textual Time Series: Structuring Narratives for Temporal AI in Healthcare
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models
by: Zbeeb, Mohammad, et al.
Published: (2025)
by: Zbeeb, Mohammad, et al.
Published: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges
by: Wang, Zi-Han, et al.
Published: (2026)
by: Wang, Zi-Han, et al.
Published: (2026)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
Instructing the Architecture Search for Spatial-temporal Sequence Forecasting with LLM
by: Xue, Xin, et al.
Published: (2025)
by: Xue, Xin, et al.
Published: (2025)
SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
From Text to Forecasts: Bridging Modality Gap with Temporal Evolution Semantic Space
by: Li, Lehui, et al.
Published: (2026)
by: Li, Lehui, et al.
Published: (2026)
Similar Items
-
Approaching Human-Level Forecasting with Language Models
by: Halawi, Danny, et al.
Published: (2024) -
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024) -
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024) -
Prompt Engineering Large Language Models' Forecasting Capabilities
by: Schoenegger, Philipp, et al.
Published: (2025) -
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
by: Halawi, Danny, et al.
Published: (2023)