FOReCAst: The Future Outcome Reasoning and Confidence Assessment Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Zhangdie, Ding, Zifeng, Vlachos, Andreas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do Language Models Update their Forecasts with New Information?
von: Yuan, Zhangdie, et al.
Veröffentlicht: (2025)
von: Yuan, Zhangdie, et al.
Veröffentlicht: (2025)
Capturing Symmetry and Antisymmetry in Language Models through Symmetry-Aware Training Objectives
von: Yuan, Zhangdie, et al.
Veröffentlicht: (2025)
von: Yuan, Zhangdie, et al.
Veröffentlicht: (2025)
Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
von: Ding, Zifeng, et al.
Veröffentlicht: (2025)
von: Ding, Zifeng, et al.
Veröffentlicht: (2025)
An LLM Feature-based Framework for Dialogue Constructiveness Assessment
von: Zhou, Lexin, et al.
Veröffentlicht: (2024)
von: Zhou, Lexin, et al.
Veröffentlicht: (2024)
Language Models as Hierarchy Encoders
von: He, Yuan, et al.
Veröffentlicht: (2024)
von: He, Yuan, et al.
Veröffentlicht: (2024)
AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
PRobELM: Plausibility Ranking Evaluation for Language Models
von: Yuan, Zhangdie, et al.
Veröffentlicht: (2024)
von: Yuan, Zhangdie, et al.
Veröffentlicht: (2024)
Confidence Geometry Reveals Trace-Level Correctness in Large Language Model Reasoning
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
Confidence in the Reasoning of Large Language Models
von: Pawitan, Yudi, et al.
Veröffentlicht: (2024)
von: Pawitan, Yudi, et al.
Veröffentlicht: (2024)
Confidence over Time: Confidence Calibration with Temporal Logic for Large Language Model Reasoning
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2026)
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2026)
Self-Training Large Language Models with Confident Reasoning
von: Jang, Hyosoon, et al.
Veröffentlicht: (2025)
von: Jang, Hyosoon, et al.
Veröffentlicht: (2025)
The Role of Ambiguity in Error Prediction via Uncertainty Quantification
von: Staliūnaitė, Ieva Raminta, et al.
Veröffentlicht: (2026)
von: Staliūnaitė, Ieva Raminta, et al.
Veröffentlicht: (2026)
Process Supervision of Confidence Margin for Calibrated LLM Reasoning
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026)
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026)
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
Ev2R: Evaluating Evidence Retrieval in Automated Fact-Checking
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2024)
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2024)
Outcome-based Exploration for LLM Reasoning
von: Song, Yuda, et al.
Veröffentlicht: (2025)
von: Song, Yuda, et al.
Veröffentlicht: (2025)
Causal Estimation of Tokenisation Bias
von: Lesci, Pietro, et al.
Veröffentlicht: (2025)
von: Lesci, Pietro, et al.
Veröffentlicht: (2025)
LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2025)
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2025)
Show Your Work with Confidence: Confidence Bands for Tuning Curves
von: Lourie, Nicholas, et al.
Veröffentlicht: (2023)
von: Lourie, Nicholas, et al.
Veröffentlicht: (2023)
Scaling Open-Ended Reasoning to Predict the Future
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2026)
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2026)
A Context-Aware Dual-Metric Framework for Confidence Estimation in Large Language Models
von: Yuan, Mingruo, et al.
Veröffentlicht: (2025)
von: Yuan, Mingruo, et al.
Veröffentlicht: (2025)
TCP: a Benchmark for Temporal Constraint-Based Planning
von: Ding, Zifeng, et al.
Veröffentlicht: (2025)
von: Ding, Zifeng, et al.
Veröffentlicht: (2025)
Early Stopping for Large Reasoning Models via Confidence Dynamics
von: Hosseini, Parsa, et al.
Veröffentlicht: (2026)
von: Hosseini, Parsa, et al.
Veröffentlicht: (2026)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
von: Lyu, Chengqi, et al.
Veröffentlicht: (2025)
von: Lyu, Chengqi, et al.
Veröffentlicht: (2025)
Benchmarking and Understanding Compositional Relational Reasoning of LLMs
von: Ni, Ruikang, et al.
Veröffentlicht: (2024)
von: Ni, Ruikang, et al.
Veröffentlicht: (2024)
ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
von: Qiao, Ziqing, et al.
Veröffentlicht: (2025)
von: Qiao, Ziqing, et al.
Veröffentlicht: (2025)
AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
von: Liu, Yuliang, et al.
Veröffentlicht: (2025)
von: Liu, Yuliang, et al.
Veröffentlicht: (2025)
AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the Web
von: Cao, Rui, et al.
Veröffentlicht: (2025)
von: Cao, Rui, et al.
Veröffentlicht: (2025)
Multicalibration for Confidence Scoring in LLMs
von: Detommaso, Gianluca, et al.
Veröffentlicht: (2024)
von: Detommaso, Gianluca, et al.
Veröffentlicht: (2024)
Benchmarking Large Language Models for Math Reasoning Tasks
von: Seßler, Kathrin, et al.
Veröffentlicht: (2024)
von: Seßler, Kathrin, et al.
Veröffentlicht: (2024)
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments
von: Chakravarty, Abhirup, et al.
Veröffentlicht: (2025)
von: Chakravarty, Abhirup, et al.
Veröffentlicht: (2025)
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
von: Ma, Zhengzhao, et al.
Veröffentlicht: (2026)
von: Ma, Zhengzhao, et al.
Veröffentlicht: (2026)
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
von: Saha, Anisha, et al.
Veröffentlicht: (2025)
von: Saha, Anisha, et al.
Veröffentlicht: (2025)
Are LLM Decisions Faithful to Verbal Confidence?
von: Wang, Jiawei, et al.
Veröffentlicht: (2026)
von: Wang, Jiawei, et al.
Veröffentlicht: (2026)
Think Dense, Not Long: Dynamic Decoupled Conditional Advantage for Efficient Reasoning
von: Peng, Keqin, et al.
Veröffentlicht: (2026)
von: Peng, Keqin, et al.
Veröffentlicht: (2026)
VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
von: Dang, Vy Tuong, et al.
Veröffentlicht: (2025)
von: Dang, Vy Tuong, et al.
Veröffentlicht: (2025)
Are Large Language Models Good Temporal Graph Learners?
von: Huang, Shenyang, et al.
Veröffentlicht: (2025)
von: Huang, Shenyang, et al.
Veröffentlicht: (2025)
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning
von: Ding, Fei, et al.
Veröffentlicht: (2026)
von: Ding, Fei, et al.
Veröffentlicht: (2026)
Can Large Language Models Generalize Procedures Across Representations?
von: Lin, Fangru, et al.
Veröffentlicht: (2026)
von: Lin, Fangru, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Do Language Models Update their Forecasts with New Information?
von: Yuan, Zhangdie, et al.
Veröffentlicht: (2025) -
Capturing Symmetry and Antisymmetry in Language Models through Symmetry-Aware Training Objectives
von: Yuan, Zhangdie, et al.
Veröffentlicht: (2025) -
Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
von: Ding, Zifeng, et al.
Veröffentlicht: (2025) -
An LLM Feature-based Framework for Dialogue Constructiveness Assessment
von: Zhou, Lexin, et al.
Veröffentlicht: (2024) -
Language Models as Hierarchy Encoders
von: He, Yuan, et al.
Veröffentlicht: (2024)