Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Hui, Teehan, Ryan, Ren, Mengye |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
by: Dai, Hui, et al.
Published: (2026)
by: Dai, Hui, et al.
Published: (2026)
Context Tuning for In-Context Optimization
by: Lu, Jack, et al.
Published: (2025)
by: Lu, Jack, et al.
Published: (2025)
CoLLEGe: Concept Embedding Generation for Large Language Models
by: Teehan, Ryan, et al.
Published: (2024)
by: Teehan, Ryan, et al.
Published: (2024)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
When Does Verification Pay Off? A Closer Look at LLMs as Solution Verifiers
by: Lu, Jack, et al.
Published: (2025)
by: Lu, Jack, et al.
Published: (2025)
Improving Autoregressive Training with Dynamic Oracles
by: Yang, Jianing, et al.
Published: (2024)
by: Yang, Jianing, et al.
Published: (2024)
A General Framework for Inference-time Scaling and Steering of Diffusion Models
by: Singhal, Raghav, et al.
Published: (2025)
by: Singhal, Raghav, et al.
Published: (2025)
ComPO: Preference Alignment via Comparison Oracles
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
by: Li, Changhao, et al.
Published: (2024)
by: Li, Changhao, et al.
Published: (2024)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
by: Chiu, Yu Ying, et al.
Published: (2024)
by: Chiu, Yu Ying, et al.
Published: (2024)
Culinary Class Wars: Evaluating LLMs using ASH in Cuisine Transfer Task
by: Lee, Hoonick, et al.
Published: (2024)
by: Lee, Hoonick, et al.
Published: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
NewsInterview: a Dataset and a Playground to Evaluate LLMs' Ground Gap via Informational Interviews
by: Spangher, Alexander, et al.
Published: (2024)
by: Spangher, Alexander, et al.
Published: (2024)
Continuous Approximations for Improving Quantization Aware Training of LLMs
by: Li, He, et al.
Published: (2024)
by: Li, He, et al.
Published: (2024)
Towards Practical Tool Usage for Continually Learning LLMs
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
by: Azmat, Muneeza, et al.
Published: (2025)
by: Azmat, Muneeza, et al.
Published: (2025)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
$L^*LM$: Learning Automata from Examples using Natural Language Oracles
by: Vazquez-Chanlatte, Marcell, et al.
Published: (2024)
by: Vazquez-Chanlatte, Marcell, et al.
Published: (2024)
PersonaGym: Evaluating Persona Agents and LLMs
by: Samuel, Vinay, et al.
Published: (2024)
by: Samuel, Vinay, et al.
Published: (2024)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
by: Davoodi, Arash Gholami, et al.
Published: (2024)
by: Davoodi, Arash Gholami, et al.
Published: (2024)
Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
by: Kawakami, Wataru, et al.
Published: (2025)
by: Kawakami, Wataru, et al.
Published: (2025)
SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMs
by: Jayasuriya, Dinithi, et al.
Published: (2025)
by: Jayasuriya, Dinithi, et al.
Published: (2025)
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
by: Cohen-Inger, Nurit, et al.
Published: (2025)
by: Cohen-Inger, Nurit, et al.
Published: (2025)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
by: Li, Aochong Oliver, et al.
Published: (2025)
by: Li, Aochong Oliver, et al.
Published: (2025)
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
by: Li, Haoran, et al.
Published: (2026)
by: Li, Haoran, et al.
Published: (2026)
EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
by: Fleshman, William, et al.
Published: (2024)
by: Fleshman, William, et al.
Published: (2024)
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
by: Hua, Andong, et al.
Published: (2025)
by: Hua, Andong, et al.
Published: (2025)
Evaluating Long Range Dependency Handling in Code Generation LLMs
by: Assogba, Yannick, et al.
Published: (2024)
by: Assogba, Yannick, et al.
Published: (2024)
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
by: Ren, Yanwei, et al.
Published: (2025)
by: Ren, Yanwei, et al.
Published: (2025)
A Unified Framework with Novel Metrics for Evaluating the Effectiveness of XAI Techniques in LLMs
by: Mersha, Melkamu Abay, et al.
Published: (2025)
by: Mersha, Melkamu Abay, et al.
Published: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Each Graph is a New Language: Graph Learning with LLMs
by: Zhou, Huachi, et al.
Published: (2025)
by: Zhou, Huachi, et al.
Published: (2025)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
by: He, Shenghua, et al.
Published: (2025)
by: He, Shenghua, et al.
Published: (2025)
ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining
by: Kim, Seonwu, et al.
Published: (2025)
by: Kim, Seonwu, et al.
Published: (2025)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
by: Yuan, Jiaqing, et al.
Published: (2024)
by: Yuan, Jiaqing, et al.
Published: (2024)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
by: Liu, Yijun, et al.
Published: (2024)
by: Liu, Yijun, et al.
Published: (2024)
Evaluating LLMs on Real-World Forecasting Against Expert Forecasters
by: Lu, Janna
Published: (2025)
by: Lu, Janna
Published: (2025)
Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
by: Abrams, Mitchell, et al.
Published: (2026)
by: Abrams, Mitchell, et al.
Published: (2026)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
by: Wang, Ganghua, et al.
Published: (2025)
by: Wang, Ganghua, et al.
Published: (2025)
Similar Items
-
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
by: Dai, Hui, et al.
Published: (2026) -
Context Tuning for In-Context Optimization
by: Lu, Jack, et al.
Published: (2025) -
CoLLEGe: Concept Embedding Generation for Large Language Models
by: Teehan, Ryan, et al.
Published: (2024) -
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025) -
When Does Verification Pay Off? A Closer Look at LLMs as Solution Verifiers
by: Lu, Jack, et al.
Published: (2025)