Pitfalls in Evaluating Language Model Forecasters
Fuente:
arXiv
Saved in:
| Main Authors: | Paleka, Daniel, Goel, Shashwat, Geiping, Jonas, Tramèr, Florian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
A Comparative Evaluation of Quantification Methods
by: Schumacher, Tobias, et al.
Published: (2021)
by: Schumacher, Tobias, et al.
Published: (2021)
Consistency Checks for Language Model Forecasters
by: Paleka, Daniel, et al.
Published: (2024)
by: Paleka, Daniel, et al.
Published: (2024)
Approaching Human-Level Forecasting with Language Models
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
Large Language Model Distilling Medication Recommendation Model
by: Liu, Qidong, et al.
Published: (2024)
by: Liu, Qidong, et al.
Published: (2024)
Metadata Extraction Leveraging Large Language Models
by: Han, Cuize, et al.
Published: (2025)
by: Han, Cuize, et al.
Published: (2025)
Language-Model Prior Overcomes Cold-Start Items
by: Wang, Shiyu, et al.
Published: (2024)
by: Wang, Shiyu, et al.
Published: (2024)
Large Language Models for Next Point-of-Interest Recommendation
by: Li, Peibo, et al.
Published: (2024)
by: Li, Peibo, et al.
Published: (2024)
PathGPT: Reframing Path Recommendation as a Natural Language Generation Task with Retrieval-Augmented Language Models
by: Marcelyn, Steeve Cuthbert, et al.
Published: (2025)
by: Marcelyn, Steeve Cuthbert, et al.
Published: (2025)
Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
by: Portes, Jacob, et al.
Published: (2025)
by: Portes, Jacob, et al.
Published: (2025)
PAP-REC: Personalized Automatic Prompt for Recommendation Language Model
by: Li, Zelong, et al.
Published: (2024)
by: Li, Zelong, et al.
Published: (2024)
Faithful Path Language Modeling for Explainable Recommendation over Knowledge Graph
by: Balloccu, Giacomo, et al.
Published: (2023)
by: Balloccu, Giacomo, et al.
Published: (2023)
Large Language Models as Recommender Systems: A Study of Popularity Bias
by: Lichtenberg, Jan Malte, et al.
Published: (2024)
by: Lichtenberg, Jan Malte, et al.
Published: (2024)
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
by: Tranheden, Wilhelm, et al.
Published: (2026)
by: Tranheden, Wilhelm, et al.
Published: (2026)
When Search Engine Services meet Large Language Models: Visions and Challenges
by: Xiong, Haoyi, et al.
Published: (2024)
by: Xiong, Haoyi, et al.
Published: (2024)
BEAR: Towards Beam-Search-Aware Optimization for Recommendation with Large Language Models
by: Yang, Weiqin, et al.
Published: (2026)
by: Yang, Weiqin, et al.
Published: (2026)
Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy
by: Zhao, Yao, et al.
Published: (2023)
by: Zhao, Yao, et al.
Published: (2023)
Modeling User Behavior from Adaptive Surveys with Supplemental Context
by: Shukla, Aman, et al.
Published: (2025)
by: Shukla, Aman, et al.
Published: (2025)
Exploring the Potential of Large Language Models in Public Transportation: San Antonio Case Study
by: Jonnala, Ramya, et al.
Published: (2025)
by: Jonnala, Ramya, et al.
Published: (2025)
STAR: A Simple Training-free Approach for Recommendations using Large Language Models
by: Lee, Dong-Ho, et al.
Published: (2024)
by: Lee, Dong-Ho, et al.
Published: (2024)
Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation
by: Li, Peibo, et al.
Published: (2025)
by: Li, Peibo, et al.
Published: (2025)
LegalPro-BERT: Classification of Legal Provisions by fine-tuning BERT Large Language Model
by: Tewari, Amit
Published: (2024)
by: Tewari, Amit
Published: (2024)
GAugLLM: Improving Graph Contrastive Learning for Text-Attributed Graphs with Large Language Models
by: Fang, Yi, et al.
Published: (2024)
by: Fang, Yi, et al.
Published: (2024)
VALUE: Value-Aware Large Language Model for Query Rewriting via Weighted Trie in Sponsored Search
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
Hybrid Intent-Aware Personalization with Machine Learning and RAG-Enabled Large Language Models for Financial Services Marketing
by: Shanivendra, Akhil Chandra
Published: (2026)
by: Shanivendra, Akhil Chandra
Published: (2026)
Evaluating and Enhancing Large Language Models for Novelty Assessment in Scholarly Publications
by: Lin, Ethan, et al.
Published: (2024)
by: Lin, Ethan, et al.
Published: (2024)
Evaluating the Robustness of Dense Retrievers in Interdisciplinary Domains
by: Chaturvedi, Sarthak, et al.
Published: (2025)
by: Chaturvedi, Sarthak, et al.
Published: (2025)
Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation
by: Klearman, Andrew, et al.
Published: (2026)
by: Klearman, Andrew, et al.
Published: (2026)
A Comprehensive Survey of Evaluation Techniques for Recommendation Systems
by: Jadon, Aryan, et al.
Published: (2023)
by: Jadon, Aryan, et al.
Published: (2023)
Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
by: Fabbri, Francesco, et al.
Published: (2025)
by: Fabbri, Francesco, et al.
Published: (2025)
Does It Look Sequential? An Analysis of Datasets for Evaluation of Sequential Recommendations
by: Klenitskiy, Anton, et al.
Published: (2024)
by: Klenitskiy, Anton, et al.
Published: (2024)
Federated Vision-Language-Recommendation with Personalized Fusion
by: Li, Zhiwei, et al.
Published: (2024)
by: Li, Zhiwei, et al.
Published: (2024)
Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
by: Beel, Joeran, et al.
Published: (2025)
by: Beel, Joeran, et al.
Published: (2025)
RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition
by: Cofala, Tim, et al.
Published: (2025)
by: Cofala, Tim, et al.
Published: (2025)
GraphMatch: Fusing Language and Graph Representations in a Dynamic Two-Sided Work Marketplace
by: Sacha, Mikołaj, et al.
Published: (2025)
by: Sacha, Mikołaj, et al.
Published: (2025)
Reasoning and Tools for Human-Level Forecasting
by: Hsieh, Elvis, et al.
Published: (2024)
by: Hsieh, Elvis, et al.
Published: (2024)
Using text embedding models as text classifiers with medical data
by: Goel, Rishabh
Published: (2024)
by: Goel, Rishabh
Published: (2024)
From AutoRecSys to AutoRecLab: A Call to Build, Evaluate, and Govern Autonomous Recommender-Systems Research Labs
by: Beel, Joeran, et al.
Published: (2025)
by: Beel, Joeran, et al.
Published: (2025)
MBD: A Model-Based Debiasing Framework Across User, Content, and Model Dimensions
by: Li, Yuantong, et al.
Published: (2026)
by: Li, Yuantong, et al.
Published: (2026)
Delayed Feedback Modeling with Influence Functions
by: Ding, Chenlu, et al.
Published: (2025)
by: Ding, Chenlu, et al.
Published: (2025)
Similar Items
-
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025) -
A Comparative Evaluation of Quantification Methods
by: Schumacher, Tobias, et al.
Published: (2021) -
Consistency Checks for Language Model Forecasters
by: Paleka, Daniel, et al.
Published: (2024) -
Approaching Human-Level Forecasting with Language Models
by: Halawi, Danny, et al.
Published: (2024) -
Large Language Model Distilling Medication Recommendation Model
by: Liu, Qidong, et al.
Published: (2024)