Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Furong, Pu, Yuan, Guo, Finn, Agrawal, Monica
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917146617970688
author Jia, Furong
Pu, Yuan
Guo, Finn
Agrawal, Monica
author_facet Jia, Furong
Pu, Yuan
Guo, Finn
Agrawal, Monica
contents Large language models (LLMs) excel on multiple-choice clinical diagnosis benchmarks, yet it is unclear how much of this performance reflects underlying probabilistic reasoning. We study this through questions from MedQA, where the task is to select the most likely diagnosis. We introduce the Frequency-Based Probabilistic Ranker (FBPR), a lightweight method that scores options with a smoothed Naive Bayes over concept-diagnosis co-occurrence statistics from a large corpus. When co-occurrence statistics were sourced from the pretraining corpora for OLMo and Llama, FBPR achieves comparable performance to the corresponding LLMs pretrained on that same corpus. Direct LLM inference and FBPR largely get different questions correct, with an overlap only slightly above random chance, indicating complementary strengths of each method. These findings highlight the continued value of explicit probabilistic baselines: they provide a meaningful performance reference point and a complementary signal for potential hybridization. While the performance of LLMs seems to be driven by a mechanism other than simple frequency aggregation, we show that an approach similar to the historically grounded, low-complexity expert systems still accounts for a substantial portion of benchmark performance.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12868
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
Jia, Furong
Pu, Yuan
Guo, Finn
Agrawal, Monica
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) excel on multiple-choice clinical diagnosis benchmarks, yet it is unclear how much of this performance reflects underlying probabilistic reasoning. We study this through questions from MedQA, where the task is to select the most likely diagnosis. We introduce the Frequency-Based Probabilistic Ranker (FBPR), a lightweight method that scores options with a smoothed Naive Bayes over concept-diagnosis co-occurrence statistics from a large corpus. When co-occurrence statistics were sourced from the pretraining corpora for OLMo and Llama, FBPR achieves comparable performance to the corresponding LLMs pretrained on that same corpus. Direct LLM inference and FBPR largely get different questions correct, with an overlap only slightly above random chance, indicating complementary strengths of each method. These findings highlight the continued value of explicit probabilistic baselines: they provide a meaningful performance reference point and a complementary signal for potential hybridization. While the performance of LLMs seems to be driven by a mechanism other than simple frequency aggregation, we show that an approach similar to the historically grounded, low-complexity expert systems still accounts for a substantial portion of benchmark performance.
title Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.12868