Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Jia, Furong, Pu, Yuan, Guo, Finn, Agrawal, Monica |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Can Transformers Count to n?
by: Yehudai, Gilad, et al.
Published: (2024)
by: Yehudai, Gilad, et al.
Published: (2024)
Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
by: Agrawal, Vishakha, et al.
Published: (2025)
by: Agrawal, Vishakha, et al.
Published: (2025)
Proving that Cryptic Crossword Clue Answers are Correct
by: Andrews, Martin, et al.
Published: (2024)
by: Andrews, Martin, et al.
Published: (2024)
Addressing LLM Diversity by Infusing Random Concepts
by: Agrawal, Pulin, et al.
Published: (2026)
by: Agrawal, Pulin, et al.
Published: (2026)
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
Every Response Counts: Quantifying Uncertainty of LLM-based Multi-Agent Systems through Tensor Decomposition
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
A Lightweight LLM Framework for Disaster Humanitarian Information Classification
by: Jinzhen, Han, et al.
Published: (2026)
by: Jinzhen, Han, et al.
Published: (2026)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability
by: Naik, Ninad
Published: (2024)
by: Naik, Ninad
Published: (2024)
Joint Detection of Fraud and Concept Drift inOnline Conversations with LLM-Assisted Judgment
by: Senol, Ali, et al.
Published: (2025)
by: Senol, Ali, et al.
Published: (2025)
Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities
by: Ball, Thomas, et al.
Published: (2024)
by: Ball, Thomas, et al.
Published: (2024)
Token-Level LLM Collaboration via FusionRoute
by: Xiong, Nuoya, et al.
Published: (2026)
by: Xiong, Nuoya, et al.
Published: (2026)
Lightweight Retrieval-Augmented Generation and Large Language Model-Based Modeling for Scalable Patient-Trial Matching
by: Li, Xiaodi, et al.
Published: (2026)
by: Li, Xiaodi, et al.
Published: (2026)
Illuminate: A novel approach for depression detection with explainable analysis and proactive therapy using prompt engineering
by: Agrawal, Aryan
Published: (2024)
by: Agrawal, Aryan
Published: (2024)
FlowRL: Matching Reward Distributions for LLM Reasoning
by: Zhu, Xuekai, et al.
Published: (2025)
by: Zhu, Xuekai, et al.
Published: (2025)
Even GPT-5.2 Can't Count to Five: The Case for Zero-Error Horizons in Trustworthy LLMs
by: Sato, Ryoma
Published: (2026)
by: Sato, Ryoma
Published: (2026)
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Reliability
by: Guo, Kevin H., et al.
Published: (2026)
by: Guo, Kevin H., et al.
Published: (2026)
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
by: Wu, Xiaoyuan, et al.
Published: (2025)
by: Wu, Xiaoyuan, et al.
Published: (2025)
LLM-Match: An Open-Sourced Patient Matching Model Based on Large Language Models and Retrieval-Augmented Generation
by: Li, Xiaodi, et al.
Published: (2025)
by: Li, Xiaodi, et al.
Published: (2025)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
by: Yuan, Yurun, et al.
Published: (2025)
by: Yuan, Yurun, et al.
Published: (2025)
Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning
by: Wu, Lixin, et al.
Published: (2025)
by: Wu, Lixin, et al.
Published: (2025)
Vidur: A Large-Scale Simulation Framework For LLM Inference
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
Can't Remember Details in Long Documents? You Need Some R&R
by: Agrawal, Devanshu, et al.
Published: (2024)
by: Agrawal, Devanshu, et al.
Published: (2024)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
by: Agrawal, Sudhanshu, et al.
Published: (2025)
by: Agrawal, Sudhanshu, et al.
Published: (2025)
SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection
by: Tang, Kexian, et al.
Published: (2026)
by: Tang, Kexian, et al.
Published: (2026)
A Data-Centric Approach To Generate Faithful and High Quality Patient Summaries with Large Language Models
by: Hegselmann, Stefan, et al.
Published: (2024)
by: Hegselmann, Stefan, et al.
Published: (2024)
No Mean Feat: Simple, Strong Baselines for Context Compression
by: Feldman, Yair, et al.
Published: (2025)
by: Feldman, Yair, et al.
Published: (2025)
Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
by: Yousuf, Raquib Bin, et al.
Published: (2025)
by: Yousuf, Raquib Bin, et al.
Published: (2025)
A Case Study of Selected PTQ Baselines for Reasoning LLMs on Ascend NPU
by: Luo, Yuchen, et al.
Published: (2026)
by: Luo, Yuchen, et al.
Published: (2026)
KL for a KL: On-Policy Distillation with Control Variate Baseline
by: Oh, Minjae, et al.
Published: (2026)
by: Oh, Minjae, et al.
Published: (2026)
NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
by: Ghimire, Rupak Raj, et al.
Published: (2026)
by: Ghimire, Rupak Raj, et al.
Published: (2026)
IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector
by: Chen, Zheng, et al.
Published: (2025)
by: Chen, Zheng, et al.
Published: (2025)
Lightweight reranking for language model generations
by: Jain, Siddhartha, et al.
Published: (2023)
by: Jain, Siddhartha, et al.
Published: (2023)
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
by: Hwang, Jaedong, et al.
Published: (2025)
by: Hwang, Jaedong, et al.
Published: (2025)
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
by: Guo, Siyuan, et al.
Published: (2024)
by: Guo, Siyuan, et al.
Published: (2024)
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
by: Zhang, Yi-Kai, et al.
Published: (2025)
by: Zhang, Yi-Kai, et al.
Published: (2025)
Reinforce LLM Reasoning through Multi-Agent Reflection
by: Yuan, Yurun, et al.
Published: (2025)
by: Yuan, Yurun, et al.
Published: (2025)
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
by: Yan, Kai, et al.
Published: (2025)
by: Yan, Kai, et al.
Published: (2025)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025)
by: Zhou, Zhi, et al.
Published: (2025)
Similar Items
-
When Can Transformers Count to n?
by: Yehudai, Gilad, et al.
Published: (2024) -
Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
by: Agrawal, Vishakha, et al.
Published: (2025) -
Proving that Cryptic Crossword Clue Answers are Correct
by: Andrews, Martin, et al.
Published: (2024) -
Addressing LLM Diversity by Infusing Random Concepts
by: Agrawal, Pulin, et al.
Published: (2026) -
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
by: Ding, Mucong, et al.
Published: (2024)