Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiang, Xin, Jiayi, Long, Qi, Su, Weijie J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Errors in AI-Assisted Retrieval of Medical Literature: A Comparative Study
by: Gao, Jenny, et al.
Published: (2026)
by: Gao, Jenny, et al.
Published: (2026)
dsld: A Socially Relevant Tool for Teaching Statistics
by: Mittal, Aditya, et al.
Published: (2024)
by: Mittal, Aditya, et al.
Published: (2024)
Do Large Language Models (Really) Need Statistical Foundations?
by: Su, Weijie
Published: (2025)
by: Su, Weijie
Published: (2025)
How many patients could we save with LLM priors?
by: Arai, Shota, et al.
Published: (2025)
by: Arai, Shota, et al.
Published: (2025)
Variance Reduction in Ratio Metrics for Efficient Online Experiments
by: Baweja, Shubham, et al.
Published: (2024)
by: Baweja, Shubham, et al.
Published: (2024)
Beyond Beats: A Recipe to Song Popularity? A machine learning approach
by: Sebastian, Niklas, et al.
Published: (2024)
by: Sebastian, Niklas, et al.
Published: (2024)
Learning Metrics that Maximise Power for Accelerated A/B-Tests
by: Jeunen, Olivier, et al.
Published: (2024)
by: Jeunen, Olivier, et al.
Published: (2024)
Multi-View Variational Autoencoder for Missing Value Imputation in Untargeted Metabolomics
by: Zhao, Chen, et al.
Published: (2023)
by: Zhao, Chen, et al.
Published: (2023)
NanoKnow: How to Know What Your Language Model Knows
by: Gu, Lingwei, et al.
Published: (2026)
by: Gu, Lingwei, et al.
Published: (2026)
Tackling Copyright Issues in AI Image Generation Through Originality Estimation and Genericization
by: Chiba-Okabe, Hiroaki, et al.
Published: (2024)
by: Chiba-Okabe, Hiroaki, et al.
Published: (2024)
A multi-language toolkit for the semi-automated checking of research outputs
by: Preen, Richard J., et al.
Published: (2022)
by: Preen, Richard J., et al.
Published: (2022)
Meta Off-Policy Estimation
by: Jeunen, Olivier
Published: (2025)
by: Jeunen, Olivier
Published: (2025)
Two-stage Risk Control with Application to Ranked Retrieval
by: Xu, Yunpeng, et al.
Published: (2024)
by: Xu, Yunpeng, et al.
Published: (2024)
Counterfactual Inference under Thompson Sampling
by: Jeunen, Olivier
Published: (2025)
by: Jeunen, Olivier
Published: (2025)
Unifying On- and Off-Policy Variance Reduction Methods
by: Jeunen, Olivier
Published: (2026)
by: Jeunen, Olivier
Published: (2026)
Evaluation of Missing Data Analytical Techniques in Longitudinal Research: Traditional and Machine Learning Approaches
by: Tang, Dandan, et al.
Published: (2024)
by: Tang, Dandan, et al.
Published: (2024)
A Simple Model to Estimate Sharing Effects in Social Networks
by: Jeunen, Olivier
Published: (2024)
by: Jeunen, Olivier
Published: (2024)
Antibiotic Resistance Microbiology Dataset (ARMD): A Resource for Antimicrobial Resistance from EHRs
by: Haredasht, Fateme Nateghi, et al.
Published: (2025)
by: Haredasht, Fateme Nateghi, et al.
Published: (2025)
LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
by: Zhang, Yunfan, et al.
Published: (2026)
by: Zhang, Yunfan, et al.
Published: (2026)
A Design-based Solution for Causal Inference with Text: Can a Language Model Be Too Large?
by: Tierney, Graham, et al.
Published: (2025)
by: Tierney, Graham, et al.
Published: (2025)
Search Arena: Analyzing Search-Augmented LLMs
by: Miroyan, Mihran, et al.
Published: (2025)
by: Miroyan, Mihran, et al.
Published: (2025)
Ranking Policy Learning via Marketplace Expected Value Estimation From Observational Data
by: Ebrahimzadeh, Ehsan, et al.
Published: (2024)
by: Ebrahimzadeh, Ehsan, et al.
Published: (2024)
Lightweight Adaptation for LLM-based Technical Service Agent: Latent Logic Augmentation and Robust Noise Reduction
by: Yu, Yi, et al.
Published: (2026)
by: Yu, Yi, et al.
Published: (2026)
Dynamic Topic Analysis in Academic Journals using Convex Non-negative Matrix Factorization Method
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
Had enough of experts? Quantitative knowledge retrieval from large language models
by: Selby, David, et al.
Published: (2024)
by: Selby, David, et al.
Published: (2024)
Efficient Inference for Noisy LLM-as-a-Judge Evaluation
by: Chen, Yiqun T, et al.
Published: (2026)
by: Chen, Yiqun T, et al.
Published: (2026)
Optimal Estimation of Watermark Proportions in Hybrid AI-Human Texts
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Spectral goodness-of-fit tests for complete and partial network data
by: Lubold, Shane, et al.
Published: (2021)
by: Lubold, Shane, et al.
Published: (2021)
FlashEvaluator: Expanding Search Space with Parallel Evaluation
by: Feng, Chao, et al.
Published: (2026)
by: Feng, Chao, et al.
Published: (2026)
Causal Structure Representation Learning of Confounders in Latent Space for Recommendation
by: Xu, Hangtong, et al.
Published: (2023)
by: Xu, Hangtong, et al.
Published: (2023)
Computational-Assisted Systematic Review and Meta-Analysis (CASMA): Effect of a Subclass of GnRH-a on Endometriosis Recurrence
by: Tsang, Sandro
Published: (2025)
by: Tsang, Sandro
Published: (2025)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
GACE: Learning Graph-Based Cross-Page Ads Embedding For Click-Through Rate Prediction
by: Wang, Haowen, et al.
Published: (2024)
by: Wang, Haowen, et al.
Published: (2024)
When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes
by: Yehudai, Asaf, et al.
Published: (2024)
by: Yehudai, Asaf, et al.
Published: (2024)
ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging
by: Verma, Neha, et al.
Published: (2026)
by: Verma, Neha, et al.
Published: (2026)
Do LLMs Understand Collaborative Signals? Diagnosis and Repair
by: Pouryousef, Shahrooz, et al.
Published: (2025)
by: Pouryousef, Shahrooz, et al.
Published: (2025)
Logistic regression models for patient-level prediction based on massive observational data: Do we need all data?
by: John, Luis H., et al.
Published: (2020)
by: John, Luis H., et al.
Published: (2020)
Not Just How Much, But Where: Decomposing Epistemic Uncertainty into Per-Class Contributions
by: Toure, Mame Diarra, et al.
Published: (2026)
by: Toure, Mame Diarra, et al.
Published: (2026)
ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities
by: Xu, Peng, et al.
Published: (2024)
by: Xu, Peng, et al.
Published: (2024)
Robust Detection of Watermarks for Large Language Models Under Human Edits
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Similar Items
-
Errors in AI-Assisted Retrieval of Medical Literature: A Comparative Study
by: Gao, Jenny, et al.
Published: (2026) -
dsld: A Socially Relevant Tool for Teaching Statistics
by: Mittal, Aditya, et al.
Published: (2024) -
Do Large Language Models (Really) Need Statistical Foundations?
by: Su, Weijie
Published: (2025) -
How many patients could we save with LLM priors?
by: Arai, Shota, et al.
Published: (2025) -
Variance Reduction in Ratio Metrics for Efficient Online Experiments
by: Baweja, Shubham, et al.
Published: (2024)