Estimating problem difficulty without ground truth using Large Language Model comparisons
Fuente:
arXiv
Saved in:
| Main Authors: | Ballon, Marthe, Algaba, Andres, Verbeken, Brecht, Ginis, Vincent |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probing the Trajectories of Reasoning Traces in Large Language Models
by: Ballon, Marthe, et al.
Published: (2026)
by: Ballon, Marthe, et al.
Published: (2026)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
by: Ballon, Marthe, et al.
Published: (2026)
by: Ballon, Marthe, et al.
Published: (2026)
The Relationship Between Reasoning and Performance in Large Language Models -- o3 (mini) Thinks Harder, Not Longer
by: Ballon, Marthe, et al.
Published: (2025)
by: Ballon, Marthe, et al.
Published: (2025)
Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking)
by: Verbeken, Brecht, et al.
Published: (2026)
by: Verbeken, Brecht, et al.
Published: (2026)
How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
by: Algaba, Andres, et al.
Published: (2025)
by: Algaba, Andres, et al.
Published: (2025)
Lexical Hints of Accuracy in LLM Reasoning Chains
by: Vanhoyweghen, Arne, et al.
Published: (2025)
by: Vanhoyweghen, Arne, et al.
Published: (2025)
Flexible Counterfactual Explanations with Generative Models
by: Hellemans, Stig, et al.
Published: (2025)
by: Hellemans, Stig, et al.
Published: (2025)
Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias
by: Algaba, Andres, et al.
Published: (2024)
by: Algaba, Andres, et al.
Published: (2024)
Scalable Classification of Course Information Sheets Using Large Language Models: A Reusable Institutional Method for Academic Quality Assurance
by: Verbeken, Brecht, et al.
Published: (2026)
by: Verbeken, Brecht, et al.
Published: (2026)
SUDO: a framework for evaluating clinical artificial intelligence systems without ground-truth annotations
by: Kiyasseh, Dani, et al.
Published: (2024)
by: Kiyasseh, Dani, et al.
Published: (2024)
Estimating the Effects of Sample Training Orders for Large Language Models without Retraining
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
Probing Graph Neural Network Activation Patterns Through Graph Topology
by: Tori, Floriano, et al.
Published: (2026)
by: Tori, Floriano, et al.
Published: (2026)
Structurally Human, Semantically Biased: Detecting LLM-Generated References with Embeddings and GNNs
by: Mobini, Melika, et al.
Published: (2026)
by: Mobini, Melika, et al.
Published: (2026)
Metro 3 in Brussels under uncertainty: scenario-based public transport accessibility analysis
by: Verbeken, Brecht, et al.
Published: (2025)
by: Verbeken, Brecht, et al.
Published: (2025)
Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments
by: Vanhoyweghen, Arne, et al.
Published: (2026)
by: Vanhoyweghen, Arne, et al.
Published: (2026)
Early evidence of how LLMs outperform traditional systems on OCR/HTR tasks for historical records
by: Kim, Seorin, et al.
Published: (2025)
by: Kim, Seorin, et al.
Published: (2025)
Ranking Large Language Models without Ground Truth
by: Dhurandhar, Amit, et al.
Published: (2024)
by: Dhurandhar, Amit, et al.
Published: (2024)
Steering Large Language Model Activations in Sparse Spaces
by: Bayat, Reza, et al.
Published: (2025)
by: Bayat, Reza, et al.
Published: (2025)
Data Descriptions from Large Language Models with Influence Estimation
by: Kim, Chaeri, et al.
Published: (2025)
by: Kim, Chaeri, et al.
Published: (2025)
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
by: Han, Bing, et al.
Published: (2025)
by: Han, Bing, et al.
Published: (2025)
KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
by: Askari, Hadi, et al.
Published: (2025)
by: Askari, Hadi, et al.
Published: (2025)
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
by: Park, Seonghyeon, et al.
Published: (2026)
by: Park, Seonghyeon, et al.
Published: (2026)
LLM4XCE: Large Language Models for Extremely Large-Scale Massive MIMO Channel Estimation
by: Li, Renbin, et al.
Published: (2025)
by: Li, Renbin, et al.
Published: (2025)
Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models
by: Wu, Jinman, et al.
Published: (2026)
by: Wu, Jinman, et al.
Published: (2026)
You can't handle the (dirty) truth: Data-centric insights improve pseudo-labeling
by: Seedat, Nabeel, et al.
Published: (2024)
by: Seedat, Nabeel, et al.
Published: (2024)
Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
by: Sharma, Mridul, et al.
Published: (2025)
by: Sharma, Mridul, et al.
Published: (2025)
Revisiting Uncertainty Estimation and Calibration of Large Language Models
by: Tao, Linwei, et al.
Published: (2025)
by: Tao, Linwei, et al.
Published: (2025)
Estimating Item Difficulty with Large Language Models as Experts
by: Kolesnikova, Diana, et al.
Published: (2026)
by: Kolesnikova, Diana, et al.
Published: (2026)
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting
by: Nam, Yunhun, et al.
Published: (2025)
by: Nam, Yunhun, et al.
Published: (2025)
Estimating the Empowerment of Language Model Agents
by: Song, Jinyeop, et al.
Published: (2025)
by: Song, Jinyeop, et al.
Published: (2025)
Estimating the Probabilities of Rare Outputs in Language Models
by: Wu, Gabriel, et al.
Published: (2024)
by: Wu, Gabriel, et al.
Published: (2024)
Generating Multiple-Choice Knowledge Questions with Interpretable Difficulty Estimation using Knowledge Graphs and Large Language Models
by: Şakiroğlu, Mehmet Can, et al.
Published: (2026)
by: Şakiroğlu, Mehmet Can, et al.
Published: (2026)
Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility
by: Wang, Mengxuan, et al.
Published: (2026)
by: Wang, Mengxuan, et al.
Published: (2026)
Large Language Model Confidence Estimation via Black-Box Access
by: Pedapati, Tejaswini, et al.
Published: (2024)
by: Pedapati, Tejaswini, et al.
Published: (2024)
Adaptable Cardiovascular Disease Risk Prediction from Heterogeneous Data using Large Language Models
by: Lübeck, Frederike, et al.
Published: (2025)
by: Lübeck, Frederike, et al.
Published: (2025)
SeRA: Self-Reviewing and Alignment of Large Language Models using Implicit Reward Margins
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
Estimating Tail Risks in Language Model Output Distributions
by: Angell, Rico, et al.
Published: (2026)
by: Angell, Rico, et al.
Published: (2026)
Learning Evolving Latent Strategies for Multi-Agent Language Systems without Model Fine-Tuning
by: Tang, Wenlong
Published: (2025)
by: Tang, Wenlong
Published: (2025)
Similar Items
-
Probing the Trajectories of Reasoning Traces in Large Language Models
by: Ballon, Marthe, et al.
Published: (2026) -
Benchmarks Saturate When The Model Gets Smarter Than The Judge
by: Ballon, Marthe, et al.
Published: (2026) -
The Relationship Between Reasoning and Performance in Large Language Models -- o3 (mini) Thinks Harder, Not Longer
by: Ballon, Marthe, et al.
Published: (2025) -
Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking)
by: Verbeken, Brecht, et al.
Published: (2026) -
How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
by: Algaba, Andres, et al.
Published: (2025)