Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
Fuente:
arXiv
Saved in:
| Main Author: | Cacioli, Jon-Paul |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
LLMs as Signal Detectors: Sensitivity, Bias, and the Temperature-Criterion Analogy
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Distilling Self-Consistency into Verbal Confidence: A Pre-Registered Negative Result and Post-Hoc Rescue on Gemma 3 4B
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Exemplar Retrieval Without Overhypothesis Induction: Limits of Distributional Sequence Learning in Early Word Learning
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Repetition Without Exclusivity: Scale Sensitivity of Referential Mechanisms in Child-Scale Language Models
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
by: Michail, Andrianos, et al.
Published: (2024)
by: Michail, Andrianos, et al.
Published: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
by: Sandan, Isik Baran, et al.
Published: (2025)
by: Sandan, Isik Baran, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
On the Effectiveness of LLM-Specific Fine-Tuning for Detecting AI-Generated Text
by: Gromadzki, Michał, et al.
Published: (2026)
by: Gromadzki, Michał, et al.
Published: (2026)
Spotlights and Blindspots: Evaluating Machine-Generated Text Detection
by: Stowe, Kevin, et al.
Published: (2026)
by: Stowe, Kevin, et al.
Published: (2026)
MedHal: An Evaluation Dataset for Medical Hallucination Detection
by: Mehenni, Gaya, et al.
Published: (2025)
by: Mehenni, Gaya, et al.
Published: (2025)
Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2026)
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2026)
Evaluating Class Membership Relations in Knowledge Graphs using Large Language Models
by: Allen, Bradley P., et al.
Published: (2024)
by: Allen, Bradley P., et al.
Published: (2024)
Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models
by: Wu, Canhui, et al.
Published: (2025)
by: Wu, Canhui, et al.
Published: (2025)
Beyond Prefixes: Graph-as-Memory Cross-Attention for Knowledge Graph Completion with Large Language Models
by: Liu, Ruitong, et al.
Published: (2025)
by: Liu, Ruitong, et al.
Published: (2025)
ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation
by: Pozzi, Riccardo, et al.
Published: (2025)
by: Pozzi, Riccardo, et al.
Published: (2025)
Persuasiveness and Bias in LLM: Investigating the Impact of Persuasiveness and Reinforcement of Bias in Language Models
by: Roy, Saumya
Published: (2025)
by: Roy, Saumya
Published: (2025)
Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
by: Salla, Rohit Kumar, et al.
Published: (2025)
by: Salla, Rohit Kumar, et al.
Published: (2025)
Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models
by: Darm, Paul, et al.
Published: (2025)
by: Darm, Paul, et al.
Published: (2025)
Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models
by: Hawkins, John
Published: (2025)
by: Hawkins, John
Published: (2025)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
University of Indonesia at SemEval-2025 Task 11: Evaluating State-of-the-Art Encoders for Multi-Label Emotion Detection
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025)
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025)
Inference to the Best Explanation in Large Language Models
by: Dalal, Dhairya, et al.
Published: (2024)
by: Dalal, Dhairya, et al.
Published: (2024)
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
TALE: A Tool-Augmented Framework for Reference-Free Evaluation of Large Language Models
by: Badshah, Sher, et al.
Published: (2025)
by: Badshah, Sher, et al.
Published: (2025)
Detecting Subtle Differences between Human and Model Languages Using Spectrum of Relative Likelihood
by: Xu, Yang, et al.
Published: (2024)
by: Xu, Yang, et al.
Published: (2024)
Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks
by: Rair, Nisrine, et al.
Published: (2026)
by: Rair, Nisrine, et al.
Published: (2026)
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models
by: Benkirane, Kenza, et al.
Published: (2024)
by: Benkirane, Kenza, et al.
Published: (2024)
Bielik 7B v0.1: A Polish Language Model -- Development, Insights, and Evaluation
by: Ociepa, Krzysztof, et al.
Published: (2024)
by: Ociepa, Krzysztof, et al.
Published: (2024)
Enhancing Sentiment Classification and Irony Detection in Large Language Models through Advanced Prompt Engineering Techniques
by: Schmitt, Marvin, et al.
Published: (2026)
by: Schmitt, Marvin, et al.
Published: (2026)
Is Our Chatbot Telling Lies? Assessing Correctness of an LLM-based Dutch Support Chatbot
by: Lassche, Herman, et al.
Published: (2024)
by: Lassche, Herman, et al.
Published: (2024)
Analyzing LLM Reasoning to Uncover Mental Health Stigma
by: Sankar, Sreehari, et al.
Published: (2026)
by: Sankar, Sreehari, et al.
Published: (2026)
Similar Items
-
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
by: Cacioli, Jon-Paul
Published: (2026) -
Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models
by: Cacioli, Jon-Paul
Published: (2026) -
LLMs as Signal Detectors: Sensitivity, Bias, and the Temperature-Criterion Analogy
by: Cacioli, Jon-Paul
Published: (2026) -
Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance
by: Cacioli, Jon-Paul
Published: (2026) -
Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging
by: Cacioli, Jon-Paul
Published: (2026)