Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Burleigh, Tyler |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Human-in-the-Loop Benchmarking of Heterogeneous LLMs for Automated Competency Assessment in Secondary Level Mathematics
von: Bhusal, Jatin, et al.
Veröffentlicht: (2026)
von: Bhusal, Jatin, et al.
Veröffentlicht: (2026)
Integrating Generative AI in Cybersecurity Education: Case Study Insights on Pedagogical Strategies, Critical Thinking, and Responsible AI Use
von: Elkhodr, Mahmoud, et al.
Veröffentlicht: (2025)
von: Elkhodr, Mahmoud, et al.
Veröffentlicht: (2025)
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
von: Du, Yishan, et al.
Veröffentlicht: (2025)
von: Du, Yishan, et al.
Veröffentlicht: (2025)
IntelliCode: A Multi-Agent LLM Tutoring System with Centralized Learner Modeling
von: David, Jones, et al.
Veröffentlicht: (2025)
von: David, Jones, et al.
Veröffentlicht: (2025)
Calibrated Confidence Estimation for Tabular Question Answering
von: Voss, Lukas
Veröffentlicht: (2026)
von: Voss, Lukas
Veröffentlicht: (2026)
MKJ at SemEval-2026 Task 9: A Comparative Study of Generalist, Specialist, and Ensemble Strategies for Multilingual Polarization
von: Jouneghani, Maziar Kianimoghadam
Veröffentlicht: (2026)
von: Jouneghani, Maziar Kianimoghadam
Veröffentlicht: (2026)
Statistical Scouting Finds Debate-Safe but Not Debate-Useful Cases: A Matched-Ceiling Study of Open-Weight LLM Reasoning Protocols
von: Hu, Julia, et al.
Veröffentlicht: (2026)
von: Hu, Julia, et al.
Veröffentlicht: (2026)
LLMs as Architects and Critics for Multi-Source Opinion Summarization
von: Attri, Anuj, et al.
Veröffentlicht: (2025)
von: Attri, Anuj, et al.
Veröffentlicht: (2025)
Why We Feel What We Feel: Joint Detection of Emotions and Their Opinion Triggers in E-commerce
von: Attri, Arnav, et al.
Veröffentlicht: (2025)
von: Attri, Arnav, et al.
Veröffentlicht: (2025)
The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks
von: Mullens, Drake, et al.
Veröffentlicht: (2026)
von: Mullens, Drake, et al.
Veröffentlicht: (2026)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
von: Ubukata, Shunsuke
Veröffentlicht: (2026)
von: Ubukata, Shunsuke
Veröffentlicht: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Navigational Thinking as an Emerging Paradigm of Computer Science in the Age of Generative AI
von: Levin, Ilya
Veröffentlicht: (2026)
von: Levin, Ilya
Veröffentlicht: (2026)
From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2026)
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2026)
The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It
von: Garcia, Gabriel
Veröffentlicht: (2026)
von: Garcia, Gabriel
Veröffentlicht: (2026)
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025)
Methodological Foundations for AI-Driven Survey Question Generation
von: Mburu, Ted K., et al.
Veröffentlicht: (2025)
von: Mburu, Ted K., et al.
Veröffentlicht: (2025)
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
von: Zahid, Anwar Hossain, et al.
Veröffentlicht: (2025)
von: Zahid, Anwar Hossain, et al.
Veröffentlicht: (2025)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
von: Han, Xudong, et al.
Veröffentlicht: (2025)
von: Han, Xudong, et al.
Veröffentlicht: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
von: Kugler, Kai
Veröffentlicht: (2025)
von: Kugler, Kai
Veröffentlicht: (2025)
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
von: Zhang, Luyan
Veröffentlicht: (2025)
von: Zhang, Luyan
Veröffentlicht: (2025)
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation
von: Bertina, Abbas, et al.
Veröffentlicht: (2025)
von: Bertina, Abbas, et al.
Veröffentlicht: (2025)
ARCHED: A Human-Centered Framework for Transparent, Responsible, and Collaborative AI-Assisted Instructional Design
von: Li, Hongming, et al.
Veröffentlicht: (2025)
von: Li, Hongming, et al.
Veröffentlicht: (2025)
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English
von: Juzek, Tom S
Veröffentlicht: (2025)
von: Juzek, Tom S
Veröffentlicht: (2025)
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
How Large Language Models Are Changing MOOC Essay Answers: A Comparison of Pre- and Post-LLM Responses
von: Leppänen, Leo, et al.
Veröffentlicht: (2025)
von: Leppänen, Leo, et al.
Veröffentlicht: (2025)
LLMs as Educational Analysts: Transforming Multimodal Data Traces into Actionable Reading Assessment Reports
von: Davalos, Eduardo, et al.
Veröffentlicht: (2025)
von: Davalos, Eduardo, et al.
Veröffentlicht: (2025)
TrueReason: An Exemplar Personalised Learning System Integrating Reasoning with Foundational Models
von: Bulathwela, Sahan, et al.
Veröffentlicht: (2025)
von: Bulathwela, Sahan, et al.
Veröffentlicht: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
von: Gwak, Jiho, et al.
Veröffentlicht: (2025)
von: Gwak, Jiho, et al.
Veröffentlicht: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
On the Influence of Discourse Relations in Persuasive Texts
von: Turk, Nawar, et al.
Veröffentlicht: (2025)
von: Turk, Nawar, et al.
Veröffentlicht: (2025)
Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus
von: Arabov, Mullosharaf K.
Veröffentlicht: (2026)
von: Arabov, Mullosharaf K.
Veröffentlicht: (2026)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
von: Resck, Lucas, et al.
Veröffentlicht: (2026)
von: Resck, Lucas, et al.
Veröffentlicht: (2026)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation
von: Li, Zilin, et al.
Veröffentlicht: (2026)
von: Li, Zilin, et al.
Veröffentlicht: (2026)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
von: Bian, Tingcheng, et al.
Veröffentlicht: (2026)
von: Bian, Tingcheng, et al.
Veröffentlicht: (2026)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
von: Young, Richard J.
Veröffentlicht: (2026)
von: Young, Richard J.
Veröffentlicht: (2026)
The AI Fiction Paradox
von: Elkins, Katherine
Veröffentlicht: (2026)
von: Elkins, Katherine
Veröffentlicht: (2026)
Distilling Self-Consistency into Verbal Confidence: A Pre-Registered Negative Result and Post-Hoc Rescue on Gemma 3 4B
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
Ähnliche Einträge
-
Human-in-the-Loop Benchmarking of Heterogeneous LLMs for Automated Competency Assessment in Secondary Level Mathematics
von: Bhusal, Jatin, et al.
Veröffentlicht: (2026) -
Integrating Generative AI in Cybersecurity Education: Case Study Insights on Pedagogical Strategies, Critical Thinking, and Responsible AI Use
von: Elkhodr, Mahmoud, et al.
Veröffentlicht: (2025) -
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
von: Du, Yishan, et al.
Veröffentlicht: (2025) -
IntelliCode: A Multi-Agent LLM Tutoring System with Centralized Learner Modeling
von: David, Jones, et al.
Veröffentlicht: (2025) -
Calibrated Confidence Estimation for Tabular Question Answering
von: Voss, Lukas
Veröffentlicht: (2026)