Certainty robustness: Evaluating LLM stability under self-challenging prompts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saadat, Mohammadreza, Nemzer, Steve |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
Fact-Checking with Large Language Models via Probabilistic Certainty and Consistency
von: Wang, Haoran, et al.
Veröffentlicht: (2026)
von: Wang, Haoran, et al.
Veröffentlicht: (2026)
Sequence-Level Certainty Reduces Hallucination In Knowledge-Grounded Dialogue Generation
von: Wan, Yixin, et al.
Veröffentlicht: (2023)
von: Wan, Yixin, et al.
Veröffentlicht: (2023)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
Auto prompting without training labels: An LLM cascade for product quality assessment in e-commerce catalogs
von: Satyadharma, Soham, et al.
Veröffentlicht: (2025)
von: Satyadharma, Soham, et al.
Veröffentlicht: (2025)
DayDreamer at CQs-Gen 2025: Generating Critical Questions through Argument Scheme Completion
von: Zhou, Wendi, et al.
Veröffentlicht: (2025)
von: Zhou, Wendi, et al.
Veröffentlicht: (2025)
Certainty-Guided Reasoning in Large Language Models: A Dynamic Thinking Budget Approach
von: Nogueira, João Paulo, et al.
Veröffentlicht: (2025)
von: Nogueira, João Paulo, et al.
Veröffentlicht: (2025)
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
von: Chadimová, Milena, et al.
Veröffentlicht: (2024)
von: Chadimová, Milena, et al.
Veröffentlicht: (2024)
Soft-prompt Tuning for Large Language Models to Evaluate Bias
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023)
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023)
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
von: Chaudhary, Manav, et al.
Veröffentlicht: (2024)
von: Chaudhary, Manav, et al.
Veröffentlicht: (2024)
Knowledge prompt chaining for semantic modeling
von: Ding, Ning Pei, et al.
Veröffentlicht: (2025)
von: Ding, Ning Pei, et al.
Veröffentlicht: (2025)
StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
von: Rozanov, Nikolai, et al.
Veröffentlicht: (2024)
von: Rozanov, Nikolai, et al.
Veröffentlicht: (2024)
MOSLIM:Align with diverse preferences in prompts through reward classification
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction
von: Pan, Ruihao, et al.
Veröffentlicht: (2026)
von: Pan, Ruihao, et al.
Veröffentlicht: (2026)
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
von: Yamamura, Atsushi, et al.
Veröffentlicht: (2024)
von: Yamamura, Atsushi, et al.
Veröffentlicht: (2024)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
von: Zeng, Jiayi, et al.
Veröffentlicht: (2025)
von: Zeng, Jiayi, et al.
Veröffentlicht: (2025)
Efficient multi-prompt evaluation of LLMs
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
von: Huang, Jiameng, et al.
Veröffentlicht: (2025)
von: Huang, Jiameng, et al.
Veröffentlicht: (2025)
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach
von: Fan, Sinan, et al.
Veröffentlicht: (2025)
von: Fan, Sinan, et al.
Veröffentlicht: (2025)
Trusting CHATGPT: how minor tweaks in the prompts lead to major differences in sentiment classification
von: Cuellar, Jaime E., et al.
Veröffentlicht: (2025)
von: Cuellar, Jaime E., et al.
Veröffentlicht: (2025)
RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning
von: Morandi, Andrea
Veröffentlicht: (2026)
von: Morandi, Andrea
Veröffentlicht: (2026)
Zero-shot prompt-based classification: topic labeling in times of foundation models in German Tweets
von: Münker, Simon, et al.
Veröffentlicht: (2024)
von: Münker, Simon, et al.
Veröffentlicht: (2024)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
von: Han, Steve, et al.
Veröffentlicht: (2025)
von: Han, Steve, et al.
Veröffentlicht: (2025)
Language hooks: a modular framework for augmenting LLM reasoning that decouples tool usage from the model and its prompt
von: de Mijolla, Damien, et al.
Veröffentlicht: (2024)
von: de Mijolla, Damien, et al.
Veröffentlicht: (2024)
RETUYT-INCO at BEA 2026 Shared Task 2: Meta-prompting in Rubric-based Scoring for German
von: Sastre, Ignacio, et al.
Veröffentlicht: (2026)
von: Sastre, Ignacio, et al.
Veröffentlicht: (2026)
Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning
von: Fuoli, Matteo, et al.
Veröffentlicht: (2025)
von: Fuoli, Matteo, et al.
Veröffentlicht: (2025)
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
von: Yuan, Dong, et al.
Veröffentlicht: (2024)
von: Yuan, Dong, et al.
Veröffentlicht: (2024)
Do LLMs Align with My Task? Evaluating Text-to-SQL via Dataset Alignment
von: Rafiei, Davood, et al.
Veröffentlicht: (2025)
von: Rafiei, Davood, et al.
Veröffentlicht: (2025)
LLM Prompt Evaluation for Educational Applications
von: Holmes, Langdon, et al.
Veröffentlicht: (2026)
von: Holmes, Langdon, et al.
Veröffentlicht: (2026)
Evil twins are not that evil: Qualitative insights into machine-generated prompts
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
DaVinci at SemEval-2024 Task 9: Few-shot prompting GPT-3.5 for Unconventional Reasoning
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024)
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024)
The Challenges of Evaluating LLM Applications: An Analysis of Automated, Human, and LLM-Based Approaches
von: Abeysinghe, Bhashithe, et al.
Veröffentlicht: (2024)
von: Abeysinghe, Bhashithe, et al.
Veröffentlicht: (2024)
Retrieval augmented generation based dynamic prompting for few-shot biomedical named entity recognition using large language models
von: Ge, Yao, et al.
Veröffentlicht: (2025)
von: Ge, Yao, et al.
Veröffentlicht: (2025)
Autorubric: Unifying Rubric-based LLM Evaluation
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
Rethinking Human Preference Evaluation of LLM Rationales
von: Li, Ziang, et al.
Veröffentlicht: (2025)
von: Li, Ziang, et al.
Veröffentlicht: (2025)
Integrated Framework for LLM Evaluation with Answer Generation
von: Lee, Sujeong, et al.
Veröffentlicht: (2025)
von: Lee, Sujeong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
von: Lu, Jinghui, et al.
Veröffentlicht: (2025) -
Fact-Checking with Large Language Models via Probabilistic Certainty and Consistency
von: Wang, Haoran, et al.
Veröffentlicht: (2026) -
Sequence-Level Certainty Reduces Hallucination In Knowledge-Grounded Dialogue Generation
von: Wan, Yixin, et al.
Veröffentlicht: (2023) -
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
von: Dong, Harry, et al.
Veröffentlicht: (2024) -
Auto prompting without training labels: An LLM cascade for product quality assessment in e-commerce catalogs
von: Satyadharma, Soham, et al.
Veröffentlicht: (2025)