Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kale, Sahil, Nadadur, Vijaykant |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
par: Kale, Sahil, et autres
Publié: (2025)
par: Kale, Sahil, et autres
Publié: (2025)
Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
par: Kale, Sahil, et autres
Publié: (2025)
par: Kale, Sahil, et autres
Publié: (2025)
Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs
par: Kale, Sahil
Publié: (2025)
par: Kale, Sahil
Publié: (2025)
KnowRL: Teaching Language Models to Know What They Know
par: Kale, Sahil, et autres
Publié: (2025)
par: Kale, Sahil, et autres
Publié: (2025)
Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
par: Kale, Sahil
Publié: (2025)
par: Kale, Sahil
Publié: (2025)
AXCEL: Automated eXplainable Consistency Evaluation using LLMs
par: Sreekar, P Aditya, et autres
Publié: (2024)
par: Sreekar, P Aditya, et autres
Publié: (2024)
Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling
par: Zheng, Hang, et autres
Publié: (2025)
par: Zheng, Hang, et autres
Publié: (2025)
SelfPrompt: Autonomously Evaluating LLM Robustness via Domain-Constrained Knowledge Guidelines and Refined Adversarial Prompts
par: Pei, Aihua, et autres
Publié: (2024)
par: Pei, Aihua, et autres
Publié: (2024)
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities
par: Hong, Zhaochen, et autres
Publié: (2025)
par: Hong, Zhaochen, et autres
Publié: (2025)
MAC-Tuning: LLM Multi-Compositional Problem Reasoning with Enhanced Knowledge Boundary Awareness
par: Huang, Junsheng, et autres
Publié: (2025)
par: Huang, Junsheng, et autres
Publié: (2025)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
par: Zhang, Kongcheng, et autres
Publié: (2025)
par: Zhang, Kongcheng, et autres
Publié: (2025)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
par: Wan, Guangya, et autres
Publié: (2024)
par: Wan, Guangya, et autres
Publié: (2024)
EMSDialog: Synthetic Multi-person Emergency Medical Service Dialogue Generation from Electronic Patient Care Reports via Multi-LLM Agents
par: Ge, Xueren, et autres
Publié: (2026)
par: Ge, Xueren, et autres
Publié: (2026)
Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models
par: Jia, Boyu, et autres
Publié: (2025)
par: Jia, Boyu, et autres
Publié: (2025)
Micro-Act: Mitigating Knowledge Conflict in LLM-based RAG via Actionable Self-Reasoning
par: Huo, Nan, et autres
Publié: (2025)
par: Huo, Nan, et autres
Publié: (2025)
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
par: Wang, Guanghui, et autres
Publié: (2025)
par: Wang, Guanghui, et autres
Publié: (2025)
Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering
par: Lee, Yanggyu, et autres
Publié: (2024)
par: Lee, Yanggyu, et autres
Publié: (2024)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
par: Sorensen, Taylor, et autres
Publié: (2023)
par: Sorensen, Taylor, et autres
Publié: (2023)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
par: Lee, Jaehyeok, et autres
Publié: (2024)
par: Lee, Jaehyeok, et autres
Publié: (2024)
Efficient Knowledge Infusion via KG-LLM Alignment
par: Jiang, Zhouyu, et autres
Publié: (2024)
par: Jiang, Zhouyu, et autres
Publié: (2024)
Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries
par: Zhang, Yuchen, et autres
Publié: (2026)
par: Zhang, Yuchen, et autres
Publié: (2026)
Confidence Improves Self-Consistency in LLMs
par: Taubenfeld, Amir, et autres
Publié: (2025)
par: Taubenfeld, Amir, et autres
Publié: (2025)
ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM
par: Su, Zhaochen, et autres
Publié: (2024)
par: Su, Zhaochen, et autres
Publié: (2024)
LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble
par: Lee, Yujeong, et autres
Publié: (2024)
par: Lee, Yujeong, et autres
Publié: (2024)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
par: Zhou, Zhi, et autres
Publié: (2025)
par: Zhou, Zhi, et autres
Publié: (2025)
Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency
par: Li, Taiji, et autres
Publié: (2024)
par: Li, Taiji, et autres
Publié: (2024)
Are LLM Belief Updates Consistent with Bayes' Theorem?
par: Imran, Sohaib, et autres
Publié: (2025)
par: Imran, Sohaib, et autres
Publié: (2025)
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
par: Li, Yanhong, et autres
Publié: (2025)
par: Li, Yanhong, et autres
Publié: (2025)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
par: Shi, Zhichao, et autres
Publié: (2025)
par: Shi, Zhichao, et autres
Publié: (2025)
Beyond Self-Consistency: Ensemble Reasoning Boosts Consistency and Accuracy of LLMs in Cancer Staging
par: Chang, Chia-Hsuan, et autres
Publié: (2024)
par: Chang, Chia-Hsuan, et autres
Publié: (2024)
Self-Consistency Boosts Calibration for Math Reasoning
par: Wang, Ante, et autres
Publié: (2024)
par: Wang, Ante, et autres
Publié: (2024)
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
par: Liang, Yu, et autres
Publié: (2026)
par: Liang, Yu, et autres
Publié: (2026)
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
par: Pandey, Atharva, et autres
Publié: (2025)
par: Pandey, Atharva, et autres
Publié: (2025)
Harnessing Consistency for Robust Test-Time LLM Ensemble
par: Zeng, Zhichen, et autres
Publié: (2025)
par: Zeng, Zhichen, et autres
Publié: (2025)
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
par: Chen, Sirui, et autres
Publié: (2026)
par: Chen, Sirui, et autres
Publié: (2026)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
par: Zhang, Xiaoying, et autres
Publié: (2024)
par: Zhang, Xiaoying, et autres
Publié: (2024)
What is in a name? Mitigating Name Bias in Text Embeddings via Anonymization
par: Manchanda, Sahil, et autres
Publié: (2025)
par: Manchanda, Sahil, et autres
Publié: (2025)
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
par: Kumar, Adarsh, et autres
Publié: (2025)
par: Kumar, Adarsh, et autres
Publié: (2025)
Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs
par: Kjorvezir, Denica, et autres
Publié: (2026)
par: Kjorvezir, Denica, et autres
Publié: (2026)
Can We Trust a Black-box LLM? LLM Untrustworthy Boundary Detection via Bias-Diffusion and Multi-Agent Reinforcement Learning
par: Zhou, Xiaotian, et autres
Publié: (2026)
par: Zhou, Xiaotian, et autres
Publié: (2026)
Documents similaires
-
TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
par: Kale, Sahil, et autres
Publié: (2025) -
Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
par: Kale, Sahil, et autres
Publié: (2025) -
Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs
par: Kale, Sahil
Publié: (2025) -
KnowRL: Teaching Language Models to Know What They Know
par: Kale, Sahil, et autres
Publié: (2025) -
Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
par: Kale, Sahil
Publié: (2025)