Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer
Fuente:
arXiv
Guardado en:
| Autores principales: | Yeom, Jewon, Sok, Jaewon, Kim, Heejun, Park, Seonghyeon, Park, Jeongjae, Kim, Taesup |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs
por: Yeom, Jewon, et al.
Publicado: (2026)
por: Yeom, Jewon, et al.
Publicado: (2026)
Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning
por: Sok, Jaewon, et al.
Publicado: (2026)
por: Sok, Jaewon, et al.
Publicado: (2026)
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
por: Park, Seonghyeon, et al.
Publicado: (2026)
por: Park, Seonghyeon, et al.
Publicado: (2026)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
por: Kim, Heejun, et al.
Publicado: (2026)
por: Kim, Heejun, et al.
Publicado: (2026)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
por: Simhi, Adi, et al.
Publicado: (2025)
por: Simhi, Adi, et al.
Publicado: (2025)
Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA
por: Chae, Kyubyung, et al.
Publicado: (2026)
por: Chae, Kyubyung, et al.
Publicado: (2026)
"Well, Keep Thinking": Enhancing LLM Reasoning with Adaptive Injection Decoding
por: Jin, Hyunbin, et al.
Publicado: (2025)
por: Jin, Hyunbin, et al.
Publicado: (2025)
PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation
por: Park, Junho, et al.
Publicado: (2026)
por: Park, Junho, et al.
Publicado: (2026)
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
por: Chae, Kyubyung, et al.
Publicado: (2024)
por: Chae, Kyubyung, et al.
Publicado: (2024)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
por: Kim, Wonjoong, et al.
Publicado: (2025)
por: Kim, Wonjoong, et al.
Publicado: (2025)
Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
por: Chae, Kyubyung, et al.
Publicado: (2025)
por: Chae, Kyubyung, et al.
Publicado: (2025)
ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge Graphs
por: Park, Minbae, et al.
Publicado: (2025)
por: Park, Minbae, et al.
Publicado: (2025)
LLM-guided Plan and Retrieval: A Strategic Alignment for Interpretable User Satisfaction Estimation in Dialogue
por: Kim, Sangyeop, et al.
Publicado: (2025)
por: Kim, Sangyeop, et al.
Publicado: (2025)
Safe-Embed: Unveiling the Safety-Critical Knowledge of Sentence Encoders
por: Kim, Jinseok, et al.
Publicado: (2024)
por: Kim, Jinseok, et al.
Publicado: (2024)
Stable On-Policy Distillation through Adaptive Target Reformulation
por: Jang, Ijun, et al.
Publicado: (2026)
por: Jang, Ijun, et al.
Publicado: (2026)
What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
por: Jang, Jinkwan, et al.
Publicado: (2026)
por: Jang, Jinkwan, et al.
Publicado: (2026)
MetaRAG: Metamorphic Testing for Hallucination Detection in RAG Systems
por: Sok, Channdeth, et al.
Publicado: (2025)
por: Sok, Channdeth, et al.
Publicado: (2025)
Breaking the Pre-Sampling Barrier: Activation-Informed Difficulty-Aware Self-Consistency
por: Yoon, Taewoong, et al.
Publicado: (2026)
por: Yoon, Taewoong, et al.
Publicado: (2026)
X-PEFT: eXtremely Parameter-Efficient Fine-Tuning for Extreme Multi-Profile Scenarios
por: Kwak, Namju, et al.
Publicado: (2024)
por: Kwak, Namju, et al.
Publicado: (2024)
Enhancing Hallucination Detection via Future Context
por: Lee, Joosung, et al.
Publicado: (2025)
por: Lee, Joosung, et al.
Publicado: (2025)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
por: Lee, Joosung, et al.
Publicado: (2026)
por: Lee, Joosung, et al.
Publicado: (2026)
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
por: Kwon, Omin, et al.
Publicado: (2026)
por: Kwon, Omin, et al.
Publicado: (2026)
Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models
por: Madhusudhan, Nishanth, et al.
Publicado: (2024)
por: Madhusudhan, Nishanth, et al.
Publicado: (2024)
DefAn: Definitive Answer Dataset for LLMs Hallucination Evaluation
por: Rahman, A B M Ashikur, et al.
Publicado: (2024)
por: Rahman, A B M Ashikur, et al.
Publicado: (2024)
Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States
por: Duan, Hanyu, et al.
Publicado: (2024)
por: Duan, Hanyu, et al.
Publicado: (2024)
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
por: Kim, Jinpyo, et al.
Publicado: (2025)
por: Kim, Jinpyo, et al.
Publicado: (2025)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
por: Park, Chanwoo, et al.
Publicado: (2025)
por: Park, Chanwoo, et al.
Publicado: (2025)
Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
por: Tan, Chenchen, et al.
Publicado: (2025)
por: Tan, Chenchen, et al.
Publicado: (2025)
Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift
por: Kim, Gihoon, et al.
Publicado: (2025)
por: Kim, Gihoon, et al.
Publicado: (2025)
Steering LLMs toward Korean Local Speech: Iterative Refinement Framework for Faithful Dialect Translation
por: Park, Keunhyeung, et al.
Publicado: (2025)
por: Park, Keunhyeung, et al.
Publicado: (2025)
Robust Domain Generalization under Divergent Marginal and Conditional Distributions
por: Yeom, Jewon, et al.
Publicado: (2026)
por: Yeom, Jewon, et al.
Publicado: (2026)
SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMs
por: Kim, Jaehyung, et al.
Publicado: (2024)
por: Kim, Jaehyung, et al.
Publicado: (2024)
Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US
por: Park, Seongchan, et al.
Publicado: (2026)
por: Park, Seongchan, et al.
Publicado: (2026)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
por: Chae, Kyubyung, et al.
Publicado: (2025)
por: Chae, Kyubyung, et al.
Publicado: (2025)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
por: Kim, Eunsu, et al.
Publicado: (2025)
por: Kim, Eunsu, et al.
Publicado: (2025)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
por: Hwang, Hyeonbin, et al.
Publicado: (2024)
por: Hwang, Hyeonbin, et al.
Publicado: (2024)
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
por: Lee, Heejun, et al.
Publicado: (2025)
por: Lee, Heejun, et al.
Publicado: (2025)
Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
por: Kim, Sewon, et al.
Publicado: (2025)
por: Kim, Sewon, et al.
Publicado: (2025)
M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs
por: Ha, Junwoo, et al.
Publicado: (2025)
por: Ha, Junwoo, et al.
Publicado: (2025)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
por: Kim, Minsang, et al.
Publicado: (2024)
por: Kim, Minsang, et al.
Publicado: (2024)
Ejemplares similares
-
EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs
por: Yeom, Jewon, et al.
Publicado: (2026) -
Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning
por: Sok, Jaewon, et al.
Publicado: (2026) -
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
por: Park, Seonghyeon, et al.
Publicado: (2026) -
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
por: Kim, Heejun, et al.
Publicado: (2026) -
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
por: Simhi, Adi, et al.
Publicado: (2025)