Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Changyue, Su, Weihang, Ai, Qingyao, Liu, Yiqun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improve Large Language Model Systems with User Logs
von: Wang, Changyue, et al.
Veröffentlicht: (2026)
von: Wang, Changyue, et al.
Veröffentlicht: (2026)
Decoupling Reasoning and Knowledge Injection for In-Context Knowledge Editing
von: Wang, Changyue, et al.
Veröffentlicht: (2025)
von: Wang, Changyue, et al.
Veröffentlicht: (2025)
Mitigating Entity-Level Hallucination in Large Language Models
von: Su, Weihang, et al.
Veröffentlicht: (2024)
von: Su, Weihang, et al.
Veröffentlicht: (2024)
Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
von: Su, Weihang, et al.
Veröffentlicht: (2024)
von: Su, Weihang, et al.
Veröffentlicht: (2024)
Towards Unification of Hallucination Detection and Fact Verification for Large Language Models
von: Su, Weihang, et al.
Veröffentlicht: (2025)
von: Su, Weihang, et al.
Veröffentlicht: (2025)
Knowledge Editing through Chain-of-Thought
von: Wang, Changyue, et al.
Veröffentlicht: (2024)
von: Wang, Changyue, et al.
Veröffentlicht: (2024)
Decoupling Knowledge and Task Subspaces for Composable Parametric Retrieval Augmented Generation
von: Su, Weihang, et al.
Veröffentlicht: (2026)
von: Su, Weihang, et al.
Veröffentlicht: (2026)
DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models
von: Su, Weihang, et al.
Veröffentlicht: (2024)
von: Su, Weihang, et al.
Veröffentlicht: (2024)
Multi-Field Tool Retrieval
von: Tang, Yichen, et al.
Veröffentlicht: (2026)
von: Tang, Yichen, et al.
Veröffentlicht: (2026)
RbFT: Robust Fine-tuning for Retrieval-Augmented Generation against Retrieval Defects
von: Tu, Yiteng, et al.
Veröffentlicht: (2025)
von: Tu, Yiteng, et al.
Veröffentlicht: (2025)
Dynamic and Parametric Retrieval-Augmented Generation
von: Su, Weihang, et al.
Veröffentlicht: (2025)
von: Su, Weihang, et al.
Veröffentlicht: (2025)
Parametric Retrieval Augmented Generation
von: Su, Weihang, et al.
Veröffentlicht: (2025)
von: Su, Weihang, et al.
Veröffentlicht: (2025)
LeKUBE: A Legal Knowledge Update BEnchmark
von: Wang, Changyue, et al.
Veröffentlicht: (2024)
von: Wang, Changyue, et al.
Veröffentlicht: (2024)
Enhancing Judgment Document Generation via Agentic Legal Information Collection and Rubric-Guided Optimization
von: Su, Weihang, et al.
Veröffentlicht: (2026)
von: Su, Weihang, et al.
Veröffentlicht: (2026)
Analytical Search
von: Tu, Yiteng, et al.
Veröffentlicht: (2026)
von: Tu, Yiteng, et al.
Veröffentlicht: (2026)
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System
von: Su, Weihang, et al.
Veröffentlicht: (2025)
von: Su, Weihang, et al.
Veröffentlicht: (2025)
PRE: A Peer Review Based Large Language Model Evaluator
von: Chu, Zhumin, et al.
Veröffentlicht: (2024)
von: Chu, Zhumin, et al.
Veröffentlicht: (2024)
Augmenting Multi-Agent Communication with State Delta Trajectory
von: Tang, Yichen, et al.
Veröffentlicht: (2025)
von: Tang, Yichen, et al.
Veröffentlicht: (2025)
LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models
von: Li, Haitao, et al.
Veröffentlicht: (2024)
von: Li, Haitao, et al.
Veröffentlicht: (2024)
Scaling Laws For Dense Retrieval
von: Fang, Yan, et al.
Veröffentlicht: (2024)
von: Fang, Yan, et al.
Veröffentlicht: (2024)
SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
von: Su, Weihang, et al.
Veröffentlicht: (2025)
von: Su, Weihang, et al.
Veröffentlicht: (2025)
Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
von: Song, Yusheng, et al.
Veröffentlicht: (2025)
von: Song, Yusheng, et al.
Veröffentlicht: (2025)
From <Answer> to <Think>: Multidimensional Supervision of Reasoning Process for LLM Optimization
von: Wang, Beining, et al.
Veröffentlicht: (2025)
von: Wang, Beining, et al.
Veröffentlicht: (2025)
Overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
von: Ai, Qingyao, et al.
Veröffentlicht: (2025)
von: Ai, Qingyao, et al.
Veröffentlicht: (2025)
TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
von: Wang, Yuhui, et al.
Veröffentlicht: (2025)
von: Wang, Yuhui, et al.
Veröffentlicht: (2025)
How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study
von: Zhu, Shuqi, et al.
Veröffentlicht: (2026)
von: Zhu, Shuqi, et al.
Veröffentlicht: (2026)
STARD: A Chinese Statute Retrieval Dataset with Real Queries Issued by Non-professionals
von: Su, Weihang, et al.
Veröffentlicht: (2024)
von: Su, Weihang, et al.
Veröffentlicht: (2024)
Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation Evaluation
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2025)
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2025)
xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
von: Chen, Ding, et al.
Veröffentlicht: (2025)
von: Chen, Ding, et al.
Veröffentlicht: (2025)
Option-ID Based Elimination For Multiple Choice Questions
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025)
von: Zhu, Zhenhao, et al.
Veröffentlicht: (2025)
Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models
von: Li, Haitao, et al.
Veröffentlicht: (2024)
von: Li, Haitao, et al.
Veröffentlicht: (2024)
Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model Capability
von: Wang, Haotian, et al.
Veröffentlicht: (2025)
von: Wang, Haotian, et al.
Veröffentlicht: (2025)
Evaluating Consistency and Reasoning Capabilities of Large Language Models
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
Mitigating Prompt-Induced Hallucinations in Large Language Models via Structured Reasoning
von: Hao, Jinbo, et al.
Veröffentlicht: (2026)
von: Hao, Jinbo, et al.
Veröffentlicht: (2026)
Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2025)
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Improve Large Language Model Systems with User Logs
von: Wang, Changyue, et al.
Veröffentlicht: (2026) -
Decoupling Reasoning and Knowledge Injection for In-Context Knowledge Editing
von: Wang, Changyue, et al.
Veröffentlicht: (2025) -
Mitigating Entity-Level Hallucination in Large Language Models
von: Su, Weihang, et al.
Veröffentlicht: (2024) -
Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
von: Su, Weihang, et al.
Veröffentlicht: (2024) -
Towards Unification of Hallucination Detection and Fact Verification for Large Language Models
von: Su, Weihang, et al.
Veröffentlicht: (2025)