DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs
Fuente:
arXiv
Saved in:
| Main Author: | Huang, Minghui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
by: Wang, Guanghui, et al.
Published: (2025)
by: Wang, Guanghui, et al.
Published: (2025)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
by: Ni, Jingwei, et al.
Published: (2024)
by: Ni, Jingwei, et al.
Published: (2024)
AlignCheck: a Semantic Open-Domain Metric for Factual Consistency Assessment
by: Aghaebrahimian, Ahmad
Published: (2025)
by: Aghaebrahimian, Ahmad
Published: (2025)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
Atomic-SNLI: Fine-Grained Natural Language Inference through Atomic Fact Decomposition
by: Huang, Minghui
Published: (2026)
by: Huang, Minghui
Published: (2026)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
PatentScore: Multi-dimensional Evaluation of LLM-Generated Patent Claims
by: Yoo, Yongmin, et al.
Published: (2025)
by: Yoo, Yongmin, et al.
Published: (2025)
Optimizing Decomposition for Optimal Claim Verification
by: Lu, Yining, et al.
Published: (2025)
by: Lu, Yining, et al.
Published: (2025)
Improving Factual Consistency of News Summarization by Contrastive Preference Optimization
by: Feng, Huawen, et al.
Published: (2023)
by: Feng, Huawen, et al.
Published: (2023)
Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
by: Ye, Yuxuan, et al.
Published: (2026)
by: Ye, Yuxuan, et al.
Published: (2026)
The Alignment Bottleneck in Decomposition-Based Claim Verification
by: Akhter, Mahmud Elahi, et al.
Published: (2026)
by: Akhter, Mahmud Elahi, et al.
Published: (2026)
A Claim Decomposition Benchmark for Long-form Answer Verification
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
Examining the Metrics for Document-Level Claim Extraction in Czech and Slovak
by: Makaiova, Lucia, et al.
Published: (2025)
by: Makaiova, Lucia, et al.
Published: (2025)
Semantic Consistency-Based Uncertainty Quantification for Factuality in Radiology Report Generation
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
From Confidence to Collapse in LLM Factual Robustness
by: Fastowski, Alina, et al.
Published: (2025)
by: Fastowski, Alina, et al.
Published: (2025)
Reasoning Factual Knowledge in Structured Data with Large Language Models
by: Huang, Sirui, et al.
Published: (2024)
by: Huang, Sirui, et al.
Published: (2024)
OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation
by: Lage, Lucas Fonseca, et al.
Published: (2025)
by: Lage, Lucas Fonseca, et al.
Published: (2025)
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
by: Shin, Hagyeong, et al.
Published: (2025)
by: Shin, Hagyeong, et al.
Published: (2025)
Distill and Align Decomposition for Enhanced Claim Verification
by: Magomere, Jabez, et al.
Published: (2026)
by: Magomere, Jabez, et al.
Published: (2026)
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
by: Xu, Yunqi, et al.
Published: (2024)
by: Xu, Yunqi, et al.
Published: (2024)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
by: Barnes, Jeremy, et al.
Published: (2025)
by: Barnes, Jeremy, et al.
Published: (2025)
InFact: Informativeness Alignment for Improved LLM Factuality
by: Cohen, Roi, et al.
Published: (2025)
by: Cohen, Roi, et al.
Published: (2025)
Less is More for Improving Automatic Evaluation of Factual Consistency
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
by: Li, Yifei, et al.
Published: (2026)
by: Li, Yifei, et al.
Published: (2026)
Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance?
by: Hu, Qisheng, et al.
Published: (2024)
by: Hu, Qisheng, et al.
Published: (2024)
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
by: He, Zhonghao, et al.
Published: (2025)
by: He, Zhonghao, et al.
Published: (2025)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
by: Liu, Hongjun, et al.
Published: (2025)
by: Liu, Hongjun, et al.
Published: (2025)
Enhancing Essay Scoring with Adversarial Weights Perturbation and Metric-specific AttentionPooling
by: Huang, Jiaxin, et al.
Published: (2024)
by: Huang, Jiaxin, et al.
Published: (2024)
LLMTaxo: Leveraging Large Language Models for Constructing Taxonomy of Factual Claims from Social Media
by: Zhang, Haiqi, et al.
Published: (2025)
by: Zhang, Haiqi, et al.
Published: (2025)
ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
by: Banyas, Peter, et al.
Published: (2025)
by: Banyas, Peter, et al.
Published: (2025)
ClaimPKG: Enhancing Claim Verification via Pseudo-Subgraph Generation with Lightweight Specialized LLM
by: Pham, Hoang, et al.
Published: (2025)
by: Pham, Hoang, et al.
Published: (2025)
Permutation-Consensus Listwise Judging for Robust Factuality Evaluation
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
by: Bae, Suyoung, et al.
Published: (2026)
by: Bae, Suyoung, et al.
Published: (2026)
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
by: Jiao, Rui, et al.
Published: (2025)
by: Jiao, Rui, et al.
Published: (2025)
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
by: He, Jinwen, et al.
Published: (2023)
by: He, Jinwen, et al.
Published: (2023)
Re-Ex: Revising after Explanation Reduces the Factual Errors in LLM Responses
by: Kim, Juyeon, et al.
Published: (2024)
by: Kim, Juyeon, et al.
Published: (2024)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2025)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2025)
Factuality on Demand: Controlling the Factuality-Informativeness Trade-off in Text Generation
by: Gong, Ziwei, et al.
Published: (2026)
by: Gong, Ziwei, et al.
Published: (2026)
Real-time Factuality Assessment from Adversarial Feedback
by: Chen, Sanxing, et al.
Published: (2024)
by: Chen, Sanxing, et al.
Published: (2024)
Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights
by: Chen, Yi, et al.
Published: (2026)
by: Chen, Yi, et al.
Published: (2026)
Similar Items
-
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
by: Wang, Guanghui, et al.
Published: (2025) -
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
by: Ni, Jingwei, et al.
Published: (2024) -
AlignCheck: a Semantic Open-Domain Metric for Factual Consistency Assessment
by: Aghaebrahimian, Ahmad
Published: (2025) -
Stress Testing Factual Consistency Metrics for Long-Document Summarization
by: Mujahid, Zain Muhammad, et al.
Published: (2025) -
Atomic-SNLI: Fine-Grained Natural Language Inference through Atomic Fact Decomposition
by: Huang, Minghui
Published: (2026)