A Claim Decomposition Benchmark for Long-form Answer Verification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhihao, Fan, Yixing, Zhang, Ruqing, Guo, Jiafeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Capacity of Citation Generation by Large Language Models
von: Qian, Haosheng, et al.
Veröffentlicht: (2024)
von: Qian, Haosheng, et al.
Veröffentlicht: (2024)
Robust Neural Information Retrieval: An Adversarial and Out-of-distribution Perspective
von: Liu, Yu-An, et al.
Veröffentlicht: (2024)
von: Liu, Yu-An, et al.
Veröffentlicht: (2024)
QUITO: Accelerating Long-Context Reasoning through Query-Guided Context Compression
von: Wang, Wenshan, et al.
Veröffentlicht: (2024)
von: Wang, Wenshan, et al.
Veröffentlicht: (2024)
Optimizing Decomposition for Optimal Claim Verification
von: Lu, Yining, et al.
Veröffentlicht: (2025)
von: Lu, Yining, et al.
Veröffentlicht: (2025)
The Alignment Bottleneck in Decomposition-Based Claim Verification
von: Akhter, Mahmud Elahi, et al.
Veröffentlicht: (2026)
von: Akhter, Mahmud Elahi, et al.
Veröffentlicht: (2026)
Distill and Align Decomposition for Enhanced Claim Verification
von: Magomere, Jabez, et al.
Veröffentlicht: (2026)
von: Magomere, Jabez, et al.
Veröffentlicht: (2026)
Controlling Risk of Retrieval-augmented Generation: A Counterfactual Prompting Framework
von: Chen, Lu, et al.
Veröffentlicht: (2024)
von: Chen, Lu, et al.
Veröffentlicht: (2024)
Robust Claim Verification Through Fact Detection
von: Jafari, Nazanin, et al.
Veröffentlicht: (2024)
von: Jafari, Nazanin, et al.
Veröffentlicht: (2024)
Claim Verification in the Age of Large Language Models: A Survey
von: Dmonte, Alphaeus, et al.
Veröffentlicht: (2024)
von: Dmonte, Alphaeus, et al.
Veröffentlicht: (2024)
BiDeV: Bilateral Defusing Verification for Complex Claim Fact-Checking
von: Liu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Liu, Yuxuan, et al.
Veröffentlicht: (2025)
HiFACTMix: A Code-Mixed Benchmark and Graph-Aware Model for EvidenceBased Political Claim Verification in Hinglish
von: Thakur, Rakesh, et al.
Veröffentlicht: (2025)
von: Thakur, Rakesh, et al.
Veröffentlicht: (2025)
QUITO-X: A New Perspective on Context Compression from the Information Bottleneck Theory
von: Wang, Yihang, et al.
Veröffentlicht: (2024)
von: Wang, Yihang, et al.
Veröffentlicht: (2024)
Evergreen: Efficient Claim Verification for Semantic Aggregates
von: Lee, Alexander W., et al.
Veröffentlicht: (2026)
von: Lee, Alexander W., et al.
Veröffentlicht: (2026)
Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
TrustRAG: An Information Assistant with Retrieval Augmented Generation
von: Fan, Yixing, et al.
Veröffentlicht: (2025)
von: Fan, Yixing, et al.
Veröffentlicht: (2025)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
von: Zhang, Yushun, et al.
Veröffentlicht: (2026)
von: Zhang, Yushun, et al.
Veröffentlicht: (2026)
Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method
von: Zhang, Weichao, et al.
Veröffentlicht: (2024)
von: Zhang, Weichao, et al.
Veröffentlicht: (2024)
ClaimPKG: Enhancing Claim Verification via Pseudo-Subgraph Generation with Lightweight Specialized LLM
von: Pham, Hoang, et al.
Veröffentlicht: (2025)
von: Pham, Hoang, et al.
Veröffentlicht: (2025)
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
von: Liu, Muxin, et al.
Veröffentlicht: (2026)
von: Liu, Muxin, et al.
Veröffentlicht: (2026)
Step-by-Step Fact Verification System for Medical Claims with Explainable Reasoning
von: Vladika, Juraj, et al.
Veröffentlicht: (2025)
von: Vladika, Juraj, et al.
Veröffentlicht: (2025)
Sandwich Reasoning: An Answer-Reasoning-Answer Approach for Low-Latency Query Correction
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs
von: Huang, Minghui
Veröffentlicht: (2025)
von: Huang, Minghui
Veröffentlicht: (2025)
Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
von: Wei, Wenda, et al.
Veröffentlicht: (2025)
von: Wei, Wenda, et al.
Veröffentlicht: (2025)
Fine-grained Claim-level RAG Benchmark for Law
von: Das, Souvick, et al.
Veröffentlicht: (2026)
von: Das, Souvick, et al.
Veröffentlicht: (2026)
Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance?
von: Hu, Qisheng, et al.
Veröffentlicht: (2024)
von: Hu, Qisheng, et al.
Veröffentlicht: (2024)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
Comparing Knowledge Sources for Open-Domain Scientific Claim Verification
von: Vladika, Juraj, et al.
Veröffentlicht: (2024)
von: Vladika, Juraj, et al.
Veröffentlicht: (2024)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
von: Yao, Siyang, et al.
Veröffentlicht: (2026)
von: Yao, Siyang, et al.
Veröffentlicht: (2026)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)
MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
von: Ng, Ming Pok, et al.
Veröffentlicht: (2025)
von: Ng, Ming Pok, et al.
Veröffentlicht: (2025)
LoGU: Long-form Generation with Uncertainty Expressions
von: Yang, Ruihan, et al.
Veröffentlicht: (2024)
von: Yang, Ruihan, et al.
Veröffentlicht: (2024)
Towards Robust Universal Information Extraction: Benchmark, Evaluation, and Solution
von: Zhu, Jizhao, et al.
Veröffentlicht: (2025)
von: Zhu, Jizhao, et al.
Veröffentlicht: (2025)
Optimizing Long-Form Clinical Text Generation with Claim-Based Rewards
von: Jhaveri, Samyak, et al.
Veröffentlicht: (2025)
von: Jhaveri, Samyak, et al.
Veröffentlicht: (2025)
Argumentative Large Language Models for Explainable and Contestable Claim Verification
von: Freedman, Gabriel, et al.
Veröffentlicht: (2024)
von: Freedman, Gabriel, et al.
Veröffentlicht: (2024)
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
A Comparative Study of Specialized LLMs as Dense Retrievers
von: Zhang, Hengran, et al.
Veröffentlicht: (2025)
von: Zhang, Hengran, et al.
Veröffentlicht: (2025)
MAPLE: Micro Analysis of Pairwise Language Evolution for Few-Shot Claim Verification
von: Zeng, Xia, et al.
Veröffentlicht: (2024)
von: Zeng, Xia, et al.
Veröffentlicht: (2024)
Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation
von: Zheng, Tianyu, et al.
Veröffentlicht: (2024)
von: Zheng, Tianyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Capacity of Citation Generation by Large Language Models
von: Qian, Haosheng, et al.
Veröffentlicht: (2024) -
Robust Neural Information Retrieval: An Adversarial and Out-of-distribution Perspective
von: Liu, Yu-An, et al.
Veröffentlicht: (2024) -
QUITO: Accelerating Long-Context Reasoning through Query-Guided Context Compression
von: Wang, Wenshan, et al.
Veröffentlicht: (2024) -
Optimizing Decomposition for Optimal Claim Verification
von: Lu, Yining, et al.
Veröffentlicht: (2025) -
The Alignment Bottleneck in Decomposition-Based Claim Verification
von: Akhter, Mahmud Elahi, et al.
Veröffentlicht: (2026)