When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Muxin, Rao, Delip, Kim, Grace, Callison-Burch, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Autorubric: Unifying Rubric-based LLM Evaluation
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
NSF-SciFy: Mining the NSF Awards Database for Scientific Claims
by: Rao, Delip, et al.
Published: (2025)
by: Rao, Delip, et al.
Published: (2025)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
by: Zhu, Andrew, et al.
Published: (2025)
by: Zhu, Andrew, et al.
Published: (2025)
You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling
by: Song, Jaewoo, et al.
Published: (2024)
by: Song, Jaewoo, et al.
Published: (2024)
WithdrarXiv: A Large-Scale Dataset for Retraction Study
by: Rao, Delip, et al.
Published: (2024)
by: Rao, Delip, et al.
Published: (2024)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
by: Zhu, Andrew, et al.
Published: (2025)
by: Zhu, Andrew, et al.
Published: (2025)
FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models
by: Zhu, Andrew, et al.
Published: (2024)
by: Zhu, Andrew, et al.
Published: (2024)
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
by: Roy, Amartya, et al.
Published: (2026)
by: Roy, Amartya, et al.
Published: (2026)
ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer
by: Horvitz, Zachary, et al.
Published: (2023)
by: Horvitz, Zachary, et al.
Published: (2023)
Concept Lancet: Image Editing with Compositional Representation Transplant
by: Luo, Jinqi, et al.
Published: (2025)
by: Luo, Jinqi, et al.
Published: (2025)
Optimizing Decomposition for Optimal Claim Verification
by: Lu, Yining, et al.
Published: (2025)
by: Lu, Yining, et al.
Published: (2025)
The Alignment Bottleneck in Decomposition-Based Claim Verification
by: Akhter, Mahmud Elahi, et al.
Published: (2026)
by: Akhter, Mahmud Elahi, et al.
Published: (2026)
Robust Claim Verification Through Fact Detection
by: Jafari, Nazanin, et al.
Published: (2024)
by: Jafari, Nazanin, et al.
Published: (2024)
ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training
by: Han, Feijiang, et al.
Published: (2025)
by: Han, Feijiang, et al.
Published: (2025)
Probabilistic Soundness Guarantees in LLM Reasoning Chains
by: You, Weiqiu, et al.
Published: (2025)
by: You, Weiqiu, et al.
Published: (2025)
Evergreen: Efficient Claim Verification for Semantic Aggregates
by: Lee, Alexander W., et al.
Published: (2026)
by: Lee, Alexander W., et al.
Published: (2026)
Distill and Align Decomposition for Enhanced Claim Verification
by: Magomere, Jabez, et al.
Published: (2026)
by: Magomere, Jabez, et al.
Published: (2026)
Claim Verification in the Age of Large Language Models: A Survey
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
A Claim Decomposition Benchmark for Long-form Answer Verification
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
Large Language Models Can Self-Improve At Web Agent Tasks
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
BiDeV: Bilateral Defusing Verification for Complex Claim Fact-Checking
by: Liu, Yuxuan, et al.
Published: (2025)
by: Liu, Yuxuan, et al.
Published: (2025)
ClaimPKG: Enhancing Claim Verification via Pseudo-Subgraph Generation with Lightweight Specialized LLM
by: Pham, Hoang, et al.
Published: (2025)
by: Pham, Hoang, et al.
Published: (2025)
Step-by-Step Fact Verification System for Medical Claims with Explainable Reasoning
by: Vladika, Juraj, et al.
Published: (2025)
by: Vladika, Juraj, et al.
Published: (2025)
How LLMs Fail to Support Fact-Checking
by: Proma, Adiba Mahbub, et al.
Published: (2025)
by: Proma, Adiba Mahbub, et al.
Published: (2025)
Evaluating Vision-Language Models on Bistable Images
by: Panagopoulou, Artemis, et al.
Published: (2024)
by: Panagopoulou, Artemis, et al.
Published: (2024)
Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
by: Rakshit, Sushrita, et al.
Published: (2026)
by: Rakshit, Sushrita, et al.
Published: (2026)
Comparing Knowledge Sources for Open-Domain Scientific Claim Verification
by: Vladika, Juraj, et al.
Published: (2024)
by: Vladika, Juraj, et al.
Published: (2024)
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
by: Lee, Seongyun, et al.
Published: (2026)
by: Lee, Seongyun, et al.
Published: (2026)
Domain Gating Ensemble Networks for AI-Generated Text Detection
by: Tripathi, Arihant, et al.
Published: (2025)
by: Tripathi, Arihant, et al.
Published: (2025)
When Claims Evolve: Evaluating and Enhancing the Robustness of Embedding Models Against Misinformation Edits
by: Magomere, Jabez, et al.
Published: (2025)
by: Magomere, Jabez, et al.
Published: (2025)
MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
by: Ng, Ming Pok, et al.
Published: (2025)
by: Ng, Ming Pok, et al.
Published: (2025)
How Transformers Reject Wrong Answers: Rotational Dynamics of Factual Constraint Processing
by: Marín, Javier
Published: (2026)
by: Marín, Javier
Published: (2026)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
by: Mehrafarin, Houman, et al.
Published: (2026)
by: Mehrafarin, Houman, et al.
Published: (2026)
Argumentative Large Language Models for Explainable and Contestable Claim Verification
by: Freedman, Gabriel, et al.
Published: (2024)
by: Freedman, Gabriel, et al.
Published: (2024)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
by: Cho, Jay Hyeon, et al.
Published: (2025)
by: Cho, Jay Hyeon, et al.
Published: (2025)
Similar Items
-
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
by: Rao, Delip, et al.
Published: (2026) -
Autorubric: Unifying Rubric-based LLM Evaluation
by: Rao, Delip, et al.
Published: (2026) -
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
by: Rao, Delip, et al.
Published: (2026) -
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
by: Rao, Delip, et al.
Published: (2026) -
NSF-SciFy: Mining the NSF Awards Database for Scientific Claims
by: Rao, Delip, et al.
Published: (2025)