DiVA: Fine-grained Factuality Verification with Agentic-Discriminative Verifier
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Hui, Yang, Muyun, Arase, Yuki |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
by: Huang, Hui, et al.
Published: (2026)
by: Huang, Hui, et al.
Published: (2026)
DiVA-DocRE: A Discriminative and Voice-Aware Paradigm for Document-Level Relation Extraction
by: Wu, Yiheng, et al.
Published: (2024)
by: Wu, Yiheng, et al.
Published: (2024)
Adaptive LoRA Merge with Parameter Pruning for Low-Resource Generation
by: Miyano, Ryota, et al.
Published: (2025)
by: Miyano, Ryota, et al.
Published: (2025)
Distilling Monolingual and Crosslingual Word-in-Context Representations
by: Arase, Yuki, et al.
Published: (2024)
by: Arase, Yuki, et al.
Published: (2024)
An In-depth Evaluation of Large Language Models in Sentence Simplification with Error-based Human Assessment
by: Wu, Xuanxin, et al.
Published: (2024)
by: Wu, Xuanxin, et al.
Published: (2024)
Hallucinated Span Detection with Multi-View Attention Features
by: Ogasa, Yuya, et al.
Published: (2025)
by: Ogasa, Yuya, et al.
Published: (2025)
Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
by: Wu, Xuanxin, et al.
Published: (2025)
by: Wu, Xuanxin, et al.
Published: (2025)
Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Aligning Sentence Simplification with ESL Learner's Proficiency for Language Acquisition
by: Li, Guanlin, et al.
Published: (2025)
by: Li, Guanlin, et al.
Published: (2025)
Edit-Constrained Decoding for Sentence Simplification
by: Zetsu, Tatsuya, et al.
Published: (2024)
by: Zetsu, Tatsuya, et al.
Published: (2024)
Improving Model Factuality with Fine-grained Critique-based Evaluator
by: Xie, Yiqing, et al.
Published: (2024)
by: Xie, Yiqing, et al.
Published: (2024)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
by: Gu, Yuzhe, et al.
Published: (2025)
by: Gu, Yuzhe, et al.
Published: (2025)
LLM-based Discriminative Reasoning for Knowledge Graph Question Answering
by: Xu, Mufan, et al.
Published: (2024)
by: Xu, Mufan, et al.
Published: (2024)
Scaling Agentic Verifier for Competitive Coding
by: Ma, Zeyao, et al.
Published: (2026)
by: Ma, Zeyao, et al.
Published: (2026)
FACTS&EVIDENCE: An Interactive Tool for Transparent Fine-Grained Factual Verification of Machine-Generated Text
by: Boonsanong, Varich, et al.
Published: (2025)
by: Boonsanong, Varich, et al.
Published: (2025)
FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
by: Chen, Xiangyan, et al.
Published: (2025)
by: Chen, Xiangyan, et al.
Published: (2025)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
by: Huang, Hui, et al.
Published: (2024)
by: Huang, Hui, et al.
Published: (2024)
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
by: Godbole, Ameya, et al.
Published: (2025)
by: Godbole, Ameya, et al.
Published: (2025)
Shrinking the Generation-Verification Gap with Weak Verifiers
by: Saad-Falcon, Jon, et al.
Published: (2025)
by: Saad-Falcon, Jon, et al.
Published: (2025)
FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
by: Chen, Mingda, et al.
Published: (2025)
by: Chen, Mingda, et al.
Published: (2025)
SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
by: Haas, Lukas, et al.
Published: (2025)
by: Haas, Lukas, et al.
Published: (2025)
DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales
by: Wan, Herun, et al.
Published: (2025)
by: Wan, Herun, et al.
Published: (2025)
From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs
by: Gong, Xuan, et al.
Published: (2025)
by: Gong, Xuan, et al.
Published: (2025)
DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
by: Wanner, Miriam, et al.
Published: (2024)
by: Wanner, Miriam, et al.
Published: (2024)
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
by: Huang, Heyuan, et al.
Published: (2025)
by: Huang, Heyuan, et al.
Published: (2025)
Agentic Verification for Ambiguous Query Disambiguation
by: Lee, Youngwon, et al.
Published: (2025)
by: Lee, Youngwon, et al.
Published: (2025)
Self-Evaluation of Large Language Model based on Glass-box Features
by: Huang, Hui, et al.
Published: (2024)
by: Huang, Hui, et al.
Published: (2024)
Fine-grained Verification via Diagnostic Reasoning Supervision for Aspect Sentiment Triplet Extraction
by: Lai, Wenna, et al.
Published: (2026)
by: Lai, Wenna, et al.
Published: (2026)
Enhancing Factual Accuracy and Citation Generation in LLMs via Multi-Stage Self-Verification
by: García, Fernando Gabriela, et al.
Published: (2025)
by: García, Fernando Gabriela, et al.
Published: (2025)
Iterate Until Retrieved: Factual Nugget Optimization for Discoverable Continual Corrections in Agentic RAG
by: Hazoom, Moshe, et al.
Published: (2026)
by: Hazoom, Moshe, et al.
Published: (2026)
Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning
by: Wang, Weiqin, et al.
Published: (2025)
by: Wang, Weiqin, et al.
Published: (2025)
Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering
by: Fan, Shicheng, et al.
Published: (2026)
by: Fan, Shicheng, et al.
Published: (2026)
Fine-Grained Self-Endorsement Improves Factuality and Reasoning
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
by: Liu, Xiaoyuan, et al.
Published: (2025)
by: Liu, Xiaoyuan, et al.
Published: (2025)
Fine-tuning Large Language Models for Improving Factuality in Legal Question Answering
by: Hu, Yinghao, et al.
Published: (2025)
by: Hu, Yinghao, et al.
Published: (2025)
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
by: Zhai, Weiqi, et al.
Published: (2026)
by: Zhai, Weiqi, et al.
Published: (2026)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
by: Zhou, Hongli, et al.
Published: (2026)
by: Zhou, Hongli, et al.
Published: (2026)
Mitigating the Bias of Large Language Model Evaluation
by: Zhou, Hongli, et al.
Published: (2024)
by: Zhou, Hongli, et al.
Published: (2024)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
by: Zhang, Jiazheng, et al.
Published: (2026)
by: Zhang, Jiazheng, et al.
Published: (2026)
Similar Items
-
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
by: Huang, Hui, et al.
Published: (2026) -
DiVA-DocRE: A Discriminative and Voice-Aware Paradigm for Document-Level Relation Extraction
by: Wu, Yiheng, et al.
Published: (2024) -
Adaptive LoRA Merge with Parameter Pruning for Low-Resource Generation
by: Miyano, Ryota, et al.
Published: (2025) -
Distilling Monolingual and Crosslingual Word-in-Context Representations
by: Arase, Yuki, et al.
Published: (2024) -
An In-depth Evaluation of Large Language Models in Sentence Simplification with Error-based Human Assessment
by: Wu, Xuanxin, et al.
Published: (2024)