Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Tianyi, Xu, Samuel, Dang, Jason Tansong, Yan, Samuel, Yin, Kimberley |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
by: Ng, Ming Pok, et al.
Published: (2025)
by: Ng, Ming Pok, et al.
Published: (2025)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
A Claim Decomposition Benchmark for Long-form Answer Verification
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
Multi-Objective Linguistic Control of Large Language Models
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
by: Huang, Heyuan, et al.
Published: (2025)
by: Huang, Heyuan, et al.
Published: (2025)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
by: Chua, James, et al.
Published: (2026)
by: Chua, James, et al.
Published: (2026)
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
by: Yao, Jihan, et al.
Published: (2024)
by: Yao, Jihan, et al.
Published: (2024)
CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
by: Yan, Tianyi Lorena, et al.
Published: (2025)
by: Yan, Tianyi Lorena, et al.
Published: (2025)
Retrieve-Refine-Calibrate: A Framework for Complex Claim Fact-Checking
by: Sun, Mingwei, et al.
Published: (2026)
by: Sun, Mingwei, et al.
Published: (2026)
Core: Robust Factual Precision with Informative Sub-Claim Identification
by: Jiang, Zhengping, et al.
Published: (2024)
by: Jiang, Zhengping, et al.
Published: (2024)
Optimizing Decomposition for Optimal Claim Verification
by: Lu, Yining, et al.
Published: (2025)
by: Lu, Yining, et al.
Published: (2025)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Exploring Modularity of Agentic Systems for Drug Discovery
by: van Weesep, Laura, et al.
Published: (2025)
by: van Weesep, Laura, et al.
Published: (2025)
Agentic Confidence Calibration
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
CAMOUFLAGE: Exploiting Misinformation Detection Systems Through LLM-driven Adversarial Claim Transformation
by: Bethany, Mazal, et al.
Published: (2025)
by: Bethany, Mazal, et al.
Published: (2025)
Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification
by: Cao, Zongwan, et al.
Published: (2026)
by: Cao, Zongwan, et al.
Published: (2026)
ClaimVer: Explainable Claim-Level Verification and Evidence Attribution of Text Through Knowledge Graphs
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
Span-Level Hallucination Detection for LLM-Generated Answers
by: Elchafei, Passant, et al.
Published: (2025)
by: Elchafei, Passant, et al.
Published: (2025)
CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading
by: Raikote, Pranav, et al.
Published: (2026)
by: Raikote, Pranav, et al.
Published: (2026)
SYN-DIGITS: A Synthetic Control Framework for Calibrated Digital Twin Simulation
by: Fan, Grace Jiarui, et al.
Published: (2026)
by: Fan, Grace Jiarui, et al.
Published: (2026)
MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation
by: Ji, Yuelyu, et al.
Published: (2026)
by: Ji, Yuelyu, et al.
Published: (2026)
Characterizing AlphaEarth Embedding Geometry for Agentic Environmental Reasoning
by: Rahman, Mashrekur, et al.
Published: (2026)
by: Rahman, Mashrekur, et al.
Published: (2026)
Credence Calibration Game? Calibrating Large Language Models through Structured Play
by: Fang, Ke, et al.
Published: (2025)
by: Fang, Ke, et al.
Published: (2025)
Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality
by: Luo, Wen, et al.
Published: (2026)
by: Luo, Wen, et al.
Published: (2026)
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning
by: He, Qi, et al.
Published: (2025)
by: He, Qi, et al.
Published: (2025)
Scientific QA System with Verifiable Answers
by: Ljajić, Adela, et al.
Published: (2024)
by: Ljajić, Adela, et al.
Published: (2024)
eTracer: Towards Traceable Text Generation via Claim-Level Grounding
by: Chu, Bohao, et al.
Published: (2026)
by: Chu, Bohao, et al.
Published: (2026)
The Master-Slave Encoder Model for Improving Patent Text Summarization: A New Approach to Combining Specifications and Claims
by: Zhou, Shu, et al.
Published: (2024)
by: Zhou, Shu, et al.
Published: (2024)
Calibrating Large Language Models Using Their Generations Only
by: Ulmer, Dennis, et al.
Published: (2024)
by: Ulmer, Dennis, et al.
Published: (2024)
Examining the Metrics for Document-Level Claim Extraction in Czech and Slovak
by: Makaiova, Lucia, et al.
Published: (2025)
by: Makaiova, Lucia, et al.
Published: (2025)
ClaimFlow: Tracing the Evolution of Scientific Claims in NLP
by: Pramanick, Aniket, et al.
Published: (2026)
by: Pramanick, Aniket, et al.
Published: (2026)
MuSciClaims: Multimodal Scientific Claim Verification
by: Lal, Yash Kumar, et al.
Published: (2025)
by: Lal, Yash Kumar, et al.
Published: (2025)
Distill and Align Decomposition for Enhanced Claim Verification
by: Magomere, Jabez, et al.
Published: (2026)
by: Magomere, Jabez, et al.
Published: (2026)
Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG
by: Qiu, Longpeng, et al.
Published: (2025)
by: Qiu, Longpeng, et al.
Published: (2025)
Precise Model Benchmarking with Only a Few Observations
by: Fogliato, Riccardo, et al.
Published: (2024)
by: Fogliato, Riccardo, et al.
Published: (2024)
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
by: Lu, Haiquan, et al.
Published: (2026)
by: Lu, Haiquan, et al.
Published: (2026)
Controllable Decontextualization of Yes/No Question and Answers into Factual Statements
by: Mo, Lingbo, et al.
Published: (2024)
by: Mo, Lingbo, et al.
Published: (2024)
Agentic Separation Logic Specification Synthesis
by: Suresh, Tarun, et al.
Published: (2026)
by: Suresh, Tarun, et al.
Published: (2026)
Similar Items
-
MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
by: Ng, Ming Pok, et al.
Published: (2025) -
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025) -
A Claim Decomposition Benchmark for Long-form Answer Verification
by: Zhang, Zhihao, et al.
Published: (2024) -
Multi-Objective Linguistic Control of Large Language Models
by: Nguyen, Dang, et al.
Published: (2024) -
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
by: Huang, Heyuan, et al.
Published: (2025)