On Finding Inconsistencies in Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Lovering, Charles J., Ebner, Seth, Smock, Brandon, Krumdick, Michael, Rabbani, Saad, Muhammad, Ahmed, Reddy, Varshini, Tanner, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
by: Krumdick, Michael, et al.
Published: (2025)
by: Krumdick, Michael, et al.
Published: (2025)
Cost-Efficient Estimation of General Abilities Across Benchmarks
by: Krumdick, Michael, et al.
Published: (2026)
by: Krumdick, Michael, et al.
Published: (2026)
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024)
by: Lai, Viet Dac, et al.
Published: (2024)
DocFinQA: A Long-Context Financial Reasoning Dataset
by: Reddy, Varshini, et al.
Published: (2024)
by: Reddy, Varshini, et al.
Published: (2024)
Tokenization with Split Trees
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
The Effect of Scripts and Formats on LLM Numeracy
by: Reddy, Varshini, et al.
Published: (2026)
by: Reddy, Varshini, et al.
Published: (2026)
Language Model Probabilities are Not Calibrated in Numeric Contexts
by: Lovering, Charles, et al.
Published: (2024)
by: Lovering, Charles, et al.
Published: (2024)
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
by: Krumdick, Michael, et al.
Published: (2026)
by: Krumdick, Michael, et al.
Published: (2026)
An Analysis of Multilingual FActScore
by: Vu, Kim Trong, et al.
Published: (2024)
by: Vu, Kim Trong, et al.
Published: (2024)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
by: Reddy, Varshini, et al.
Published: (2025)
by: Reddy, Varshini, et al.
Published: (2025)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
by: Schmidt, Craig W., et al.
Published: (2025)
by: Schmidt, Craig W., et al.
Published: (2025)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
by: Hu, Yifan, et al.
Published: (2025)
by: Hu, Yifan, et al.
Published: (2025)
Not How Many, But Which: Parameter Placement in Low-Rank Adaptation
by: Sehanobish, Arijit, et al.
Published: (2026)
by: Sehanobish, Arijit, et al.
Published: (2026)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
by: Chang, Yapei, et al.
Published: (2025)
by: Chang, Yapei, et al.
Published: (2025)
Tokenization Is More Than Compression
by: Schmidt, Craig W., et al.
Published: (2024)
by: Schmidt, Craig W., et al.
Published: (2024)
A Closer Look at Claim Decomposition
by: Wanner, Miriam, et al.
Published: (2024)
by: Wanner, Miriam, et al.
Published: (2024)
FIZZ: Factual Inconsistency Detection by Zoom-in Summary and Zoom-out Document
by: Yang, Joonho, et al.
Published: (2024)
by: Yang, Joonho, et al.
Published: (2024)
Misleading through Inconsistency: A Benchmark for Political Inconsistencies Detection
by: Sagimbayeva, Nursulu, et al.
Published: (2025)
by: Sagimbayeva, Nursulu, et al.
Published: (2025)
Faster Superword Tokenization
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs
by: Tan, Nelvin, et al.
Published: (2026)
by: Tan, Nelvin, et al.
Published: (2026)
Fast and Accurate Factual Inconsistency Detection Over Long Documents
by: Lattimer, Barrett Martin, et al.
Published: (2023)
by: Lattimer, Barrett Martin, et al.
Published: (2023)
Automated Quality Control for Language Documentation: Detecting Phonotactic Inconsistencies in a Kokborok Wordlist
by: van Dam, Kellen Parker, et al.
Published: (2025)
by: van Dam, Kellen Parker, et al.
Published: (2025)
Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization
by: Zhong, Yang, et al.
Published: (2025)
by: Zhong, Yang, et al.
Published: (2025)
Finding a Needle in the Adversarial Haystack: A Targeted Paraphrasing Approach For Uncovering Edge Cases with Minimal Distribution Distortion
by: Kassem, Aly M., et al.
Published: (2024)
by: Kassem, Aly M., et al.
Published: (2024)
Multi Class Depression Detection Through Tweets using Artificial Intelligence
by: Nusrat, Muhammad Osama, et al.
Published: (2024)
by: Nusrat, Muhammad Osama, et al.
Published: (2024)
Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing
by: Ahn, Jihyun Janice, et al.
Published: (2025)
by: Ahn, Jihyun Janice, et al.
Published: (2025)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
by: Uzan, Omri, et al.
Published: (2024)
by: Uzan, Omri, et al.
Published: (2024)
DRIFT: Detecting Representational Inconsistencies for Factual Truthfulness
by: Bhatnagar, Rohan, et al.
Published: (2026)
by: Bhatnagar, Rohan, et al.
Published: (2026)
Localizing Factual Inconsistencies in Attributable Text Generation
by: Cattan, Arie, et al.
Published: (2024)
by: Cattan, Arie, et al.
Published: (2024)
Inconsistencies in Masked Language Models
by: Young, Tom, et al.
Published: (2022)
by: Young, Tom, et al.
Published: (2022)
Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models
by: Semnani, Sina J., et al.
Published: (2025)
by: Semnani, Sina J., et al.
Published: (2025)
Core: Robust Factual Precision with Informative Sub-Claim Identification
by: Jiang, Zhengping, et al.
Published: (2024)
by: Jiang, Zhengping, et al.
Published: (2024)
Diagnosing and Mitigating Semantic Inconsistencies in Wikidata's Classification Hierarchy
by: Zhao, Shixiong, et al.
Published: (2025)
by: Zhao, Shixiong, et al.
Published: (2025)
Measuring the Inconsistency of Large Language Models in Preferential Ranking
by: Zhao, Xiutian, et al.
Published: (2024)
by: Zhao, Xiutian, et al.
Published: (2024)
Inconsistent dialogue responses and how to recover from them
by: Zhang, Mian, et al.
Published: (2024)
by: Zhang, Mian, et al.
Published: (2024)
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
by: Rabbani, Parisa, et al.
Published: (2025)
by: Rabbani, Parisa, et al.
Published: (2025)
Project MOSLA: Recording Every Moment of Second Language Acquisition
by: Hagiwara, Masato, et al.
Published: (2024)
by: Hagiwara, Masato, et al.
Published: (2024)
EnronQA: Towards Personalized RAG over Private Documents
by: Ryan, Michael J., et al.
Published: (2025)
by: Ryan, Michael J., et al.
Published: (2025)
Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
by: Haldar, Rajarshi, et al.
Published: (2025)
by: Haldar, Rajarshi, et al.
Published: (2025)
Similar Items
-
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
by: Krumdick, Michael, et al.
Published: (2025) -
Cost-Efficient Estimation of General Abilities Across Benchmarks
by: Krumdick, Michael, et al.
Published: (2026) -
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023) -
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024) -
DocFinQA: A Long-Context Financial Reasoning Dataset
by: Reddy, Varshini, et al.
Published: (2024)