Saved in:
Bibliographic Details
Main Authors: Kumar, Babu, Kumar, Gaurav, Garg, Ayush, Kishore, Aditya, Patro, Jasabanta
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.26755
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911725296549888
author Kumar, Babu
Kumar, Gaurav
Garg, Ayush
Kishore, Aditya
Patro, Jasabanta
author_facet Kumar, Babu
Kumar, Gaurav
Garg, Ayush
Kishore, Aditya
Patro, Jasabanta
contents Multilingual fact verification requires evidence that is both relevant and sufficiently complete for reliable factuality prediction. However, existing systems often rely on search snippets, sentence-level evidence, or locally segmented passages, which can miss decisive context and produce fragmented evidence. To overcome these limitations, we propose SEEK, a Semantic Evidence Extraction with an adaptive chunKing framework that constructs coherent evidence chunks from full fact-checking articles by identifying semantic topic transitions and preserving local verification context. The constructed chunks are encoded using a multilingual encoder and then multilingual LLMs are finetuned using LoRA adapter for veracity prediction. Experiments on X-FACT and RU22Fact show that SEEK improves macro-f1 by up to 10% over semantic chunking, 19% over sentence chunking, and 20% over search-snippet baselines. Evidence completeness and significance analyses further show that SEEK preserves richer verification context and enables more reliable multilingual fact-checking.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26755
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SEEK: Semantic Evidence Extraction via Adaptive ChunKing for Multilingual Fact-Checking
Kumar, Babu
Kumar, Gaurav
Garg, Ayush
Kishore, Aditya
Patro, Jasabanta
Computation and Language
Multilingual fact verification requires evidence that is both relevant and sufficiently complete for reliable factuality prediction. However, existing systems often rely on search snippets, sentence-level evidence, or locally segmented passages, which can miss decisive context and produce fragmented evidence. To overcome these limitations, we propose SEEK, a Semantic Evidence Extraction with an adaptive chunKing framework that constructs coherent evidence chunks from full fact-checking articles by identifying semantic topic transitions and preserving local verification context. The constructed chunks are encoded using a multilingual encoder and then multilingual LLMs are finetuned using LoRA adapter for veracity prediction. Experiments on X-FACT and RU22Fact show that SEEK improves macro-f1 by up to 10% over semantic chunking, 19% over sentence chunking, and 20% over search-snippet baselines. Evidence completeness and significance analyses further show that SEEK preserves richer verification context and enables more reliable multilingual fact-checking.
title SEEK: Semantic Evidence Extraction via Adaptive ChunKing for Multilingual Fact-Checking
topic Computation and Language
url https://arxiv.org/abs/2605.26755