"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
Fuente:
arXiv
Saved in:
| Main Authors: | Rathva, Harsh, Mishra, Pruthwik, Malviya, Shrikant |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A RoBERTa-Based Functional Syntax Annotation Model for Chinese Texts
by: Xiaohui, Han, et al.
Published: (2025)
by: Xiaohui, Han, et al.
Published: (2025)
EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System
by: Chaybouti, Sofian, et al.
Published: (2021)
by: Chaybouti, Sofian, et al.
Published: (2021)
Logits-Constrained Framework with RoBERTa for Ancient Chinese NER
by: Hua, Wenjie, et al.
Published: (2025)
by: Hua, Wenjie, et al.
Published: (2025)
Controllable Text Summarization: Unraveling Challenges, Approaches, and Prospects -- A Survey
by: Urlana, Ashok, et al.
Published: (2023)
by: Urlana, Ashok, et al.
Published: (2023)
Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning
by: Bhaskar, Yash, et al.
Published: (2025)
by: Bhaskar, Yash, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
by: Liu, Zhongxin, et al.
Published: (2025)
by: Liu, Zhongxin, et al.
Published: (2025)
DESS: DeBERTa Enhanced Syntactic-Semantic Aspect Sentiment Triplet Extraction
by: Thenuwara, Vishal, et al.
Published: (2025)
by: Thenuwara, Vishal, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Socially Responsible Data for Large Multilingual Language Models
by: Smart, Andrew, et al.
Published: (2024)
by: Smart, Andrew, et al.
Published: (2024)
Boosting Accuracy and Interpretability in Multilingual Hate Speech Detection Through Layer Freezing and Explainable AI
by: Bilehsavar, Meysam Shirdel, et al.
Published: (2026)
by: Bilehsavar, Meysam Shirdel, et al.
Published: (2026)
Exploring the Maze of Multilingual Modeling
by: Nezhad, Sina Bagheri, et al.
Published: (2023)
by: Nezhad, Sina Bagheri, et al.
Published: (2023)
HACK: Hallucinations Along Certainty and Knowledge Axes
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Distinguishing Ignorance from Error in LLM Hallucinations
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
Grade Guard: A Smart System for Short Answer Automated Grading
by: Dadu, Niharika, et al.
Published: (2025)
by: Dadu, Niharika, et al.
Published: (2025)
Towards Massive Multilingual Holistic Bias
by: Tan, Xiaoqing Ellen, et al.
Published: (2024)
by: Tan, Xiaoqing Ellen, et al.
Published: (2024)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
by: Cherif, Ahmed
Published: (2026)
by: Cherif, Ahmed
Published: (2026)
What Drives Performance in Multilingual Language Models?
by: Nezhad, Sina Bagheri, et al.
Published: (2024)
by: Nezhad, Sina Bagheri, et al.
Published: (2024)
Embedded Safety-Aligned Intelligence via Differentiable Internal Alignment Embeddings
by: Rathva, Harsh, et al.
Published: (2025)
by: Rathva, Harsh, et al.
Published: (2025)
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models
by: Benkirane, Kenza, et al.
Published: (2024)
by: Benkirane, Kenza, et al.
Published: (2024)
idT5: Indonesian Version of Multilingual T5 Transformer
by: Fuadi, Mukhlish, et al.
Published: (2023)
by: Fuadi, Mukhlish, et al.
Published: (2023)
DimStance: Multilingual Datasets for Dimensional Stance Analysis
by: Becker, Jonas, et al.
Published: (2026)
by: Becker, Jonas, et al.
Published: (2026)
Ensembling Multilingual Transformers for Robust Sentiment Analysis of Tweets
by: Bilehsavar, Meysam Shirdel, et al.
Published: (2025)
by: Bilehsavar, Meysam Shirdel, et al.
Published: (2025)
MedHal: An Evaluation Dataset for Medical Hallucination Detection
by: Mehenni, Gaya, et al.
Published: (2025)
by: Mehenni, Gaya, et al.
Published: (2025)
Enhancing Hate Speech Detection on Social Media: A Comparative Analysis of Machine Learning Models and Text Transformation Approaches
by: Mishra, Saurabh, et al.
Published: (2026)
by: Mishra, Saurabh, et al.
Published: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
ML-Promise: A Multilingual Dataset for Corporate Promise Verification
by: Seki, Yohei, et al.
Published: (2024)
by: Seki, Yohei, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
by: Kodali, Prashant, et al.
Published: (2025)
by: Kodali, Prashant, et al.
Published: (2025)
German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset
by: Mascarell, Laura, et al.
Published: (2024)
by: Mascarell, Laura, et al.
Published: (2024)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
by: Lee, Lung-Hao, et al.
Published: (2026)
by: Lee, Lung-Hao, et al.
Published: (2026)
Self-Supervised Borrowing Detection on Multilingual Wordlists
by: Wientzek, Tim
Published: (2025)
by: Wientzek, Tim
Published: (2025)
I run as fast as a rabbit, can you? A Multilingual Simile Dialogue Dataset
by: Ma, Longxuan, et al.
Published: (2023)
by: Ma, Longxuan, et al.
Published: (2023)
Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
by: Dang, Thao Anh, et al.
Published: (2024)
by: Dang, Thao Anh, et al.
Published: (2024)
Multilingual jailbreaking of LLMs using low-resource languages
by: Marx, Dylan, et al.
Published: (2026)
by: Marx, Dylan, et al.
Published: (2026)
Seeing Through the Fog: A Cost-Effectiveness Analysis of Hallucination Detection Systems
by: Thomas, Alexander, et al.
Published: (2024)
by: Thomas, Alexander, et al.
Published: (2024)
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Similar Items
-
A RoBERTa-Based Functional Syntax Annotation Model for Chinese Texts
by: Xiaohui, Han, et al.
Published: (2025) -
EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System
by: Chaybouti, Sofian, et al.
Published: (2021) -
Logits-Constrained Framework with RoBERTa for Ancient Chinese NER
by: Hua, Wenjie, et al.
Published: (2025) -
Controllable Text Summarization: Unraveling Challenges, Approaches, and Prospects -- A Survey
by: Urlana, Ashok, et al.
Published: (2023) -
Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning
by: Bhaskar, Yash, et al.
Published: (2025)