BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Adib, Shefayat E Shams, Sani, Ahmed Alfey, Esham, Ekramul Alam, Abrar, Ajwad, Tashdeed, Ishmam, Chowdhury, Md Taukir Azam |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
par: Adib, Shefayat E Shams, et autres
Publié: (2026)
par: Adib, Shefayat E Shams, et autres
Publié: (2026)
LinguIUTics at PsyDefDetect: Iterative Imbalance-Aware Fine-tuning of Qwen3-8B for Psychological Defense Mechanism Classification
par: Adib, Shefayat E Shams, et autres
Publié: (2026)
par: Adib, Shefayat E Shams, et autres
Publié: (2026)
Addressing Data Scarcity in Bangla Fake News Detection: An LLM-Based Dataset Augmentation Approach
par: Sani, Ahmed Alfey, et autres
Publié: (2026)
par: Sani, Ahmed Alfey, et autres
Publié: (2026)
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
par: Rafid, Ahmed, et autres
Publié: (2026)
par: Rafid, Ahmed, et autres
Publié: (2026)
Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
par: Tashdeed, Ishmam, et autres
Publié: (2025)
par: Tashdeed, Ishmam, et autres
Publié: (2025)
Progressive Code Integration for Abstractive Bug Report Summarization
par: Karim, Shaira Sadia, et autres
Publié: (2025)
par: Karim, Shaira Sadia, et autres
Publié: (2025)
Visual Robustness Benchmark for Visual Question Answering (VQA)
par: Ishmam, Md Farhan, et autres
Publié: (2024)
par: Ishmam, Md Farhan, et autres
Publié: (2024)
MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification
par: Alam, Kazi Samin Yasar, et autres
Publié: (2026)
par: Alam, Kazi Samin Yasar, et autres
Publié: (2026)
BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis
par: Alam, Sadia, et autres
Publié: (2024)
par: Alam, Sadia, et autres
Publié: (2024)
BengaliSent140: A Large-Scale Bengali Binary Sentiment Dataset for Hate and Non-Hate Speech Classification
par: Islam, Akif, et autres
Publié: (2026)
par: Islam, Akif, et autres
Publié: (2026)
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems
par: Saadat, Asir, et autres
Publié: (2024)
par: Saadat, Asir, et autres
Publié: (2024)
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
par: Alansari, Aisha, et autres
Publié: (2025)
par: Alansari, Aisha, et autres
Publié: (2025)
Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization
par: Abrar, Ajwad, et autres
Publié: (2025)
par: Abrar, Ajwad, et autres
Publié: (2025)
HalluSearch at SemEval-2025 Task 3: A Search-Enhanced RAG Pipeline for Hallucination Detection
par: Abdallah, Mohamed A., et autres
Publié: (2025)
par: Abdallah, Mohamed A., et autres
Publié: (2025)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
par: Hosseini, Mohammad, et autres
Publié: (2025)
par: Hosseini, Mohammad, et autres
Publié: (2025)
CogniAlign: Survivability-Grounded Multi-Agent Moral Reasoning for Safe and Transparent AI
par: Ali, Hasin Jawad, et autres
Publié: (2025)
par: Ali, Hasin Jawad, et autres
Publié: (2025)
Faithful Summarization of Consumer Health Queries: A Cross-Lingual Framework with LLMs
par: Abrar, Ajwad, et autres
Publié: (2025)
par: Abrar, Ajwad, et autres
Publié: (2025)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
par: Kabir, Mohsinul, et autres
Publié: (2025)
par: Kabir, Mohsinul, et autres
Publié: (2025)
HalluClean: A Unified Framework to Combat Hallucinations in LLMs
par: Zhao, Yaxin, et autres
Publié: (2025)
par: Zhao, Yaxin, et autres
Publié: (2025)
HalluLens: LLM Hallucination Benchmark
par: Bang, Yejin, et autres
Publié: (2025)
par: Bang, Yejin, et autres
Publié: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
par: Cherif, Ahmed
Publié: (2026)
par: Cherif, Ahmed
Publié: (2026)
Ontology-Guided Query Expansion for Biomedical Document Retrieval using Large Language Models
par: Nazi, Zabir Al, et autres
Publié: (2025)
par: Nazi, Zabir Al, et autres
Publié: (2025)
From Chat to Checkup: Can Large Language Models Assist in Diabetes Prediction?
par: Sakib, Shadman, et autres
Publié: (2025)
par: Sakib, Shadman, et autres
Publié: (2025)
Retrieval Augmented Enhanced Dual Co-Attention Framework for Target Aware Multimodal Bengali Hateful Meme Detection
par: Tanvir, Raihan, et autres
Publié: (2026)
par: Tanvir, Raihan, et autres
Publié: (2026)
HausaNLP at SemEval-2025 Task 3: Towards a Fine-Grained Model-Aware Hallucination Detection
par: Bala, Maryam, et autres
Publié: (2025)
par: Bala, Maryam, et autres
Publié: (2025)
Enhancing UAV Security Through Zero Trust Architecture: An Advanced Deep Learning and Explainable AI Analysis
par: Haque, Ekramul, et autres
Publié: (2024)
par: Haque, Ekramul, et autres
Publié: (2024)
BenLLMEval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP
par: Kabir, Mohsinul, et autres
Publié: (2023)
par: Kabir, Mohsinul, et autres
Publié: (2023)
When Words Don't Mean What They Say: Figurative Understanding in Bengali Idioms
par: Sakhawat, Adib, et autres
Publié: (2026)
par: Sakhawat, Adib, et autres
Publié: (2026)
Motamot: A Dataset for Revealing the Supremacy of Large Language Models over Transformer Models in Bengali Political Sentiment Analysis
par: Faria, Fatema Tuj Johora, et autres
Publié: (2024)
par: Faria, Fatema Tuj Johora, et autres
Publié: (2024)
Tackling Fake News in Bengali: Unraveling the Impact of Summarization vs. Augmentation on Pre-trained Language Models
par: Chowdhury, Arman Sakif, et autres
Publié: (2023)
par: Chowdhury, Arman Sakif, et autres
Publié: (2023)
HalluZig: Hallucination Detection using Zigzag Persistence
par: Samaga, Shreyas N., et autres
Publié: (2026)
par: Samaga, Shreyas N., et autres
Publié: (2026)
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
par: Emery, Deanna, et autres
Publié: (2025)
par: Emery, Deanna, et autres
Publié: (2025)
HausaNLP at SemEval-2025 Task 11: Hausa Text Emotion Detection
par: Sani, Sani Abdullahi, et autres
Publié: (2025)
par: Sani, Sani Abdullahi, et autres
Publié: (2025)
BanglaMedQA and BanglaMMedBench: Evaluating Retrieval-Augmented Generation Strategies for Bangla Biomedical Question Answering
par: Sultana, Sadia, et autres
Publié: (2025)
par: Sultana, Sadia, et autres
Publié: (2025)
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
par: Urlana, Ashok, et autres
Publié: (2025)
par: Urlana, Ashok, et autres
Publié: (2025)
HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents
par: Jin, Chao, et autres
Publié: (2026)
par: Jin, Chao, et autres
Publié: (2026)
HalluCana: Fixing LLM Hallucination with A Canary Lookahead
par: Li, Tianyi, et autres
Publié: (2024)
par: Li, Tianyi, et autres
Publié: (2024)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
par: Fan, Dongyang, et autres
Publié: (2026)
par: Fan, Dongyang, et autres
Publié: (2026)
HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
par: Yeh, Min-Hsuan, et autres
Publié: (2025)
par: Yeh, Min-Hsuan, et autres
Publié: (2025)
Improving Multi-turn Task Completion in Task-Oriented Dialog Systems via Prompt Chaining and Fine-Grained Feedback
par: Fereidouni, Moghis, et autres
Publié: (2025)
par: Fereidouni, Moghis, et autres
Publié: (2025)
Documents similaires
-
Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
par: Adib, Shefayat E Shams, et autres
Publié: (2026) -
LinguIUTics at PsyDefDetect: Iterative Imbalance-Aware Fine-tuning of Qwen3-8B for Psychological Defense Mechanism Classification
par: Adib, Shefayat E Shams, et autres
Publié: (2026) -
Addressing Data Scarcity in Bangla Fake News Detection: An LLM-Based Dataset Augmentation Approach
par: Sani, Ahmed Alfey, et autres
Publié: (2026) -
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
par: Rafid, Ahmed, et autres
Publié: (2026) -
Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
par: Tashdeed, Ishmam, et autres
Publié: (2025)