Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abdaljalil, Samir, Sharma, Parichit, Serpedin, Erchin, Kurban, Hasan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SINdex: Semantic INconsistency Index for Hallucination Detection in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Audit-of-Understanding: Posterior-Constrained Inference for Mathematical Reasoning in Language Models
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation
von: Mahmood, Alhasan, et al.
Veröffentlicht: (2026)
von: Mahmood, Alhasan, et al.
Veröffentlicht: (2026)
4D Synchronized Fields: Motion-Language Gaussian Splatting for Temporal Scene Understanding
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2026)
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2026)
Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning
von: Polat, Can, et al.
Veröffentlicht: (2025)
von: Polat, Can, et al.
Veröffentlicht: (2025)
QuantumCanvas: A Multimodal Benchmark for Visual Learning of Atomic Interactions
von: Polat, Can, et al.
Veröffentlicht: (2025)
von: Polat, Can, et al.
Veröffentlicht: (2025)
SCALAR: Quantifying Structural Hallucination, Consistency, and Reasoning Gaps in Materials Foundation Models
von: Polat, Can, et al.
Veröffentlicht: (2026)
von: Polat, Can, et al.
Veröffentlicht: (2026)
EMPATHIA: Multi-Faceted Human-AI Collaboration for Refugee Integration
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2025)
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2025)
C2NP: A Benchmark for Learning Scale-Dependent Geometric Invariances in 3D Materials Generation
von: Polat, Can, et al.
Veröffentlicht: (2026)
von: Polat, Can, et al.
Veröffentlicht: (2026)
A multilingual hallucination benchmark: MultiWikiQHalluA
von: Thoresen, Freja, et al.
Veröffentlicht: (2026)
von: Thoresen, Freja, et al.
Veröffentlicht: (2026)
Probabilistic distances-based hallucination detection in LLMs with RAG
von: Oblovatny, Rodion, et al.
Veröffentlicht: (2025)
von: Oblovatny, Rodion, et al.
Veröffentlicht: (2025)
xChemAgents: Agentic AI for Explainable Quantum Chemistry
von: Polat, Can, et al.
Veröffentlicht: (2025)
von: Polat, Can, et al.
Veröffentlicht: (2025)
Beyond Atomic Geometry Representations in Materials Science: A Human-in-the-Loop Multimodal Framework
von: Polat, Can, et al.
Veröffentlicht: (2025)
von: Polat, Can, et al.
Veröffentlicht: (2025)
How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science
von: Polat, Can, et al.
Veröffentlicht: (2026)
von: Polat, Can, et al.
Veröffentlicht: (2026)
Understanding the Capabilities of Molecular Graph Neural Networks in Materials Science Through Multimodal Learning and Physical Context Encoding
von: Polat, Can, et al.
Veröffentlicht: (2025)
von: Polat, Can, et al.
Veröffentlicht: (2025)
A novel hallucination classification framework
von: Zavhorodnii, Maksym, et al.
Veröffentlicht: (2025)
von: Zavhorodnii, Maksym, et al.
Veröffentlicht: (2025)
IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video
von: Khanbayov, Rasul, et al.
Veröffentlicht: (2026)
von: Khanbayov, Rasul, et al.
Veröffentlicht: (2026)
Scalable multilingual PII annotation for responsible AI in LLMs
von: Meena, Bharti, et al.
Veröffentlicht: (2025)
von: Meena, Bharti, et al.
Veröffentlicht: (2025)
Beyond Textual Context: Structural Graph Encoding with Adaptive Space Alignment to alleviate the hallucination of LLMs
von: Zhang, Yifang, et al.
Veröffentlicht: (2025)
von: Zhang, Yifang, et al.
Veröffentlicht: (2025)
Joint Sensor Deployment and Physics-Informed Graph Transformer for Smart Grid Attack Detection
von: Elnour, Mariam, et al.
Veröffentlicht: (2026)
von: Elnour, Mariam, et al.
Veröffentlicht: (2026)
Do LLM hallucination detectors suffer from low-resource effect?
von: Datta, Debtanu, et al.
Veröffentlicht: (2026)
von: Datta, Debtanu, et al.
Veröffentlicht: (2026)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
NEU-ESC: A Comprehensive Vietnamese dataset for Educational Sentiment analysis and topic Classification toward multitask learning
von: Mai, Phan Quoc Hung, et al.
Veröffentlicht: (2025)
von: Mai, Phan Quoc Hung, et al.
Veröffentlicht: (2025)
Ask-EDA: A Design Assistant Empowered by LLM, Hybrid RAG and Abbreviation De-hallucination
von: Shi, Luyao, et al.
Veröffentlicht: (2024)
von: Shi, Luyao, et al.
Veröffentlicht: (2024)
Are LLMs Ready to Replace Bangla Annotators?
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2026)
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2026)
The impact of fine tuning in LLaMA on hallucinations for named entity extraction in legal documentation
von: Vargas, Francisco, et al.
Veröffentlicht: (2025)
von: Vargas, Francisco, et al.
Veröffentlicht: (2025)
ACL-Verbatim: hallucination-free question answering for research
von: Recski, Gábor, et al.
Veröffentlicht: (2026)
von: Recski, Gábor, et al.
Veröffentlicht: (2026)
Retrieval-augmented generation in multilingual settings
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
A comprehensive taxonomy of hallucinations in Large Language Models
von: Cossio, Manuel
Veröffentlicht: (2025)
von: Cossio, Manuel
Veröffentlicht: (2025)
A multimodal multiplex of the mental lexicon for multilingual individuals
von: Huynh, Maria, et al.
Veröffentlicht: (2025)
von: Huynh, Maria, et al.
Veröffentlicht: (2025)
SLPL SHROOM at SemEval2024 Task 06: A comprehensive study on models ability to detect hallucination
von: Fallah, Pouya, et al.
Veröffentlicht: (2024)
von: Fallah, Pouya, et al.
Veröffentlicht: (2024)
Towards a resource for multilingual lexicons: an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation
von: Han, Lifeng, et al.
Veröffentlicht: (2020)
von: Han, Lifeng, et al.
Veröffentlicht: (2020)
Strong hallucinations from negation and how to fix them
von: Asher, Nicholas, et al.
Veröffentlicht: (2024)
von: Asher, Nicholas, et al.
Veröffentlicht: (2024)
Disentangling concept semantics via multilingual averaging in Sparse Autoencoders
von: O'Reilly, Cliff, et al.
Veröffentlicht: (2025)
von: O'Reilly, Cliff, et al.
Veröffentlicht: (2025)
Suvach -- Generated Hindi QA benchmark
von: Narayanan, Vaishak, et al.
Veröffentlicht: (2024)
von: Narayanan, Vaishak, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SINdex: Semantic INconsistency Index for Hallucination Detection in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025) -
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2026) -
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025) -
Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025) -
Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)