Hallucination Detection in Large Language Models with Metamorphic Relations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Borui, Mamun, Md Afif Al, Zhang, Jie M., Uddin, Gias |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large Language Models in Software Documentation and Modeling: A Literature Review and Findings
von: Radosky, Lukas, et al.
Veröffentlicht: (2026)
von: Radosky, Lukas, et al.
Veröffentlicht: (2026)
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
von: Nwokocha, Caleb Princewill
Veröffentlicht: (2022)
von: Nwokocha, Caleb Princewill
Veröffentlicht: (2022)
Supporting software engineering tasks with agentic AI: Demonstration on document retrieval and test scenario generation
von: Kica, Marian, et al.
Veröffentlicht: (2026)
von: Kica, Marian, et al.
Veröffentlicht: (2026)
Judgment2vec: Apply Graph Analytics to Searching and Recommendation of Similar Judgments
von: Shao, Hsuan-Lei
Veröffentlicht: (2024)
von: Shao, Hsuan-Lei
Veröffentlicht: (2024)
A transfer learning approach for automatic conflicts detection in software requirement sentence pairs based on dual encoders
von: Wang, Yizheng, et al.
Veröffentlicht: (2025)
von: Wang, Yizheng, et al.
Veröffentlicht: (2025)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
von: Wang, Fali, et al.
Veröffentlicht: (2025)
von: Wang, Fali, et al.
Veröffentlicht: (2025)
Linguistic Collapse: Neural Collapse in (Large) Language Models
von: Wu, Robert, et al.
Veröffentlicht: (2024)
von: Wu, Robert, et al.
Veröffentlicht: (2024)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
von: Abhishek, Alok, et al.
Veröffentlicht: (2026)
von: Abhishek, Alok, et al.
Veröffentlicht: (2026)
SocialX: A Modular Platform for Multi-Source Big Data Research in Indonesia
von: Saputra, Muhammad Apriandito Arya, et al.
Veröffentlicht: (2026)
von: Saputra, Muhammad Apriandito Arya, et al.
Veröffentlicht: (2026)
Large Language Models are Inconsistent and Biased Evaluators
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
Generative AI Models: Opportunities and Risks for Industry and Authorities
von: Alt, Tobias, et al.
Veröffentlicht: (2024)
von: Alt, Tobias, et al.
Veröffentlicht: (2024)
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
von: Gupta, Aayush
Veröffentlicht: (2025)
von: Gupta, Aayush
Veröffentlicht: (2025)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
von: Radosky, Lukas, et al.
Veröffentlicht: (2026)
von: Radosky, Lukas, et al.
Veröffentlicht: (2026)
See-Saw Generative Mechanism for Scalable Recursive Code Generation with Generative AI
von: Vsevolodovna, Ruslan Idelfonso Magaña
Veröffentlicht: (2024)
von: Vsevolodovna, Ruslan Idelfonso Magaña
Veröffentlicht: (2024)
TRAWL: Tensor Reduced and Approximated Weights for Large Language Models
von: Luo, Yiran, et al.
Veröffentlicht: (2024)
von: Luo, Yiran, et al.
Veröffentlicht: (2024)
Data and AI governance: Promoting equity, ethics, and fairness in large language models
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
Do Reasoning Models Enhance Embedding Models?
von: Chan, Wun Yu, et al.
Veröffentlicht: (2026)
von: Chan, Wun Yu, et al.
Veröffentlicht: (2026)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
von: Kuz, Mykola, et al.
Veröffentlicht: (2025)
von: Kuz, Mykola, et al.
Veröffentlicht: (2025)
Reasoning Promotes Robustness in Theory of Mind Tasks
von: de Haan, Ian B., et al.
Veröffentlicht: (2026)
von: de Haan, Ian B., et al.
Veröffentlicht: (2026)
A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
von: Wang, Fali, et al.
Veröffentlicht: (2024)
von: Wang, Fali, et al.
Veröffentlicht: (2024)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
von: Lei, Xiang, et al.
Veröffentlicht: (2025)
von: Lei, Xiang, et al.
Veröffentlicht: (2025)
Make Literature-Based Discovery Great Again through Reproducible Pipelines
von: Cestnik, Bojan, et al.
Veröffentlicht: (2025)
von: Cestnik, Bojan, et al.
Veröffentlicht: (2025)
Semantic Retention and Extreme Compression in LLMs: Can We Have Both?
von: Laborde, Stanislas, et al.
Veröffentlicht: (2025)
von: Laborde, Stanislas, et al.
Veröffentlicht: (2025)
Text Clustering with Large Language Model Embeddings
von: Petukhova, Alina, et al.
Veröffentlicht: (2024)
von: Petukhova, Alina, et al.
Veröffentlicht: (2024)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025)
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025)
DeepSeek-V3, GPT-4, Phi-4, and LLaMA-3.3 generate correct code for LoRaWAN-related engineering tasks
von: Fernandes, Daniel, et al.
Veröffentlicht: (2025)
von: Fernandes, Daniel, et al.
Veröffentlicht: (2025)
A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy Bangla Texts
von: Elahi, Kazi Toufique, et al.
Veröffentlicht: (2024)
von: Elahi, Kazi Toufique, et al.
Veröffentlicht: (2024)
Uncovering Uncertainty in Transformer Inference
von: Brothers, Greyson, et al.
Veröffentlicht: (2024)
von: Brothers, Greyson, et al.
Veröffentlicht: (2024)
Model Surgery: Modulating LLM's Behavior Via Simple Parameter Editing
von: Wang, Huanqian, et al.
Veröffentlicht: (2024)
von: Wang, Huanqian, et al.
Veröffentlicht: (2024)
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
Quantum NLP models on Natural Language Inference
von: Sun, Ling, et al.
Veröffentlicht: (2025)
von: Sun, Ling, et al.
Veröffentlicht: (2025)
DRS-OSS: Practical Diff Risk Scoring with LLMs
von: Sayedsalehi, Ali, et al.
Veröffentlicht: (2025)
von: Sayedsalehi, Ali, et al.
Veröffentlicht: (2025)
Kodezi Chronos: A Debugging-First Language Model for Repository-Scale Code Understanding
von: Khan, Ishraq, et al.
Veröffentlicht: (2025)
von: Khan, Ishraq, et al.
Veröffentlicht: (2025)
Investigating Neuron Ablation in Attention Heads: The Case for Peak Activation Centering
von: Pochinkov, Nicholas, et al.
Veröffentlicht: (2024)
von: Pochinkov, Nicholas, et al.
Veröffentlicht: (2024)
Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2025)
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2025)
Correctness is not Faithfulness in RAG Attributions
von: Wallat, Jonas, et al.
Veröffentlicht: (2024)
von: Wallat, Jonas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Large Language Models in Software Documentation and Modeling: A Literature Review and Findings
von: Radosky, Lukas, et al.
Veröffentlicht: (2026) -
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
von: Nwokocha, Caleb Princewill
Veröffentlicht: (2022) -
Supporting software engineering tasks with agentic AI: Demonstration on document retrieval and test scenario generation
von: Kica, Marian, et al.
Veröffentlicht: (2026) -
Judgment2vec: Apply Graph Analytics to Searching and Recommendation of Similar Judgments
von: Shao, Hsuan-Lei
Veröffentlicht: (2024) -
A transfer learning approach for automatic conflicts detection in software requirement sentence pairs based on dual encoders
von: Wang, Yizheng, et al.
Veröffentlicht: (2025)