Detection Without Correction: A Robust Asymmetry in Activation-Based Hallucination Probing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Roy, Dip, Misra, Rajiv, Singh, Sanjay Kumar, Roy, Anisha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
von: Roy, Dip, et al.
Veröffentlicht: (2025)
von: Roy, Dip, et al.
Veröffentlicht: (2025)
Posterior-Calibrated Causal Circuits in Variational Autoencoders: Why Image-Domain Interpretability Fails on Tabular Data
von: Roy, Dip, et al.
Veröffentlicht: (2026)
von: Roy, Dip, et al.
Veröffentlicht: (2026)
Fundamental Limits of Neural Network Sparsification: Evidence from Catastrophic Interpretability Collapse
von: Roy, Dip, et al.
Veröffentlicht: (2026)
von: Roy, Dip, et al.
Veröffentlicht: (2026)
MemGuard-Alpha: Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting via Membership Inference and Cross-Model Disagreement
von: Roy, Anisha, et al.
Veröffentlicht: (2026)
von: Roy, Anisha, et al.
Veröffentlicht: (2026)
Hallucination Detection via Activations of Open-Weight Proxy Analyzers
von: Singh, Akshita, et al.
Veröffentlicht: (2026)
von: Singh, Akshita, et al.
Veröffentlicht: (2026)
Bayesian Autoencoder for Medical Anomaly Detection: Uncertainty-Aware Approach for Brain 2 MRI Analysis
von: Roy, Dip
Veröffentlicht: (2025)
von: Roy, Dip
Veröffentlicht: (2025)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
von: Kossen, Jannik, et al.
Veröffentlicht: (2024)
von: Kossen, Jannik, et al.
Veröffentlicht: (2024)
Temporal Graph Network: Hallucination Detection in Multi-Turn Conversation
von: Rathore, Vidhi, et al.
Veröffentlicht: (2026)
von: Rathore, Vidhi, et al.
Veröffentlicht: (2026)
Detecting Token-Level Hallucinations Using Variance Signals: A Reference-Free Approach
von: Kumar, Keshav
Veröffentlicht: (2025)
von: Kumar, Keshav
Veröffentlicht: (2025)
GPT-3 Powered Information Extraction for Building Robust Knowledge Bases
von: Choudhury, Ritabrata Roy, et al.
Veröffentlicht: (2024)
von: Choudhury, Ritabrata Roy, et al.
Veröffentlicht: (2024)
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
von: Zhou, Hanhan, et al.
Veröffentlicht: (2026)
von: Zhou, Hanhan, et al.
Veröffentlicht: (2026)
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
von: Saha, Anisha, et al.
Veröffentlicht: (2025)
von: Saha, Anisha, et al.
Veröffentlicht: (2025)
Temporal Alignment of Time Sensitive Facts with Activation Engineering
von: Govindan, Sanjay, et al.
Veröffentlicht: (2025)
von: Govindan, Sanjay, et al.
Veröffentlicht: (2025)
Hallucination is Inevitable: An Innate Limitation of Large Language Models
von: Xu, Ziwei, et al.
Veröffentlicht: (2024)
von: Xu, Ziwei, et al.
Veröffentlicht: (2024)
Supporting Assessment of Novelty of Design Problems Using Concept of Problem SAPPhIRE
von: Singh, Sanjay, et al.
Veröffentlicht: (2024)
von: Singh, Sanjay, et al.
Veröffentlicht: (2024)
Security Without Detection: Economic Denial as a Primitive for Edge and IoT Defense
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2025)
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2025)
Generalization Gaps in Political Fake News Detection: An Empirical Study on the LIAR Dataset
von: Hasan, S Mahmudul, et al.
Veröffentlicht: (2025)
von: Hasan, S Mahmudul, et al.
Veröffentlicht: (2025)
Prompt-Based Bias Calibration for Better Zero/Few-Shot Learning of Language Models
von: He, Kang, et al.
Veröffentlicht: (2024)
von: He, Kang, et al.
Veröffentlicht: (2024)
A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages
von: Das, Susmita, et al.
Veröffentlicht: (2024)
von: Das, Susmita, et al.
Veröffentlicht: (2024)
Detecting AI Hallucinations in Finance: An Information-Theoretic Method Cuts Hallucination Rate by 92%
von: Singha, Mainak
Veröffentlicht: (2025)
von: Singha, Mainak
Veröffentlicht: (2025)
HU at SemEval-2024 Task 8A: Can Contrastive Learning Learn Embeddings to Detect Machine-Generated Text?
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2024)
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2024)
HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection
von: Du, Xuefeng, et al.
Veröffentlicht: (2024)
von: Du, Xuefeng, et al.
Veröffentlicht: (2024)
Leveraging Graph Structures to Detect Hallucinations in Large Language Models
von: Nonkes, Noa, et al.
Veröffentlicht: (2024)
von: Nonkes, Noa, et al.
Veröffentlicht: (2024)
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
von: Mitsuzawa, Kensuke, et al.
Veröffentlicht: (2025)
von: Mitsuzawa, Kensuke, et al.
Veröffentlicht: (2025)
Wisdom of the Crowds in Forecasting: Forecast Summarization for Supporting Future Event Prediction
von: Saha, Anisha, et al.
Veröffentlicht: (2025)
von: Saha, Anisha, et al.
Veröffentlicht: (2025)
ZClip: Adaptive Spike Mitigation for LLM Pre-Training
von: Kumar, Abhay, et al.
Veröffentlicht: (2025)
von: Kumar, Abhay, et al.
Veröffentlicht: (2025)
Variance Control via Weight Rescaling in LLM Pre-training
von: Owen, Louis, et al.
Veröffentlicht: (2025)
von: Owen, Louis, et al.
Veröffentlicht: (2025)
SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution
von: He, Kang, et al.
Veröffentlicht: (2026)
von: He, Kang, et al.
Veröffentlicht: (2026)
TraceNAS: Zero-shot LLM Pruning via Gradient Trace Correlation
von: Malettira, Prajna G., et al.
Veröffentlicht: (2026)
von: Malettira, Prajna G., et al.
Veröffentlicht: (2026)
Lost in State Space: Probing Frozen Mamba Representations
von: Wagh, Bhagyashree, et al.
Veröffentlicht: (2026)
von: Wagh, Bhagyashree, et al.
Veröffentlicht: (2026)
Detecting and Preventing Hallucinations in Large Vision Language Models
von: Gunjal, Anisha, et al.
Veröffentlicht: (2023)
von: Gunjal, Anisha, et al.
Veröffentlicht: (2023)
RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration
von: Ridder, Fabian, et al.
Veröffentlicht: (2026)
von: Ridder, Fabian, et al.
Veröffentlicht: (2026)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
von: Binkowski, Jakub, et al.
Veröffentlicht: (2026)
von: Binkowski, Jakub, et al.
Veröffentlicht: (2026)
InterrogateLLM: Zero-Resource Hallucination Detection in LLM-Generated Answers
von: Yehuda, Yakir, et al.
Veröffentlicht: (2024)
von: Yehuda, Yakir, et al.
Veröffentlicht: (2024)
Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection
von: Li, Jerry, et al.
Veröffentlicht: (2025)
von: Li, Jerry, et al.
Veröffentlicht: (2025)
A Unified Definition of Hallucination: It's The World Model, Stupid!
von: Liu, Emmy, et al.
Veröffentlicht: (2025)
von: Liu, Emmy, et al.
Veröffentlicht: (2025)
A Robust Autoencoder Ensemble-Based Approach for Anomaly Detection in Text
von: Pantin, Jeremie, et al.
Veröffentlicht: (2024)
von: Pantin, Jeremie, et al.
Veröffentlicht: (2024)
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
von: Liu, Emmy, et al.
Veröffentlicht: (2026)
von: Liu, Emmy, et al.
Veröffentlicht: (2026)
Why Fine-Tuning Encourages Hallucinations and How to Fix It
von: Kaplan, Guy, et al.
Veröffentlicht: (2026)
von: Kaplan, Guy, et al.
Veröffentlicht: (2026)
Learning to Reason for Hallucination Span Detection
von: Su, Hsuan, et al.
Veröffentlicht: (2025)
von: Su, Hsuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
von: Roy, Dip, et al.
Veröffentlicht: (2025) -
Posterior-Calibrated Causal Circuits in Variational Autoencoders: Why Image-Domain Interpretability Fails on Tabular Data
von: Roy, Dip, et al.
Veröffentlicht: (2026) -
Fundamental Limits of Neural Network Sparsification: Evidence from Catastrophic Interpretability Collapse
von: Roy, Dip, et al.
Veröffentlicht: (2026) -
MemGuard-Alpha: Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting via Membership Inference and Cross-Model Disagreement
von: Roy, Anisha, et al.
Veröffentlicht: (2026) -
Hallucination Detection via Activations of Open-Weight Proxy Analyzers
von: Singh, Akshita, et al.
Veröffentlicht: (2026)