Squish and Release: Exposing Hidden Hallucinations by Making Them Surface as Safety Signals
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Oh, Nathaniel, Attie, Paul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
First Hallucination Tokens Are Different from Conditional Ones
von: Snel, Jakob, et al.
Veröffentlicht: (2025)
von: Snel, Jakob, et al.
Veröffentlicht: (2025)
Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals
von: Susanto, Lucky, et al.
Veröffentlicht: (2025)
von: Susanto, Lucky, et al.
Veröffentlicht: (2025)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
von: Foerster, Hanna, et al.
Veröffentlicht: (2025)
von: Foerster, Hanna, et al.
Veröffentlicht: (2025)
Fantastic Pretraining Optimizers and Where to Find Them
von: Wen, Kaiyue, et al.
Veröffentlicht: (2025)
von: Wen, Kaiyue, et al.
Veröffentlicht: (2025)
Low Rank Gradients and Where to Find Them
von: Sonthalia, Rishi, et al.
Veröffentlicht: (2025)
von: Sonthalia, Rishi, et al.
Veröffentlicht: (2025)
Computational Safety for Generative AI: A Signal Processing Perspective
von: Chen, Pin-Yu
Veröffentlicht: (2025)
von: Chen, Pin-Yu
Veröffentlicht: (2025)
Making AI-Assisted Grant Evaluation Auditable without Exposing the Model
von: Bicakci, Kemal
Veröffentlicht: (2026)
von: Bicakci, Kemal
Veröffentlicht: (2026)
GNN Explanations that do not Explain and How to find Them
von: Azzolin, Steve, et al.
Veröffentlicht: (2026)
von: Azzolin, Steve, et al.
Veröffentlicht: (2026)
What Cohort INRs Encode and Where to Freeze Them
von: Sideri-Lampretsa, Vasiliki, et al.
Veröffentlicht: (2026)
von: Sideri-Lampretsa, Vasiliki, et al.
Veröffentlicht: (2026)
One Filter to Deploy Them All: Robust Safety for Quadrupedal Navigation in Unknown Environments
von: Lin, Albert, et al.
Veröffentlicht: (2024)
von: Lin, Albert, et al.
Veröffentlicht: (2024)
Sea-cret Agents: Maritime Abduction for Region Generation to Expose Dark Vessel Trajectories
von: Bavikadi, Divyagna, et al.
Veröffentlicht: (2025)
von: Bavikadi, Divyagna, et al.
Veröffentlicht: (2025)
Explaining Predictive Uncertainty by Exposing Second-Order Effects
von: Bley, Florian, et al.
Veröffentlicht: (2024)
von: Bley, Florian, et al.
Veröffentlicht: (2024)
How to Square Tensor Networks and Circuits Without Squaring Them
von: Loconte, Lorenzo, et al.
Veröffentlicht: (2025)
von: Loconte, Lorenzo, et al.
Veröffentlicht: (2025)
Transcendence: Generative Models Can Outperform The Experts That Train Them
von: Zhang, Edwin, et al.
Veröffentlicht: (2024)
von: Zhang, Edwin, et al.
Veröffentlicht: (2024)
Calibrated Language Models and How to Find Them with Label Smoothing
von: Huang, Jerry, et al.
Veröffentlicht: (2025)
von: Huang, Jerry, et al.
Veröffentlicht: (2025)
Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal
von: Yuan, Aojie, et al.
Veröffentlicht: (2026)
von: Yuan, Aojie, et al.
Veröffentlicht: (2026)
REBEL: Hidden Knowledge Recovery via Evolutionary-Based Evaluation Loop
von: Rybak, Patryk, et al.
Veröffentlicht: (2026)
von: Rybak, Patryk, et al.
Veröffentlicht: (2026)
When Privacy Isn't Synthetic: Hidden Data Leakage in Generative AI Models
von: Mustaqim, S. M., et al.
Veröffentlicht: (2025)
von: Mustaqim, S. M., et al.
Veröffentlicht: (2025)
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
One Wave To Explain Them All: A Unifying Perspective On Feature Attribution
von: Kasmi, Gabriel, et al.
Veröffentlicht: (2024)
von: Kasmi, Gabriel, et al.
Veröffentlicht: (2024)
Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)
von: Prinster, Drew, et al.
Veröffentlicht: (2024)
von: Prinster, Drew, et al.
Veröffentlicht: (2024)
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
von: Jin, Jiahe, et al.
Veröffentlicht: (2025)
von: Jin, Jiahe, et al.
Veröffentlicht: (2025)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
von: Huang, Audrey, et al.
Veröffentlicht: (2025)
von: Huang, Audrey, et al.
Veröffentlicht: (2025)
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
What LLMs Think When You Don't Tell Them What to Think About?
von: Kwon, Yongchan, et al.
Veröffentlicht: (2026)
von: Kwon, Yongchan, et al.
Veröffentlicht: (2026)
On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm Perspective
von: Xie, Zeke, et al.
Veröffentlicht: (2020)
von: Xie, Zeke, et al.
Veröffentlicht: (2020)
When Stability Fails: Hidden Failure Modes Of LLMS in Data-Constrained Scientific Decision-Making
von: Riasat, Nazia
Veröffentlicht: (2026)
von: Riasat, Nazia
Veröffentlicht: (2026)
Learning with Hidden Factorial Structure
von: Arnal, Charles, et al.
Veröffentlicht: (2024)
von: Arnal, Charles, et al.
Veröffentlicht: (2024)
Toward a Metrology for Artificial Intelligence: Hidden-Rule Environments and Reinforcement Learning
von: Mathew, Christo, et al.
Veröffentlicht: (2025)
von: Mathew, Christo, et al.
Veröffentlicht: (2025)
CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right
von: Lee, Jaebok, et al.
Veröffentlicht: (2025)
von: Lee, Jaebok, et al.
Veröffentlicht: (2025)
The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
von: Lan, Yifan, et al.
Veröffentlicht: (2026)
von: Lan, Yifan, et al.
Veröffentlicht: (2026)
StyleShield: Exposing the Fragility of AIGC Detectors through Continuous Controllable Style Transfer
von: Zheng, Guantian
Veröffentlicht: (2026)
von: Zheng, Guantian
Veröffentlicht: (2026)
Trained Models Tell Us How to Make Them Robust to Spurious Correlation without Group Annotation
von: Ghaznavi, Mahdi, et al.
Veröffentlicht: (2024)
von: Ghaznavi, Mahdi, et al.
Veröffentlicht: (2024)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
Exposing propaganda: an analysis of stylistic cues comparing human annotations and machine classification
von: Faye, Géraud, et al.
Veröffentlicht: (2024)
von: Faye, Géraud, et al.
Veröffentlicht: (2024)
Rank-1 LoRAs Encode Interpretable Reasoning Signals
von: Ward, Jake, et al.
Veröffentlicht: (2025)
von: Ward, Jake, et al.
Veröffentlicht: (2025)
ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language Models
von: Oh, Jio, et al.
Veröffentlicht: (2024)
von: Oh, Jio, et al.
Veröffentlicht: (2024)
Translating Subgraphs to Nodes Makes Simple GNNs Strong and Efficient for Subgraph Representation Learning
von: Kim, Dongkwan, et al.
Veröffentlicht: (2022)
von: Kim, Dongkwan, et al.
Veröffentlicht: (2022)
Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2026)
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
First Hallucination Tokens Are Different from Conditional Ones
von: Snel, Jakob, et al.
Veröffentlicht: (2025) -
Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals
von: Susanto, Lucky, et al.
Veröffentlicht: (2025) -
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
von: Foerster, Hanna, et al.
Veröffentlicht: (2025) -
Fantastic Pretraining Optimizers and Where to Find Them
von: Wen, Kaiyue, et al.
Veröffentlicht: (2025) -
Low Rank Gradients and Where to Find Them
von: Sonthalia, Rishi, et al.
Veröffentlicht: (2025)