Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates
Fuente:
arXiv
Saved in:
| Main Author: | Scrivens, Arsenios |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Information-Theoretic Limits of Safety Verification for Self-Improving Systems
by: Scrivens, Arsenios
Published: (2026)
by: Scrivens, Arsenios
Published: (2026)
Holographic Invariant Storage: Design-Time Safety Contracts via Vector Symbolic Architectures
by: Scrivens, Arsenios
Published: (2026)
by: Scrivens, Arsenios
Published: (2026)
NeuroAI for AI Safety
by: Mineault, Patrick, et al.
Published: (2024)
by: Mineault, Patrick, et al.
Published: (2024)
Bridging Efficiency and Safety: Formal Verification of Neural Networks with Early Exits
by: Elboher, Yizhak Yisrael, et al.
Published: (2025)
by: Elboher, Yizhak Yisrael, et al.
Published: (2025)
TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift
by: Mahato, Dipesh Tharu, et al.
Published: (2025)
by: Mahato, Dipesh Tharu, et al.
Published: (2025)
Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce Grokking
by: Lyu, Kaifeng, et al.
Published: (2023)
by: Lyu, Kaifeng, et al.
Published: (2023)
Divergence of Empirical Neural Tangent Kernel in Classification Problems
by: Yu, Zixiong, et al.
Published: (2025)
by: Yu, Zixiong, et al.
Published: (2025)
Bridging Neural ODE and ResNet: A Formal Error Bound for Safety Verification
by: Sayed, Abdelrahman Sayed, et al.
Published: (2025)
by: Sayed, Abdelrahman Sayed, et al.
Published: (2025)
Shutdown Safety Valves for Advanced AI
by: Conitzer, Vincent
Published: (2026)
by: Conitzer, Vincent
Published: (2026)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
Holistic Safety and Responsibility Evaluations of Advanced AI Models
by: Weidinger, Laura, et al.
Published: (2024)
by: Weidinger, Laura, et al.
Published: (2024)
Self-Improving Safety Performance of Reinforcement Learning Based Driving with Black-Box Verification Algorithms
by: Dagdanov, Resul, et al.
Published: (2022)
by: Dagdanov, Resul, et al.
Published: (2022)
Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability
by: Liu, Yushen, et al.
Published: (2026)
by: Liu, Yushen, et al.
Published: (2026)
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
by: Liu, Chenruo, et al.
Published: (2025)
by: Liu, Chenruo, et al.
Published: (2025)
Computational Safety for Generative AI: A Signal Processing Perspective
by: Chen, Pin-Yu
Published: (2025)
by: Chen, Pin-Yu
Published: (2025)
Effective Controllable Bias Mitigation for Classification and Retrieval using Gate Adapters
by: Masoudian, Shahed, et al.
Published: (2024)
by: Masoudian, Shahed, et al.
Published: (2024)
Spectral Gating Networks
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
Does Faithfulness Conflict with Plausibility? An Empirical Study in Explainable AI across NLP Tasks
by: Lu, Xiaolei, et al.
Published: (2024)
by: Lu, Xiaolei, et al.
Published: (2024)
Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2
by: Martra, Pere
Published: (2025)
by: Martra, Pere
Published: (2025)
NASP-T: A Fuzzy Neuro-Symbolic Transformer for Logic-Constrained Aviation Safety Report Classification
by: Machot, Fadi Al, et al.
Published: (2025)
by: Machot, Fadi Al, et al.
Published: (2025)
A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety
by: Lee, Hyunin, et al.
Published: (2024)
by: Lee, Hyunin, et al.
Published: (2024)
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
by: Overman, William, et al.
Published: (2025)
by: Overman, William, et al.
Published: (2025)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
DB-FGA-Net: Dual Backbone Frequency Gated Attention Network for Multi-Class Brain Tumor Classification with Grad-CAM Interpretability
by: Shreya, Saraf Anzum, et al.
Published: (2025)
by: Shreya, Saraf Anzum, et al.
Published: (2025)
International AI Safety Report
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
Domain Gating Ensemble Networks for AI-Generated Text Detection
by: Tripathi, Arihant, et al.
Published: (2025)
by: Tripathi, Arihant, et al.
Published: (2025)
xEEGNet: Towards Explainable AI in EEG Dementia Classification
by: Zanola, Andrea, et al.
Published: (2025)
by: Zanola, Andrea, et al.
Published: (2025)
Lightweight Safety Classification Using Pruned Language Models
by: Sawtell, Mason, et al.
Published: (2024)
by: Sawtell, Mason, et al.
Published: (2024)
Green MLOps to Green GenOps: An Empirical Study of Energy Consumption in Discriminative and Generative AI Operations
by: Sánchez-Mompó, Adrián, et al.
Published: (2025)
by: Sánchez-Mompó, Adrián, et al.
Published: (2025)
Rethinking Gating Mechanism in Sparse MoE: Handling Arbitrary Modality Inputs with Confidence-Guided Gate
by: Zheng, Liangwei Nathan, et al.
Published: (2025)
by: Zheng, Liangwei Nathan, et al.
Published: (2025)
Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
by: Li, Guanlin, et al.
Published: (2025)
by: Li, Guanlin, et al.
Published: (2025)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Spatially-Delineated Domain-Adapted AI Classification: An Application for Oncology Data
by: Farhadloo, Majid, et al.
Published: (2025)
by: Farhadloo, Majid, et al.
Published: (2025)
DynamicGate MLP Conditional Computation via Learned Structural Dropout and Input Dependent Gating for Functional Plasticity
by: Choi, Yong Il
Published: (2026)
by: Choi, Yong Il
Published: (2026)
Generative AI Agents in Autonomous Machines: A Safety Perspective
by: Jabbour, Jason, et al.
Published: (2024)
by: Jabbour, Jason, et al.
Published: (2024)
Automated Conjecture Resolution with Formal Verification
by: Ju, Haocheng, et al.
Published: (2026)
by: Ju, Haocheng, et al.
Published: (2026)
Neural Network Verification with PyRAT
by: Lemesle, Augustin, et al.
Published: (2024)
by: Lemesle, Augustin, et al.
Published: (2024)
From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems
by: Stefani, Thomas, et al.
Published: (2026)
by: Stefani, Thomas, et al.
Published: (2026)
Fitting Multilinear Polynomials for Logic Gate Networks
by: Kim, Youngsung
Published: (2026)
by: Kim, Youngsung
Published: (2026)
Uncertainty Estimation using Variance-Gated Distributions
by: Gillis, H. Martin, et al.
Published: (2025)
by: Gillis, H. Martin, et al.
Published: (2025)
Similar Items
-
Information-Theoretic Limits of Safety Verification for Self-Improving Systems
by: Scrivens, Arsenios
Published: (2026) -
Holographic Invariant Storage: Design-Time Safety Contracts via Vector Symbolic Architectures
by: Scrivens, Arsenios
Published: (2026) -
NeuroAI for AI Safety
by: Mineault, Patrick, et al.
Published: (2024) -
Bridging Efficiency and Safety: Formal Verification of Neural Networks with Early Exits
by: Elboher, Yizhak Yisrael, et al.
Published: (2025) -
TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift
by: Mahato, Dipesh Tharu, et al.
Published: (2025)