Information-Theoretic Limits of Safety Verification for Self-Improving Systems
Fuente:
arXiv
Saved in:
| Main Author: | Scrivens, Arsenios |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates
by: Scrivens, Arsenios
Published: (2026)
by: Scrivens, Arsenios
Published: (2026)
Holographic Invariant Storage: Design-Time Safety Contracts via Vector Symbolic Architectures
by: Scrivens, Arsenios
Published: (2026)
by: Scrivens, Arsenios
Published: (2026)
Self-Improving Safety Performance of Reinforcement Learning Based Driving with Black-Box Verification Algorithms
by: Dagdanov, Resul, et al.
Published: (2022)
by: Dagdanov, Resul, et al.
Published: (2022)
On The Statistical Limits of Self-Improving Agents
by: Wang, Charles L., et al.
Published: (2025)
by: Wang, Charles L., et al.
Published: (2025)
Information-Theoretic Dual Memory System for Continual Learning
by: Wu, RunQing, et al.
Published: (2025)
by: Wu, RunQing, et al.
Published: (2025)
On the Theoretical Limitations of Embedding-based Link Prediction
by: Badreddine, Samy, et al.
Published: (2025)
by: Badreddine, Samy, et al.
Published: (2025)
Harnessing Neuron Stability to Improve DNN Verification
by: Duong, Hai, et al.
Published: (2024)
by: Duong, Hai, et al.
Published: (2024)
Bridging Efficiency and Safety: Formal Verification of Neural Networks with Early Exits
by: Elboher, Yizhak Yisrael, et al.
Published: (2025)
by: Elboher, Yizhak Yisrael, et al.
Published: (2025)
TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift
by: Mahato, Dipesh Tharu, et al.
Published: (2025)
by: Mahato, Dipesh Tharu, et al.
Published: (2025)
Balancing Graph Embedding Smoothness in Self-Supervised Learning via Information-Theoretic Decomposition
by: Jung, Heesoo, et al.
Published: (2025)
by: Jung, Heesoo, et al.
Published: (2025)
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
by: Shen, Sicheng, et al.
Published: (2026)
by: Shen, Sicheng, et al.
Published: (2026)
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Bridging Neural ODE and ResNet: A Formal Error Bound for Safety Verification
by: Sayed, Abdelrahman Sayed, et al.
Published: (2025)
by: Sayed, Abdelrahman Sayed, et al.
Published: (2025)
The Geometry of Self-Verification in a Task-Specific Reasoning Model
by: Lee, Andrew, et al.
Published: (2025)
by: Lee, Andrew, et al.
Published: (2025)
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification
by: Liang, Zhenwen, et al.
Published: (2024)
by: Liang, Zhenwen, et al.
Published: (2024)
Towards Theoretical Understandings of Self-Consuming Generative Models
by: Fu, Shi, et al.
Published: (2024)
by: Fu, Shi, et al.
Published: (2024)
Reasoning on the Manifold: Bidirectional Consistency for Self-Verification in Diffusion Language Models
by: Ruan, Jiaoyang, et al.
Published: (2026)
by: Ruan, Jiaoyang, et al.
Published: (2026)
On Information-Theoretic Measures of Predictive Uncertainty
by: Schweighofer, Kajetan, et al.
Published: (2024)
by: Schweighofer, Kajetan, et al.
Published: (2024)
Information-Theoretic Safe Bayesian Optimization
by: Bottero, Alessandro G., et al.
Published: (2024)
by: Bottero, Alessandro G., et al.
Published: (2024)
Information-Theoretic Foundations for Machine Learning
by: Jeon, Hong Jun, et al.
Published: (2024)
by: Jeon, Hong Jun, et al.
Published: (2024)
An Information Theoretic Approach to Machine Unlearning
by: Foster, Jack, et al.
Published: (2024)
by: Foster, Jack, et al.
Published: (2024)
Theoretical Benefit and Limitation of Diffusion Language Model
by: Feng, Guhao, et al.
Published: (2025)
by: Feng, Guhao, et al.
Published: (2025)
SafeCoT: Improving VLM Safety with Minimal Reasoning
by: Ma, Jiachen, et al.
Published: (2025)
by: Ma, Jiachen, et al.
Published: (2025)
An Information Theoretic Evaluation Metric For Strong Unlearning
by: Jeon, Dongjae, et al.
Published: (2024)
by: Jeon, Dongjae, et al.
Published: (2024)
Information-Theoretic Foundations for Neural Scaling Laws
by: Jeon, Hong Jun, et al.
Published: (2024)
by: Jeon, Hong Jun, et al.
Published: (2024)
Physics-Informed Deep Learning for Entropy Prediction in Heterogeneous Systems: Thermodynamic and Information-Theoretic Case Studies
by: Sahoo, Biswajeet, et al.
Published: (2026)
by: Sahoo, Biswajeet, et al.
Published: (2026)
Provably Invincible Adversarial Attacks on Reinforcement Learning Systems: A Rate-Distortion Information-Theoretic Approach
by: Lu, Ziqing, et al.
Published: (2025)
by: Lu, Ziqing, et al.
Published: (2025)
Self-Trained Verification for Training- and Test-Time Self-Improvement
by: Wu, Chen Henry, et al.
Published: (2026)
by: Wu, Chen Henry, et al.
Published: (2026)
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025)
by: Zhou, Zhi, et al.
Published: (2025)
Self-Improving Robust Preference Optimization
by: Choi, Eugene, et al.
Published: (2024)
by: Choi, Eugene, et al.
Published: (2024)
What Limits Agentic Systems Efficiency?
by: Bian, Song, et al.
Published: (2025)
by: Bian, Song, et al.
Published: (2025)
On the Limits of Self-Improving in Large Language Models: The Singularity Is Not Near Without Symbolic Model Synthesis
by: Zenil, Hector
Published: (2026)
by: Zenil, Hector
Published: (2026)
Backtracking Improves Generation Safety
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
by: Li, Zhuo, et al.
Published: (2025)
by: Li, Zhuo, et al.
Published: (2025)
Measuring Leakage in Concept-Based Methods: An Information Theoretic Approach
by: Makonnen, Mikael, et al.
Published: (2025)
by: Makonnen, Mikael, et al.
Published: (2025)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
Hexa: Self-Improving for Knowledge-Grounded Dialogue System
by: Jo, Daejin, et al.
Published: (2023)
by: Jo, Daejin, et al.
Published: (2023)
Optimizing Deep Neural Networks using Safety-Guided Self Compression
by: Zbeeb, Mohammad, et al.
Published: (2025)
by: Zbeeb, Mohammad, et al.
Published: (2025)
Probabilistic Runtime Verification, Evaluation and Risk Assessment of Visual Deep Learning Systems
by: Torpmann-Hagen, Birk, et al.
Published: (2025)
by: Torpmann-Hagen, Birk, et al.
Published: (2025)
Similar Items
-
Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates
by: Scrivens, Arsenios
Published: (2026) -
Holographic Invariant Storage: Design-Time Safety Contracts via Vector Symbolic Architectures
by: Scrivens, Arsenios
Published: (2026) -
Self-Improving Safety Performance of Reinforcement Learning Based Driving with Black-Box Verification Algorithms
by: Dagdanov, Resul, et al.
Published: (2022) -
On The Statistical Limits of Self-Improving Agents
by: Wang, Charles L., et al.
Published: (2025) -
Information-Theoretic Dual Memory System for Continual Learning
by: Wu, RunQing, et al.
Published: (2025)