Fundamental Limits of Black-Box Safety Evaluation: Information-Theoretic and Computational Barriers from Latent Context Conditioning
Fuente:
arXiv
Saved in:
| Main Author: | Srivastava, Vishal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Information-Theoretic Limits of Safety Verification for Self-Improving Systems
by: Scrivens, Arsenios
Published: (2026)
by: Scrivens, Arsenios
Published: (2026)
Beyond the Black Box: A Statistical Model for LLM Reasoning and Inference
by: Dalal, Siddhartha, et al.
Published: (2024)
by: Dalal, Siddhartha, et al.
Published: (2024)
The Case for Developing a Foundation Model for Planning-like Tasks from Scratch
by: Srivastava, Biplav, et al.
Published: (2024)
by: Srivastava, Biplav, et al.
Published: (2024)
In-Context Black-Box Optimization with Unreliable Feedback
by: Blumer, Nicolas Samuel, et al.
Published: (2026)
by: Blumer, Nicolas Samuel, et al.
Published: (2026)
From Black-Box to White-Box: Control-Theoretic Neural Network Interpretability
by: Moon, Jihoon
Published: (2025)
by: Moon, Jihoon
Published: (2025)
Quantifying Automation Risk in High-Automation AI Systems: A Bayesian Framework for Failure Propagation and Optimal Oversight
by: Srivastava, Vishal, et al.
Published: (2026)
by: Srivastava, Vishal, et al.
Published: (2026)
Discovering and Learning Probabilistic Models of Black-Box AI Capabilities
by: Bramblett, Daniel, et al.
Published: (2025)
by: Bramblett, Daniel, et al.
Published: (2025)
Reinforced In-Context Black-Box Optimization
by: Song, Lei, et al.
Published: (2024)
by: Song, Lei, et al.
Published: (2024)
Capabilities and Fundamental Limits of Latent Chain-of-Thought
by: Zou, Jiaxuan, et al.
Published: (2026)
by: Zou, Jiaxuan, et al.
Published: (2026)
Advanced Black-Box Tuning of Large Language Models with Limited API Calls
by: Xie, Zhikang, et al.
Published: (2025)
by: Xie, Zhikang, et al.
Published: (2025)
FABLE: A Novel Data-Flow Analysis Benchmark on Procedural Text for Large Language Model Evaluation
by: Pallagani, Vishal, et al.
Published: (2025)
by: Pallagani, Vishal, et al.
Published: (2025)
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
by: Shen, Lingfeng, et al.
Published: (2024)
by: Shen, Lingfeng, et al.
Published: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)
by: Kroeger, Nicholas, et al.
Published: (2023)
Context-Former: Stitching via Latent Conditioned Sequence Modeling
by: Zhang, Ziqi, et al.
Published: (2024)
by: Zhang, Ziqi, et al.
Published: (2024)
Surrogate-Guided Quantum Discovery in Black-Box Landscapes with Latent-Quadratic Interaction Embedding Transformers
by: Gopalakrishnan, Saisubramaniam, et al.
Published: (2026)
by: Gopalakrishnan, Saisubramaniam, et al.
Published: (2026)
The Query Channel: Information-Theoretic Limits of Masking-Based Explanations
by: Karakaya, Erciyes, et al.
Published: (2026)
by: Karakaya, Erciyes, et al.
Published: (2026)
An Information-Theoretic Framework for Comparing Voice and Text Explainability
by: Rajhans, Mona, et al.
Published: (2026)
by: Rajhans, Mona, et al.
Published: (2026)
Stein Variational Black-Box Combinatorial Optimization
by: Landais, Thomas, et al.
Published: (2026)
by: Landais, Thomas, et al.
Published: (2026)
LongSafety: Evaluating Long-Context Safety of Large Language Models
by: Lu, Yida, et al.
Published: (2025)
by: Lu, Yida, et al.
Published: (2025)
Learning to Decide with Just Enough: Information-Theoretic Context Summarization for CMDPs
by: Liu, Peidong, et al.
Published: (2025)
by: Liu, Peidong, et al.
Published: (2025)
PLANTS: A Novel Problem and Dataset for Summarization of Planning-Like (PL) Tasks
by: Pallagani, Vishal, et al.
Published: (2024)
by: Pallagani, Vishal, et al.
Published: (2024)
Instance Generation for Meta-Black-Box Optimization through Latent Space Reverse Engineering
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Contract And Conquer: How to Provably Compute Adversarial Examples for a Black-Box Model?
by: Chistyakova, Anna, et al.
Published: (2026)
by: Chistyakova, Anna, et al.
Published: (2026)
Self-Improving Safety Performance of Reinforcement Learning Based Driving with Black-Box Verification Algorithms
by: Dagdanov, Resul, et al.
Published: (2022)
by: Dagdanov, Resul, et al.
Published: (2022)
Evaluation of Black-Box XAI Approaches for Predictors of Values of Boolean Formulae
by: Armoni-Friedmann, Stav, et al.
Published: (2025)
by: Armoni-Friedmann, Stav, et al.
Published: (2025)
From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models
by: Li, Xinyang, et al.
Published: (2025)
by: Li, Xinyang, et al.
Published: (2025)
ASBI: Leveraging Informative Real-World Data for Active Black-Box Simulator Tuning
by: Kim, Gahee, et al.
Published: (2025)
by: Kim, Gahee, et al.
Published: (2025)
Safety Evaluation of DeepSeek Models in Chinese Contexts
by: Zhang, Wenjing, et al.
Published: (2025)
by: Zhang, Wenjing, et al.
Published: (2025)
Hallucination Stations: On Some Basic Limitations of Transformer-Based Language Models
by: Sikka, Varin, et al.
Published: (2025)
by: Sikka, Varin, et al.
Published: (2025)
Sharpness-Aware Black-Box Optimization
by: Ye, Feiyang, et al.
Published: (2024)
by: Ye, Feiyang, et al.
Published: (2024)
Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs
by: Li, Zhuo, et al.
Published: (2024)
by: Li, Zhuo, et al.
Published: (2024)
BarrierBench: Evaluating Large Language Models for Safety Verification in Dynamical Systems
by: Taheri, Ali, et al.
Published: (2025)
by: Taheri, Ali, et al.
Published: (2025)
On Sample-Efficient Generalized Planning via Learned Transition Models
by: Gupta, Nitin, et al.
Published: (2026)
by: Gupta, Nitin, et al.
Published: (2026)
BarrierSteer: LLM Safety via Learning Barrier Steering
by: Tran, Thanh Q., et al.
Published: (2026)
by: Tran, Thanh Q., et al.
Published: (2026)
Black Box Model Explanations and the Human Interpretability Expectations -- An Analysis in the Context of Homicide Prediction
by: Ribeiro, José, et al.
Published: (2022)
by: Ribeiro, José, et al.
Published: (2022)
Turning Black Box into White Box: Dataset Distillation Leaks
by: Chen, Huajie, et al.
Published: (2026)
by: Chen, Huajie, et al.
Published: (2026)
Theoretical Barriers in Bellman-Based Reinforcement Learning
by: Pinon, Brieuc, et al.
Published: (2025)
by: Pinon, Brieuc, et al.
Published: (2025)
Benchmarking Instance-Centric Counterfactual Algorithms for XAI: From White Box to Black Box
by: Moreira, Catarina, et al.
Published: (2022)
by: Moreira, Catarina, et al.
Published: (2022)
Comparative Analysis of Black-Box and White-Box Machine Learning Model in Phishing Detection
by: Fajar, Abdullah, et al.
Published: (2024)
by: Fajar, Abdullah, et al.
Published: (2024)
Predictive Monitoring of Black-Box Dynamical Systems
by: Henzinger, Thomas A., et al.
Published: (2024)
by: Henzinger, Thomas A., et al.
Published: (2024)
Similar Items
-
Information-Theoretic Limits of Safety Verification for Self-Improving Systems
by: Scrivens, Arsenios
Published: (2026) -
Beyond the Black Box: A Statistical Model for LLM Reasoning and Inference
by: Dalal, Siddhartha, et al.
Published: (2024) -
The Case for Developing a Foundation Model for Planning-like Tasks from Scratch
by: Srivastava, Biplav, et al.
Published: (2024) -
In-Context Black-Box Optimization with Unreliable Feedback
by: Blumer, Nicolas Samuel, et al.
Published: (2026) -
From Black-Box to White-Box: Control-Theoretic Neural Network Interpretability
by: Moon, Jihoon
Published: (2025)