Interpretability Guarantees with Merlin-Arthur Classifiers
Fuente:
arXiv
Saved in:
| Main Authors: | Wäldchen, Stephan, Sharma, Kartikey, Turan, Berkant, Zimmer, Max, Pokutta, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses
by: Głuch, Grzegorz, et al.
Published: (2024)
by: Głuch, Grzegorz, et al.
Published: (2024)
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
by: Turan, Berkant, et al.
Published: (2025)
by: Turan, Berkant, et al.
Published: (2025)
Quantifying Behavioral Dissimilarity Between Mathematical Expressions
by: Mežnar, Sebastian, et al.
Published: (2024)
by: Mežnar, Sebastian, et al.
Published: (2024)
Feature Relevancy, Necessity and Usefulness: Complexity and Algorithms
by: Capdevielle, Tomás, et al.
Published: (2025)
by: Capdevielle, Tomás, et al.
Published: (2025)
A Theoretical Framework for Adaptive Utility-Weighted Benchmarking
by: Waggoner, Philip
Published: (2026)
by: Waggoner, Philip
Published: (2026)
From Language Models to Practical Self-Improving Computer Agents
by: Sheng, Alex
Published: (2024)
by: Sheng, Alex
Published: (2024)
A Taxonomy of Omnicidal Futures Involving Artificial Intelligence
by: Critch, Andrew, et al.
Published: (2025)
by: Critch, Andrew, et al.
Published: (2025)
Intelligence as Computation
by: Brock, Oliver
Published: (2024)
by: Brock, Oliver
Published: (2024)
CyGATE: Game-Theoretic Cyber Attack-Defense Engine for Patch Strategy Optimization
by: Jiang, Yuning, et al.
Published: (2025)
by: Jiang, Yuning, et al.
Published: (2025)
The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment
by: Spizzirri, Austin
Published: (2025)
by: Spizzirri, Austin
Published: (2025)
Instilling Organisational Values in Firefighters through Simulation-Based Training
by: Osman, Nardine, et al.
Published: (2025)
by: Osman, Nardine, et al.
Published: (2025)
Value-Aware Multiagent Systems
by: Osman, Nardine
Published: (2025)
by: Osman, Nardine
Published: (2025)
Dynamic Observation Policies in Observation Cost-Sensitive Reinforcement Learning
by: Bellinger, Colin, et al.
Published: (2023)
by: Bellinger, Colin, et al.
Published: (2023)
On the Invariants of Softmax Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
Heckerthoughts
by: Heckerman, David
Published: (2023)
by: Heckerman, David
Published: (2023)
Charting the Future of Scholarly Knowledge with AI: A Community Perspective
by: Jiomekong, Azanzi, et al.
Published: (2025)
by: Jiomekong, Azanzi, et al.
Published: (2025)
Achieving Distributive Justice in Federated Learning via Uncertainty Quantification
by: Carey, Alycia, et al.
Published: (2025)
by: Carey, Alycia, et al.
Published: (2025)
Return of the Schema: Building Complete Datasets for Machine Learning and Reasoning on Knowledge Graphs
by: Diliso, Ivan, et al.
Published: (2026)
by: Diliso, Ivan, et al.
Published: (2026)
ATEX-CF: Attack-Informed Counterfactual Explanations for Graph Neural Networks
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
by: Schaeffer, Joachim, et al.
Published: (2026)
by: Schaeffer, Joachim, et al.
Published: (2026)
Compressible Softmax-Attended Language under Incompressible Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
Online Decision Making with Generative Action Sets
by: Xu, Jianyu, et al.
Published: (2025)
by: Xu, Jianyu, et al.
Published: (2025)
Reasoning Promotes Robustness in Theory of Mind Tasks
by: de Haan, Ian B., et al.
Published: (2026)
by: de Haan, Ian B., et al.
Published: (2026)
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
by: Shafieinejad, Masoumeh, et al.
Published: (2026)
by: Shafieinejad, Masoumeh, et al.
Published: (2026)
Bridging Voting and Deliberation with Algorithms: Field Insights from vTaiwan and Kultur Komitee
by: Yang, Joshua C., et al.
Published: (2025)
by: Yang, Joshua C., et al.
Published: (2025)
ECG-FM: An Open Electrocardiogram Foundation Model
by: McKeen, Kaden, et al.
Published: (2024)
by: McKeen, Kaden, et al.
Published: (2024)
Enhancing PyKEEN with Multiple Negative Sampling Solutions for Knowledge Graph Embedding Models
by: d'Amato, Claudia, et al.
Published: (2025)
by: d'Amato, Claudia, et al.
Published: (2025)
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
by: Huang, Ningyuan, et al.
Published: (2025)
by: Huang, Ningyuan, et al.
Published: (2025)
Driving down Poisson error can offset classification error in clinical tasks
by: Delahunt, Charles B., et al.
Published: (2024)
by: Delahunt, Charles B., et al.
Published: (2024)
Machine Learning for Physical Simulation Challenge Results and Retrospective Analysis: Power Grid Use Case
by: Leyli-Abadi, Milad, et al.
Published: (2025)
by: Leyli-Abadi, Milad, et al.
Published: (2025)
Computing Game Symmetries and Equilibria That Respect Them
by: Tewolde, Emanuel, et al.
Published: (2025)
by: Tewolde, Emanuel, et al.
Published: (2025)
Ideological Isolation in Online Social Networks: A Survey of Computational Definitions, Metrics, and Mitigation Strategies
by: Wang, Xiaodan, et al.
Published: (2026)
by: Wang, Xiaodan, et al.
Published: (2026)
Combination of Weak Learners eXplanations to Improve Random Forest eXplicability Robustness
by: Pala, Riccardo, et al.
Published: (2024)
by: Pala, Riccardo, et al.
Published: (2024)
Revenue-Sharing as Infrastructure: A Distributed Business Model for Generative AI Platforms
by: Mondjo, Ghislain Dorian Tchuente
Published: (2026)
by: Mondjo, Ghislain Dorian Tchuente
Published: (2026)
Vibe-Creation: The Epistemology of Human-AI Emergent Cognition
by: Levin, Ilya
Published: (2026)
by: Levin, Ilya
Published: (2026)
Dancing in the Shadows: Harnessing Ambiguity for Fairer Classifiers
by: Barrainkua, Ainhize, et al.
Published: (2024)
by: Barrainkua, Ainhize, et al.
Published: (2024)
Report of the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science
by: McInnes, Lois Curfman, et al.
Published: (2025)
by: McInnes, Lois Curfman, et al.
Published: (2025)
Proceedings of the 20th International Conference on Knowledge, Information and Creativity Support Systems (KICSS 2025)
by: Hayama, Edited by Tessai, et al.
Published: (2025)
by: Hayama, Edited by Tessai, et al.
Published: (2025)
An Automatic Text Classification Method Based on Hierarchical Taxonomies, Neural Networks and Document Embedding: The NETHIC Tool
by: Lomasto, Luigi, et al.
Published: (2026)
by: Lomasto, Luigi, et al.
Published: (2026)
A ZeNN architecture to avoid the Gaussian trap
by: Carvalho, Luís, et al.
Published: (2025)
by: Carvalho, Luís, et al.
Published: (2025)
Similar Items
-
The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses
by: Głuch, Grzegorz, et al.
Published: (2024) -
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
by: Turan, Berkant, et al.
Published: (2025) -
Quantifying Behavioral Dissimilarity Between Mathematical Expressions
by: Mežnar, Sebastian, et al.
Published: (2024) -
Feature Relevancy, Necessity and Usefulness: Complexity and Algorithms
by: Capdevielle, Tomás, et al.
Published: (2025) -
A Theoretical Framework for Adaptive Utility-Weighted Benchmarking
by: Waggoner, Philip
Published: (2026)