Classification with Conceptual Safeguards
Fuente:
arXiv
Guardado en:
| Autores principales: | Joren, Hailey, Marx, Charles, Ustun, Berk |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Observational Multiplicity
por: George, Erin, et al.
Publicado: (2025)
por: George, Erin, et al.
Publicado: (2025)
Excess Description Length of Learning Generalizable Predictors
por: Donoway, Elizabeth, et al.
Publicado: (2026)
por: Donoway, Elizabeth, et al.
Publicado: (2026)
Explanations are a Means to an End: Decision Theoretic Explanation Evaluation
por: Guo, Ziyang, et al.
Publicado: (2025)
por: Guo, Ziyang, et al.
Publicado: (2025)
Feature Responsiveness Scores: Model-Agnostic Explanations for Recourse
por: Cheon, Harry, et al.
Publicado: (2024)
por: Cheon, Harry, et al.
Publicado: (2024)
Regretful Decisions under Label Noise
por: Nagaraj, Sujay, et al.
Publicado: (2025)
por: Nagaraj, Sujay, et al.
Publicado: (2025)
Prediction without Preclusion: Recourse Verification with Reachable Sets
por: Kothari, Avni, et al.
Publicado: (2023)
por: Kothari, Avni, et al.
Publicado: (2023)
Concept Bottleneck Large Language Models
por: Sun, Chung-En, et al.
Publicado: (2024)
por: Sun, Chung-En, et al.
Publicado: (2024)
Understanding Fixed Predictions via Confined Regions
por: Lawless, Connor, et al.
Publicado: (2025)
por: Lawless, Connor, et al.
Publicado: (2025)
Statistical Inference for Responsiveness Verification
por: Cheon, Seung Hyun, et al.
Publicado: (2025)
por: Cheon, Seung Hyun, et al.
Publicado: (2025)
FINEST: Stabilizing Recommendations by Rank-Preserving Fine-Tuning
por: Oh, Sejoon, et al.
Publicado: (2024)
por: Oh, Sejoon, et al.
Publicado: (2024)
Learning under Temporal Label Noise
por: Nagaraj, Sujay, et al.
Publicado: (2024)
por: Nagaraj, Sujay, et al.
Publicado: (2024)
Predictive Churn with the Set of Good Models
por: Watson-Daniels, Jamelle, et al.
Publicado: (2024)
por: Watson-Daniels, Jamelle, et al.
Publicado: (2024)
Calibrated Probabilistic Forecasts for Arbitrary Sequences
por: Marx, Charles, et al.
Publicado: (2024)
por: Marx, Charles, et al.
Publicado: (2024)
Calibrated Regression Against An Adversary Without Regret
por: Deshpande, Shachi, et al.
Publicado: (2023)
por: Deshpande, Shachi, et al.
Publicado: (2023)
Safety in Graph Machine Learning: Threats and Safeguards
por: Wang, Song, et al.
Publicado: (2024)
por: Wang, Song, et al.
Publicado: (2024)
NeuralSentinel: Safeguarding Neural Network Reliability and Trustworthiness
por: Echeberria-Barrio, Xabier, et al.
Publicado: (2024)
por: Echeberria-Barrio, Xabier, et al.
Publicado: (2024)
Hyperparameter Tuning Through Pessimistic Bilevel Optimization
por: Ustun, Meltem Apaydin, et al.
Publicado: (2024)
por: Ustun, Meltem Apaydin, et al.
Publicado: (2024)
Boosting Digital Safeguards: Blending Cryptography and Steganography
por: Maiti, Anamitra, et al.
Publicado: (2024)
por: Maiti, Anamitra, et al.
Publicado: (2024)
An Example Safety Case for Safeguards Against Misuse
por: Clymer, Joshua, et al.
Publicado: (2025)
por: Clymer, Joshua, et al.
Publicado: (2025)
Explainable AI Through a Democratic Lens: DhondtXAI for D'Hondt-Projected Feature Attribution
por: Donmez, Turker Berk
Publicado: (2024)
por: Donmez, Turker Berk
Publicado: (2024)
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
por: Youstra, Jack, et al.
Publicado: (2025)
por: Youstra, Jack, et al.
Publicado: (2025)
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?
por: Markgraf, Hannah, et al.
Publicado: (2025)
por: Markgraf, Hannah, et al.
Publicado: (2025)
Online Calibrated and Conformal Prediction Improves Bayesian Optimization
por: Deshpande, Shachi, et al.
Publicado: (2021)
por: Deshpande, Shachi, et al.
Publicado: (2021)
Lightweight Conceptual Dictionary Learning for Text Classification Using Information Compression
por: Wan, Li, et al.
Publicado: (2024)
por: Wan, Li, et al.
Publicado: (2024)
GDeR: Safeguarding Efficiency, Balancing, and Robustness via Prototypical Graph Pruning
por: Zhang, Guibin, et al.
Publicado: (2024)
por: Zhang, Guibin, et al.
Publicado: (2024)
PFGuard: A Generative Framework with Privacy and Fairness Safeguards
por: Kim, Soyeon, et al.
Publicado: (2024)
por: Kim, Soyeon, et al.
Publicado: (2024)
MolMark: Safeguarding Molecular Structures through Learnable Atom-Level Watermarking
por: Hu, Runwen, et al.
Publicado: (2025)
por: Hu, Runwen, et al.
Publicado: (2025)
Zeroth-Order Non-Log-Concave Sampling with Variance Reduction and Applications to Inverse Problems
por: Sahin, M. Berk, et al.
Publicado: (2026)
por: Sahin, M. Berk, et al.
Publicado: (2026)
A Data-Driven Discretized CS:GO Simulation Environment to Facilitate Strategic Multi-Agent Planning Research
por: Wang, Yunzhe, et al.
Publicado: (2025)
por: Wang, Yunzhe, et al.
Publicado: (2025)
Predicting Air Temperature from Volumetric Urban Morphology with Machine Learning
por: Kıvılcım, Berk, et al.
Publicado: (2025)
por: Kıvılcım, Berk, et al.
Publicado: (2025)
Skewness-Robust Causal Discovery in Location-Scale Noise Models
por: Klippert, Daniel, et al.
Publicado: (2025)
por: Klippert, Daniel, et al.
Publicado: (2025)
X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
por: Wang, Yunzhe, et al.
Publicado: (2025)
por: Wang, Yunzhe, et al.
Publicado: (2025)
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning
por: Jiang, Zhiheng, et al.
Publicado: (2026)
por: Jiang, Zhiheng, et al.
Publicado: (2026)
Safeguarding Graph Neural Networks against Topology Inference Attacks
por: Fu, Jie, et al.
Publicado: (2025)
por: Fu, Jie, et al.
Publicado: (2025)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
por: Fu, Yao, et al.
Publicado: (2025)
por: Fu, Yao, et al.
Publicado: (2025)
SWaRL: Safeguard Code Watermarking via Reinforcement Learning
por: Javidnia, Neusha, et al.
Publicado: (2026)
por: Javidnia, Neusha, et al.
Publicado: (2026)
Speeding up Policy Simulation in Supply Chain RL
por: Farias, Vivek, et al.
Publicado: (2024)
por: Farias, Vivek, et al.
Publicado: (2024)
Progressive Safeguards for Safe and Model-Agnostic Reinforcement Learning
por: Omi, Nabil, et al.
Publicado: (2024)
por: Omi, Nabil, et al.
Publicado: (2024)
On Prompt-Driven Safeguarding for Large Language Models
por: Zheng, Chujie, et al.
Publicado: (2024)
por: Zheng, Chujie, et al.
Publicado: (2024)
Tamper-Resistant Safeguards for Open-Weight LLMs
por: Tamirisa, Rishub, et al.
Publicado: (2024)
por: Tamirisa, Rishub, et al.
Publicado: (2024)
Ejemplares similares
-
Observational Multiplicity
por: George, Erin, et al.
Publicado: (2025) -
Excess Description Length of Learning Generalizable Predictors
por: Donoway, Elizabeth, et al.
Publicado: (2026) -
Explanations are a Means to an End: Decision Theoretic Explanation Evaluation
por: Guo, Ziyang, et al.
Publicado: (2025) -
Feature Responsiveness Scores: Model-Agnostic Explanations for Recourse
por: Cheon, Harry, et al.
Publicado: (2024) -
Regretful Decisions under Label Noise
por: Nagaraj, Sujay, et al.
Publicado: (2025)