Model Guidance via Robust Feature Attribution
Fuente:
arXiv
Saved in:
| Main Authors: | Ghitu, Mihnea, Piratla, Vihari, Wicker, Matthew |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Poisoning Robustness Certification for Natural Language Generation
by: Ghitu, Mihnea, et al.
Published: (2026)
by: Ghitu, Mihnea, et al.
Published: (2026)
XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks
by: Jain, Purvam, et al.
Published: (2026)
by: Jain, Purvam, et al.
Published: (2026)
Rethinking Cross-lingual Gaps from a Statistical Viewpoint
by: Piratla, Vihari, et al.
Published: (2025)
by: Piratla, Vihari, et al.
Published: (2025)
Estimation of Concept Explanations Should be Uncertainty Aware
by: Piratla, Vihari, et al.
Published: (2023)
by: Piratla, Vihari, et al.
Published: (2023)
Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts Transformers
by: Li, Albus Yizhuo, et al.
Published: (2026)
by: Li, Albus Yizhuo, et al.
Published: (2026)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
by: Anisimov, Maksim, et al.
Published: (2026)
by: Anisimov, Maksim, et al.
Published: (2026)
Abstract Gradient Training: A Unified Certification Framework for Data Poisoning, Unlearning, and Differential Privacy
by: Sosnin, Philip, et al.
Published: (2025)
by: Sosnin, Philip, et al.
Published: (2025)
Provably Safe Model Updates
by: Elmecker-Plakolm, Leo, et al.
Published: (2025)
by: Elmecker-Plakolm, Leo, et al.
Published: (2025)
The Attribution Contract: Feature Attribution for Generative Language Models
by: Nguyen, Giang
Published: (2026)
by: Nguyen, Giang
Published: (2026)
Practical Attribution Guidance for Rashomon Sets
by: Li, Sichao, et al.
Published: (2024)
by: Li, Sichao, et al.
Published: (2024)
On the Robustness of Bayesian Neural Networks to Adversarial Attacks
by: Bortolussi, Luca, et al.
Published: (2022)
by: Bortolussi, Luca, et al.
Published: (2022)
Model Monitoring in the Absence of Labeled Data via Feature Attributions Distributions
by: Mougan, Carlos
Published: (2025)
by: Mougan, Carlos
Published: (2025)
Certified Robustness to Data Poisoning in Gradient-Based Training
by: Sosnin, Philip, et al.
Published: (2024)
by: Sosnin, Philip, et al.
Published: (2024)
Certified $\ell_2$ Attribution Robustness via Uniformly Smoothed Attributions
by: Wang, Fan, et al.
Published: (2024)
by: Wang, Fan, et al.
Published: (2024)
Topology Guidance: Controlling the Outputs of Generative Models via Vector Field Topology
by: Wang, Xiaohan, et al.
Published: (2025)
by: Wang, Xiaohan, et al.
Published: (2025)
RoSHAP: A Distributional Framework and Robust Metric for Stable Feature Attribution
by: Xiang, Lanxin, et al.
Published: (2026)
by: Xiang, Lanxin, et al.
Published: (2026)
Enhancing Visual Feature Attribution via Weighted Integrated Gradients
by: Tuan, Kien Tran Duc, et al.
Published: (2025)
by: Tuan, Kien Tran Duc, et al.
Published: (2025)
AttributionLab: Faithfulness of Feature Attribution Under Controllable Environments
by: Zhang, Yang, et al.
Published: (2023)
by: Zhang, Yang, et al.
Published: (2023)
GRAFT: Auditing Graph Neural Networks via Global Feature Attribution
by: Sahoo, Rishi Raj, et al.
Published: (2026)
by: Sahoo, Rishi Raj, et al.
Published: (2026)
Exploring the Relationship Between Feature Attribution Methods and Model Performance
by: Silva, Priscylla, et al.
Published: (2024)
by: Silva, Priscylla, et al.
Published: (2024)
Generalized Attention Flow: Feature Attribution for Transformer Models via Maximum Flow
by: Azarkhalili, Behrooz, et al.
Published: (2025)
by: Azarkhalili, Behrooz, et al.
Published: (2025)
Feature Attribution from First Principles
by: Taimeskhanov, Magamed, et al.
Published: (2025)
by: Taimeskhanov, Magamed, et al.
Published: (2025)
Disentangling Interactions and Dependencies in Feature Attribution
by: König, Gunnar, et al.
Published: (2024)
by: König, Gunnar, et al.
Published: (2024)
Training Feature Attribution for Vision Models
by: Bacha, Aziz, et al.
Published: (2025)
by: Bacha, Aziz, et al.
Published: (2025)
Energy-Based Model for Accurate Estimation of Shapley Values in Feature Attribution
by: Lu, Cheng, et al.
Published: (2024)
by: Lu, Cheng, et al.
Published: (2024)
DeepACTIF: Efficient Feature Attribution via Activation Traces in Neural Sequence Models
by: Hosp, Benedikt W.
Published: (2025)
by: Hosp, Benedikt W.
Published: (2025)
Invariant Features in Language Models: Geometric Characterization and Model Attribution
by: Dasgupta, Agnibh, et al.
Published: (2026)
by: Dasgupta, Agnibh, et al.
Published: (2026)
Explaining Concept Shift with Interpretable Feature Attribution
by: Lyu, Ruiqi, et al.
Published: (2025)
by: Lyu, Ruiqi, et al.
Published: (2025)
Boundary-Aware Uncertainty for Feature Attribution Explainers
by: Hill, Davin, et al.
Published: (2022)
by: Hill, Davin, et al.
Published: (2022)
Missingness Bias Calibration in Feature Attribution Explanations
by: Sridhar, Shailesh, et al.
Published: (2026)
by: Sridhar, Shailesh, et al.
Published: (2026)
Impossibility Theorems for Feature Attribution
by: Bilodeau, Blair, et al.
Published: (2022)
by: Bilodeau, Blair, et al.
Published: (2022)
Enabling Asymmetric Knowledge Transfer in Multi-Task Learning with Self-Auxiliaries
by: Graffeuille, Olivier, et al.
Published: (2024)
by: Graffeuille, Olivier, et al.
Published: (2024)
Model-Based Counterfactual Explanations Incorporating Feature Space Attributes for Tabular Data
by: Sumiya, Yuta, et al.
Published: (2024)
by: Sumiya, Yuta, et al.
Published: (2024)
On the Robustness of Distribution Support under Diffusion Guidance
by: Cao, Ruijia, et al.
Published: (2026)
by: Cao, Ruijia, et al.
Published: (2026)
Hybrid Attribution Priors for Explainable and Robust Model Training
by: Zhang, Zhuoran, et al.
Published: (2025)
by: Zhang, Zhuoran, et al.
Published: (2025)
FairGen: Controlling Sensitive Attributes for Fair Generations in Diffusion Models via Adaptive Latent Guidance
by: Kang, Mintong, et al.
Published: (2025)
by: Kang, Mintong, et al.
Published: (2025)
Probabilistic Stability Guarantees for Feature Attributions
by: Jin, Helen, et al.
Published: (2025)
by: Jin, Helen, et al.
Published: (2025)
Feature Attribution with Necessity and Sufficiency via Dual-stage Perturbation Test for Causal Explanation
by: Chen, Xuexin, et al.
Published: (2024)
by: Chen, Xuexin, et al.
Published: (2024)
SPRINT: Robust Model Attribution of Generated Images via Secret Pixel Reconstruction
by: Yao, Kai, et al.
Published: (2025)
by: Yao, Kai, et al.
Published: (2025)
Rethinking Robustness: A New Approach to Evaluating Feature Attribution Methods
by: Kiourti, Panagiota, et al.
Published: (2025)
by: Kiourti, Panagiota, et al.
Published: (2025)
Similar Items
-
Towards Poisoning Robustness Certification for Natural Language Generation
by: Ghitu, Mihnea, et al.
Published: (2026) -
XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks
by: Jain, Purvam, et al.
Published: (2026) -
Rethinking Cross-lingual Gaps from a Statistical Viewpoint
by: Piratla, Vihari, et al.
Published: (2025) -
Estimation of Concept Explanations Should be Uncertainty Aware
by: Piratla, Vihari, et al.
Published: (2023) -
Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts Transformers
by: Li, Albus Yizhuo, et al.
Published: (2026)