Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Yijun, Shi, Taiwei, Zhao, Jieyu, Ma, Jiaqi W. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
von: Shi, Taiwei, et al.
Veröffentlicht: (2025)
von: Shi, Taiwei, et al.
Veröffentlicht: (2025)
Daunce: Data Attribution through Uncertainty Estimation
von: Pan, Xingyuan, et al.
Veröffentlicht: (2025)
von: Pan, Xingyuan, et al.
Veröffentlicht: (2025)
Exploring Training Data Attribution under Limited Access Constraints
von: Zhang, Shiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Shiyuan, et al.
Veröffentlicht: (2025)
Efficient Ensembles Improve Training Data Attribution
von: Deng, Junwei, et al.
Veröffentlicht: (2024)
von: Deng, Junwei, et al.
Veröffentlicht: (2024)
Adversarial Attacks on Data Attribution
von: Wang, Xinhe, et al.
Veröffentlicht: (2024)
von: Wang, Xinhe, et al.
Veröffentlicht: (2024)
Measuring Fine-Grained Relatedness in Multitask Learning via Data Attribution
von: Tu, Yiwen, et al.
Veröffentlicht: (2025)
von: Tu, Yiwen, et al.
Veröffentlicht: (2025)
$\texttt{dattri}$: A Library for Efficient Data Attribution
von: Deng, Junwei, et al.
Veröffentlicht: (2024)
von: Deng, Junwei, et al.
Veröffentlicht: (2024)
Enhancing Training Data Attribution with Representational Optimization
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
A Versatile Influence Function for Data Attribution with Non-Decomposable Loss
von: Deng, Junwei, et al.
Veröffentlicht: (2024)
von: Deng, Junwei, et al.
Veröffentlicht: (2024)
Skill Reuse as Compression in Agentic RL
von: Xu, Zhikun, et al.
Veröffentlicht: (2026)
von: Xu, Zhikun, et al.
Veröffentlicht: (2026)
Experiential Reinforcement Learning
von: Shi, Taiwei, et al.
Veröffentlicht: (2026)
von: Shi, Taiwei, et al.
Veröffentlicht: (2026)
GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
von: Hu, Pingbang, et al.
Veröffentlicht: (2025)
von: Hu, Pingbang, et al.
Veröffentlicht: (2025)
Safer-Instruct: Aligning Language Models with Automated Preference Data
von: Shi, Taiwei, et al.
Veröffentlicht: (2023)
von: Shi, Taiwei, et al.
Veröffentlicht: (2023)
No Safe Dose: How Training Data Drives Unsafe Image Generation
von: Friedrich, Felix, et al.
Veröffentlicht: (2026)
von: Friedrich, Felix, et al.
Veröffentlicht: (2026)
Training Data Attribution via Approximate Unrolled Differentiation
von: Bae, Juhan, et al.
Veröffentlicht: (2024)
von: Bae, Juhan, et al.
Veröffentlicht: (2024)
Taming Hyperparameter Sensitivity in Data Attribution: Practical Selection Without Costly Retraining
von: Wang, Weiyi, et al.
Veröffentlicht: (2025)
von: Wang, Weiyi, et al.
Veröffentlicht: (2025)
A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning
von: Hu, Yuzheng, et al.
Veröffentlicht: (2025)
von: Hu, Yuzheng, et al.
Veröffentlicht: (2025)
Controllable Pareto Trade-off between Fairness and Accuracy
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines
von: Deng, Junwei, et al.
Veröffentlicht: (2026)
von: Deng, Junwei, et al.
Veröffentlicht: (2026)
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)
FairSHAP: Preprocessing for Fairness Through Attribution-Based Data Augmentation
von: Zhu, Lin, et al.
Veröffentlicht: (2025)
von: Zhu, Lin, et al.
Veröffentlicht: (2025)
Better Training Data Attribution via Better Inverse Hessian-Vector Products
von: Wang, Andrew, et al.
Veröffentlicht: (2025)
von: Wang, Andrew, et al.
Veröffentlicht: (2025)
Scalable Data Attribution via Forward-Only Test-Time Inference
von: Ma, Sibo, et al.
Veröffentlicht: (2025)
von: Ma, Sibo, et al.
Veröffentlicht: (2025)
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
von: Chen, Jianhui, et al.
Veröffentlicht: (2026)
von: Chen, Jianhui, et al.
Veröffentlicht: (2026)
A Data-Driven Gaussian Process Filter for Electrocardiogram Denoising
von: Dumitru, Mircea, et al.
Veröffentlicht: (2023)
von: Dumitru, Mircea, et al.
Veröffentlicht: (2023)
Data Denoising and Derivative Estimation for Data-Driven Modeling of Nonlinear Dynamical Systems
von: Yao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Yao, Jiaqi, et al.
Veröffentlicht: (2025)
Efficient Sketches for Training Data Attribution and Studying the Loss Landscape
von: Schioppa, Andrea
Veröffentlicht: (2024)
von: Schioppa, Andrea
Veröffentlicht: (2024)
Diffusion-Scheduled Denoising Autoencoders for Anomaly Detection in Tabular Data
von: Sattarov, Timur, et al.
Veröffentlicht: (2025)
von: Sattarov, Timur, et al.
Veröffentlicht: (2025)
LLM-Powered Text-Attributed Graph Anomaly Detection via Retrieval-Augmented Reasoning
von: Xu, Haoyan, et al.
Veröffentlicht: (2025)
von: Xu, Haoyan, et al.
Veröffentlicht: (2025)
Nonparametric Data Attribution for Diffusion Models
von: Zhao, Yutian, et al.
Veröffentlicht: (2025)
von: Zhao, Yutian, et al.
Veröffentlicht: (2025)
Learning to Weight Parameters for Training Data Attribution
von: Li, Shuangqi, et al.
Veröffentlicht: (2025)
von: Li, Shuangqi, et al.
Veröffentlicht: (2025)
GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning
von: Murata, Naoki, et al.
Veröffentlicht: (2026)
von: Murata, Naoki, et al.
Veröffentlicht: (2026)
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
von: Bareeva, Dilyara, et al.
Veröffentlicht: (2024)
von: Bareeva, Dilyara, et al.
Veröffentlicht: (2024)
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
von: Eichin, Florian, et al.
Veröffentlicht: (2025)
von: Eichin, Florian, et al.
Veröffentlicht: (2025)
LoRIF: Low-Rank Influence Functions for Scalable Training Data Attribution
von: Li, Shuangqi, et al.
Veröffentlicht: (2026)
von: Li, Shuangqi, et al.
Veröffentlicht: (2026)
Multimodal Data Curation via Object Detection and Filter Ensembles
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)
ReTrack: Data Unlearning in Diffusion Models through Redirecting the Denoising Trajectory
von: Shi, Qitan, et al.
Veröffentlicht: (2025)
von: Shi, Qitan, et al.
Veröffentlicht: (2025)
ShieldNN: A Provably Safe NN Filter for Unsafe NN Controllers
von: Ferlez, James, et al.
Veröffentlicht: (2020)
von: Ferlez, James, et al.
Veröffentlicht: (2020)
Data Attribution in Adaptive Learning
von: Rege, Amit Kiran
Veröffentlicht: (2026)
von: Rege, Amit Kiran
Veröffentlicht: (2026)
Accumulative SGD Influence Estimation for Data Attribution
von: Shi, Yunxiao, et al.
Veröffentlicht: (2025)
von: Shi, Yunxiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
von: Shi, Taiwei, et al.
Veröffentlicht: (2025) -
Daunce: Data Attribution through Uncertainty Estimation
von: Pan, Xingyuan, et al.
Veröffentlicht: (2025) -
Exploring Training Data Attribution under Limited Access Constraints
von: Zhang, Shiyuan, et al.
Veröffentlicht: (2025) -
Efficient Ensembles Improve Training Data Attribution
von: Deng, Junwei, et al.
Veröffentlicht: (2024) -
Adversarial Attacks on Data Attribution
von: Wang, Xinhe, et al.
Veröffentlicht: (2024)