Influence-based Attributions can be Manipulated
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yadav, Chhavi, Wu, Ruihan, Chaudhuri, Kamalika |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FairProof : Confidential and Certifiable Fairness for Neural Networks
von: Yadav, Chhavi, et al.
Veröffentlicht: (2024)
von: Yadav, Chhavi, et al.
Veröffentlicht: (2024)
ExpProof : Operationalizing Explanations for Confidential Models with ZKPs
von: Yadav, Chhavi, et al.
Veröffentlicht: (2025)
von: Yadav, Chhavi, et al.
Veröffentlicht: (2025)
Can We Infer Confidential Properties of Training Data from LLMs?
von: Huang, Pengrun, et al.
Veröffentlicht: (2025)
von: Huang, Pengrun, et al.
Veröffentlicht: (2025)
Curriculum Learning for Safety Alignment
von: Kumar, Sandeep, et al.
Veröffentlicht: (2026)
von: Kumar, Sandeep, et al.
Veröffentlicht: (2026)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
von: Wang, Erchi, et al.
Veröffentlicht: (2026)
von: Wang, Erchi, et al.
Veröffentlicht: (2026)
Better Membership Inference Privacy Measurement through Discrepancy
von: Wu, Ruihan, et al.
Veröffentlicht: (2024)
von: Wu, Ruihan, et al.
Veröffentlicht: (2024)
Learning-Time Encoding Shapes Unlearning in LLMs
von: Wu, Ruihan, et al.
Veröffentlicht: (2025)
von: Wu, Ruihan, et al.
Veröffentlicht: (2025)
Evaluating Deep Unlearning in Large Language Models
von: Wu, Ruihan, et al.
Veröffentlicht: (2024)
von: Wu, Ruihan, et al.
Veröffentlicht: (2024)
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
von: Koga, Tatsuki, et al.
Veröffentlicht: (2024)
von: Koga, Tatsuki, et al.
Veröffentlicht: (2024)
Revisiting Data Attribution for Influence Functions
von: Zhu, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhu, Hongbo, et al.
Veröffentlicht: (2025)
Integrated Influence: Data Attribution with Baseline
von: Yang, Linxiao, et al.
Veröffentlicht: (2025)
von: Yang, Linxiao, et al.
Veröffentlicht: (2025)
Interaction-Aware Influence Functions for Group Attribution
von: Heo, Jaeseung, et al.
Veröffentlicht: (2026)
von: Heo, Jaeseung, et al.
Veröffentlicht: (2026)
Accumulative SGD Influence Estimation for Data Attribution
von: Shi, Yunxiao, et al.
Veröffentlicht: (2025)
von: Shi, Yunxiao, et al.
Veröffentlicht: (2025)
Influence Functions for Scalable Data Attribution in Diffusion Models
von: Mlodozeniec, Bruno, et al.
Veröffentlicht: (2024)
von: Mlodozeniec, Bruno, et al.
Veröffentlicht: (2024)
Diffusion Attribution Score: Evaluating Training Data Influence in Diffusion Models
von: Lin, Jinxu, et al.
Veröffentlicht: (2024)
von: Lin, Jinxu, et al.
Veröffentlicht: (2024)
Distributional Training Data Attribution: What do Influence Functions Sample?
von: Mlodozeniec, Bruno, et al.
Veröffentlicht: (2025)
von: Mlodozeniec, Bruno, et al.
Veröffentlicht: (2025)
Z0-Inf: Zeroth Order Approximation for Data Influence
von: Kokhlikyan, Narine, et al.
Veröffentlicht: (2025)
von: Kokhlikyan, Narine, et al.
Veröffentlicht: (2025)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
BoolGebra: Attributed Graph-learning for Boolean Algebraic Manipulation
von: Li, Yingjie, et al.
Veröffentlicht: (2024)
von: Li, Yingjie, et al.
Veröffentlicht: (2024)
Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation Models
von: Wen, Yeming, et al.
Veröffentlicht: (2024)
von: Wen, Yeming, et al.
Veröffentlicht: (2024)
Harmonic Mobile Manipulation
von: Yang, Ruihan, et al.
Veröffentlicht: (2023)
von: Yang, Ruihan, et al.
Veröffentlicht: (2023)
VRAIL: Vectorized Reward-based Attribution for Interpretable Learning
von: Kim, Jina, et al.
Veröffentlicht: (2025)
von: Kim, Jina, et al.
Veröffentlicht: (2025)
Full-Atom Peptide Design based on Multi-modal Flow Matching
von: Li, Jiahan, et al.
Veröffentlicht: (2024)
von: Li, Jiahan, et al.
Veröffentlicht: (2024)
AIMM: An AI-Driven Multimodal Framework for Detecting Social-Media-Influenced Stock Market Manipulation
von: Neela, Sandeep
Veröffentlicht: (2025)
von: Neela, Sandeep
Veröffentlicht: (2025)
Machine Learning with Privacy for Protected Attributes
von: Mahloujifar, Saeed, et al.
Veröffentlicht: (2025)
von: Mahloujifar, Saeed, et al.
Veröffentlicht: (2025)
Quantifying Manifolds: Do the manifolds learned by Generative Adversarial Networks converge to the real data manifold
von: Chaudhuri, Anupam, et al.
Veröffentlicht: (2024)
von: Chaudhuri, Anupam, et al.
Veröffentlicht: (2024)
Distribution Learning with Valid Outputs Beyond the Worst-Case
von: Rittler, Nick, et al.
Veröffentlicht: (2024)
von: Rittler, Nick, et al.
Veröffentlicht: (2024)
A Closer Look at the Learnability of Out-of-Distribution (OOD) Detection
von: Garov, Konstantin, et al.
Veröffentlicht: (2025)
von: Garov, Konstantin, et al.
Veröffentlicht: (2025)
Pinpointing crucial steps: Attribution-based Credit Assignment for Verifiable Reinforcement Learning
von: Yin, Junxi, et al.
Veröffentlicht: (2025)
von: Yin, Junxi, et al.
Veröffentlicht: (2025)
Dynamic Influence Tracker: Measuring Time-Varying Sample Influence During Training
von: Xu, Jie, et al.
Veröffentlicht: (2025)
von: Xu, Jie, et al.
Veröffentlicht: (2025)
Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private
von: Wu, Ruihan, et al.
Veröffentlicht: (2025)
von: Wu, Ruihan, et al.
Veröffentlicht: (2025)
NN-Former: Rethinking Graph Structure in Neural Architecture Representation
von: Xu, Ruihan, et al.
Veröffentlicht: (2025)
von: Xu, Ruihan, et al.
Veröffentlicht: (2025)
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
Understanding and Improving Noisy Embedding Techniques in Instruction Finetuning
von: Yadav, Abhay
Veröffentlicht: (2026)
von: Yadav, Abhay
Veröffentlicht: (2026)
SPICED: Syntactical Bug and Trojan Pattern Identification in A/MS Circuits using LLM-Enhanced Detection
von: Chaudhuri, Jayeeta, et al.
Veröffentlicht: (2024)
von: Chaudhuri, Jayeeta, et al.
Veröffentlicht: (2024)
Prompt Tuning Strikes Back: Customizing Foundation Models with Low-Rank Prompt Adaptation
von: Jain, Abhinav, et al.
Veröffentlicht: (2024)
von: Jain, Abhinav, et al.
Veröffentlicht: (2024)
Generalized Group Data Attribution
von: Ley, Dan, et al.
Veröffentlicht: (2024)
von: Ley, Dan, et al.
Veröffentlicht: (2024)
Impossibility Theorems for Feature Attribution
von: Bilodeau, Blair, et al.
Veröffentlicht: (2022)
von: Bilodeau, Blair, et al.
Veröffentlicht: (2022)
Batched Low-Rank Adaptation of Foundation Models
von: Wen, Yeming, et al.
Veröffentlicht: (2023)
von: Wen, Yeming, et al.
Veröffentlicht: (2023)
Enhancing LLMs for Physics Problem-Solving using Reinforcement Learning with Human-AI Feedback
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FairProof : Confidential and Certifiable Fairness for Neural Networks
von: Yadav, Chhavi, et al.
Veröffentlicht: (2024) -
ExpProof : Operationalizing Explanations for Confidential Models with ZKPs
von: Yadav, Chhavi, et al.
Veröffentlicht: (2025) -
Can We Infer Confidential Properties of Training Data from LLMs?
von: Huang, Pengrun, et al.
Veröffentlicht: (2025) -
Curriculum Learning for Safety Alignment
von: Kumar, Sandeep, et al.
Veröffentlicht: (2026) -
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
von: Wang, Erchi, et al.
Veröffentlicht: (2026)