Helpful or Harmful Data? Fine-tuning-free Shapley Attribution for Explaining Language Model Predictions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jingtan, Lin, Xiaoqiang, Qiao, Rui, Foo, Chuan-Sheng, Low, Bryan Kian Hsiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Source Attribution for Large Language Model-Generated Data
von: Wang, Jingtan, et al.
Veröffentlicht: (2023)
von: Wang, Jingtan, et al.
Veröffentlicht: (2023)
Group-robust Sample Reweighting for Subpopulation Shifts via Influence Functions
von: Qiao, Rui, et al.
Veröffentlicht: (2025)
von: Qiao, Rui, et al.
Veröffentlicht: (2025)
Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025)
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025)
Fine-tuning Language Models with Generative Adversarial Reward Modelling
von: Yu, Zhang Ze, et al.
Veröffentlicht: (2023)
von: Yu, Zhang Ze, et al.
Veröffentlicht: (2023)
DUET: Optimizing Training Data Mixtures via Feedback from Unseen Evaluation Tasks
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025)
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025)
DETAIL: Task DEmonsTration Attribution for Interpretable In-context Learning
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
von: Tian, Xiao, et al.
Veröffentlicht: (2026)
von: Tian, Xiao, et al.
Veröffentlicht: (2026)
Data Distribution Valuation
von: Xu, Xinyi, et al.
Veröffentlicht: (2024)
von: Xu, Xinyi, et al.
Veröffentlicht: (2024)
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
von: Verma, Arun, et al.
Veröffentlicht: (2024)
von: Verma, Arun, et al.
Veröffentlicht: (2024)
Understanding Domain Generalization: A Noise Robustness Perspective
von: Qiao, Rui, et al.
Veröffentlicht: (2024)
von: Qiao, Rui, et al.
Veröffentlicht: (2024)
REFRAG: Rethinking RAG based Decoding
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2025)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2025)
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2024)
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2024)
Provably Adaptive Linear Approximation for the Shapley Value and Beyond
von: Li, Weida, et al.
Veröffentlicht: (2026)
von: Li, Weida, et al.
Veröffentlicht: (2026)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2025)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2025)
Prompt Optimization with Human Feedback
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2024)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2024)
Data value estimation on private gradients
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
The Chicken and Egg Dilemma: Co-optimizing Data and Model Configurations for LLMs
von: Chen, Zhiliang, et al.
Veröffentlicht: (2026)
von: Chen, Zhiliang, et al.
Veröffentlicht: (2026)
Dependency Structure Search Bayesian Optimization for Decision Making Models
von: Rajpal, Mohit, et al.
Veröffentlicht: (2023)
von: Rajpal, Mohit, et al.
Veröffentlicht: (2023)
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
DeRDaVa: Deletion-Robust Data Valuation for Machine Learning
von: Tian, Xiao, et al.
Veröffentlicht: (2023)
von: Tian, Xiao, et al.
Veröffentlicht: (2023)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
von: Liu, Guozhi, et al.
Veröffentlicht: (2025)
von: Liu, Guozhi, et al.
Veröffentlicht: (2025)
Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2026)
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2026)
Ferret: Federated Full-Parameter Tuning at Scale for Large Language Models
von: Shu, Yao, et al.
Veröffentlicht: (2024)
von: Shu, Yao, et al.
Veröffentlicht: (2024)
Paid with Models: Optimal Contract Design for Collaborative Machine Learning
von: Wang, Bingchen, et al.
Veröffentlicht: (2024)
von: Wang, Bingchen, et al.
Veröffentlicht: (2024)
Self-Interested Agents in Collaborative Machine Learning: An Incentivized Adaptive Data-Centric Framework
von: Vijayan, Nithia, et al.
Veröffentlicht: (2024)
von: Vijayan, Nithia, et al.
Veröffentlicht: (2024)
Uncovering Scaling Laws for Large Language Models via Inverse Problems
von: Verma, Arun, et al.
Veröffentlicht: (2025)
von: Verma, Arun, et al.
Veröffentlicht: (2025)
Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies
von: Verma, Arun, et al.
Veröffentlicht: (2024)
von: Verma, Arun, et al.
Veröffentlicht: (2024)
COBRA: Contextual Bandit Algorithm for Ensuring Truthful Strategic Agents
von: Verma, Arun, et al.
Veröffentlicht: (2025)
von: Verma, Arun, et al.
Veröffentlicht: (2025)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
Global-to-Local Support Spectrums for Language Model Explainability
von: Agussurja, Lucas, et al.
Veröffentlicht: (2024)
von: Agussurja, Lucas, et al.
Veröffentlicht: (2024)
Fine-tuning can Help Detect Pretraining Data from Large Language Models
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2024)
Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with Transformers
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2023)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2023)
Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
Active Human Feedback Collection via Neural Contextual Dueling Bandits
von: Verma, Arun, et al.
Veröffentlicht: (2025)
von: Verma, Arun, et al.
Veröffentlicht: (2025)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
BoLT: A Benchmark to Democratize Black-box Optimization Research for Expensive LLM Tasks
von: Chew, Ruth Wan Theng, et al.
Veröffentlicht: (2026)
von: Chew, Ruth Wan Theng, et al.
Veröffentlicht: (2026)
INO-SGD: Addressing Utility Imbalance under Individualized Differential Privacy
von: Tian, Xiao, et al.
Veröffentlicht: (2026)
von: Tian, Xiao, et al.
Veröffentlicht: (2026)
Decentralized Sum-of-Nonconvex Optimization
von: Liu, Zhuanghua, et al.
Veröffentlicht: (2024)
von: Liu, Zhuanghua, et al.
Veröffentlicht: (2024)
TRACE: TRansformer-based Attribution using Contrastive Embeddings in LLMs
von: Wang, Cheng, et al.
Veröffentlicht: (2024)
von: Wang, Cheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Source Attribution for Large Language Model-Generated Data
von: Wang, Jingtan, et al.
Veröffentlicht: (2023) -
Group-robust Sample Reweighting for Subpopulation Shifts via Influence Functions
von: Qiao, Rui, et al.
Veröffentlicht: (2025) -
Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025) -
Fine-tuning Language Models with Generative Adversarial Reward Modelling
von: Yu, Zhang Ze, et al.
Veröffentlicht: (2023) -
DUET: Optimizing Training Data Mixtures via Feedback from Unseen Evaluation Tasks
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025)