Trustworthy AI: Safety, Bias, and Privacy -- A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fang, Xingli, Li, Jianwei, Mulchandani, Varun, Kim, Jung-Eun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
Decoupling Generalizability and Membership Privacy Risks in Neural Networks
von: Fang, Xingli, et al.
Veröffentlicht: (2026)
von: Fang, Xingli, et al.
Veröffentlicht: (2026)
Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights
von: Fang, Xingli, et al.
Veröffentlicht: (2026)
von: Fang, Xingli, et al.
Veröffentlicht: (2026)
Superficial Safety Alignment Hypothesis
von: Li, Jianwei, et al.
Veröffentlicht: (2024)
von: Li, Jianwei, et al.
Veröffentlicht: (2024)
Representation Magnitude has a Liability to Privacy Vulnerability
von: Fang, Xingli, et al.
Veröffentlicht: (2024)
von: Fang, Xingli, et al.
Veröffentlicht: (2024)
Center-Based Relaxed Learning Against Membership Inference Attacks
von: Fang, Xingli, et al.
Veröffentlicht: (2024)
von: Fang, Xingli, et al.
Veröffentlicht: (2024)
Position: Retire the "Positive Backdoor" Label -- Secret Alignment Requires Strict and Systematic Evaluation
von: Li, Jianwei, et al.
Veröffentlicht: (2026)
von: Li, Jianwei, et al.
Veröffentlicht: (2026)
Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference
von: Li, Jianwei, et al.
Veröffentlicht: (2026)
von: Li, Jianwei, et al.
Veröffentlicht: (2026)
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
PBa-LLM: Privacy- and Bias-aware NLP using Named-Entity Recognition (NER)
von: Mancera, Gonzalo, et al.
Veröffentlicht: (2025)
von: Mancera, Gonzalo, et al.
Veröffentlicht: (2025)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
von: Wang, Kun, et al.
Veröffentlicht: (2025)
von: Wang, Kun, et al.
Veröffentlicht: (2025)
Preserving Privacy in Large Language Models: A Survey on Current Threats and Solutions
von: Miranda, Michele, et al.
Veröffentlicht: (2024)
von: Miranda, Michele, et al.
Veröffentlicht: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
Intent Laundering: AI Safety Datasets Are Not What They Seem
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
DP-MemArc: Differential Privacy Transfer Learning for Memory Efficient Language Models
von: Liu, Yanming, et al.
Veröffentlicht: (2024)
von: Liu, Yanming, et al.
Veröffentlicht: (2024)
Position: Privacy Is Not Just Memorization!
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2025)
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2025)
On the Role of Attention Heads in Large Language Model Safety
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
von: Parmar, Manojkumar, et al.
Veröffentlicht: (2025)
von: Parmar, Manojkumar, et al.
Veröffentlicht: (2025)
Do Phone-Use Agents Respect Your Privacy?
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
On the Trustworthiness Landscape of State-of-the-art Generative Models: A Survey and Outlook
von: Fan, Mingyuan, et al.
Veröffentlicht: (2023)
von: Fan, Mingyuan, et al.
Veröffentlicht: (2023)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
Fine-Tuning Language Models with Differential Privacy through Adaptive Noise Allocation
von: Li, Xianzhi, et al.
Veröffentlicht: (2024)
von: Li, Xianzhi, et al.
Veröffentlicht: (2024)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
von: Li, Lijun, et al.
Veröffentlicht: (2024)
von: Li, Lijun, et al.
Veröffentlicht: (2024)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
Clio: Privacy-Preserving Insights into Real-World AI Use
von: Tamkin, Alex, et al.
Veröffentlicht: (2024)
von: Tamkin, Alex, et al.
Veröffentlicht: (2024)
Learnable Privacy Neurons Localization in Language Models
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
Lifelong Safety Alignment for Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
von: Kan, Chun Yan Ryan, et al.
Veröffentlicht: (2026)
von: Kan, Chun Yan Ryan, et al.
Veröffentlicht: (2026)
How Private is Your Attention? Bridging Privacy with In-Context Learning
von: Bonnerjee, Soham, et al.
Veröffentlicht: (2025)
von: Bonnerjee, Soham, et al.
Veröffentlicht: (2025)
HARMONIC: Harnessing LLMs for Tabular Data Synthesis and Privacy Protection
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
Trustworthy Distributed AI Systems: Robustness, Privacy, and Governance
von: Wei, Wenqi, et al.
Veröffentlicht: (2024)
von: Wei, Wenqi, et al.
Veröffentlicht: (2024)
Bridging Privacy and Robustness for Trustworthy Machine Learning
von: Zhang, Xiaojin, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaojin, et al.
Veröffentlicht: (2024)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
von: Vega, Jason, et al.
Veröffentlicht: (2023)
von: Vega, Jason, et al.
Veröffentlicht: (2023)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
von: Xu, Chejian, et al.
Veröffentlicht: (2025)
von: Xu, Chejian, et al.
Veröffentlicht: (2025)
FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
von: Sarwar, Nobin, et al.
Veröffentlicht: (2025)
von: Sarwar, Nobin, et al.
Veröffentlicht: (2025)
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
von: Li, Jianwei, et al.
Veröffentlicht: (2025) -
Decoupling Generalizability and Membership Privacy Risks in Neural Networks
von: Fang, Xingli, et al.
Veröffentlicht: (2026) -
Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights
von: Fang, Xingli, et al.
Veröffentlicht: (2026) -
Superficial Safety Alignment Hypothesis
von: Li, Jianwei, et al.
Veröffentlicht: (2024) -
Representation Magnitude has a Liability to Privacy Vulnerability
von: Fang, Xingli, et al.
Veröffentlicht: (2024)