Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts
Fuente:
arXiv
Saved in:
| Main Authors: | Zarlenga, Mateo Espinosa, Dominici, Gabriele, Barbiero, Pietro, Shams, Zohreh, Jamnik, Mateja |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent
by: Zhang, Zhiping, et al.
Published: (2024)
by: Zhang, Zhiping, et al.
Published: (2024)
Learning to Receive Help: Intervention-Aware Concept Embedding Models
by: Zarlenga, Mateo Espinosa, et al.
Published: (2023)
by: Zarlenga, Mateo Espinosa, et al.
Published: (2023)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
by: Wang, Jiongxiao, et al.
Published: (2023)
by: Wang, Jiongxiao, et al.
Published: (2023)
An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments
by: Yang, Hongjang, et al.
Published: (2026)
by: Yang, Hongjang, et al.
Published: (2026)
BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos
by: Lin, Lehao, et al.
Published: (2025)
by: Lin, Lehao, et al.
Published: (2025)
PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy
by: Kundu, Ripan Kumar, et al.
Published: (2025)
by: Kundu, Ripan Kumar, et al.
Published: (2025)
Adversarial VR: An Open-Source Testbed for Evaluating Adversarial Robustness of VR Cybersickness Detection and Mitigation
by: Ahmed, Istiak, et al.
Published: (2025)
by: Ahmed, Istiak, et al.
Published: (2025)
Agentic AI and the Industrialization of Cyber Offense: Forecast, Consequences, and Defensive Priorities for Enterprises and the Mittelstand
by: Koch, Christopher
Published: (2026)
by: Koch, Christopher
Published: (2026)
Decision-Aware Trust Signal Alignment for SOC Alert Triage
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
Current state of LLM Risks and AI Guardrails
by: Ayyamperumal, Suriya Ganesh, et al.
Published: (2024)
by: Ayyamperumal, Suriya Ganesh, et al.
Published: (2024)
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots
by: Zhou, Jijie, et al.
Published: (2024)
by: Zhou, Jijie, et al.
Published: (2024)
"Impressively Scary:" Exploring User Perceptions and Reactions to Unraveling Machine Learning Models in Social Media Applications
by: West, Jack, et al.
Published: (2025)
by: West, Jack, et al.
Published: (2025)
Empowering Users in Digital Privacy Management through Interactive LLM-Based Agents
by: Sun, Bolun, et al.
Published: (2024)
by: Sun, Bolun, et al.
Published: (2024)
"It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents
by: Zhang, Zhiping, et al.
Published: (2023)
by: Zhang, Zhiping, et al.
Published: (2023)
MeAJOR Corpus: A Multi-Source Dataset for Phishing Email Detection
by: Mendes, Paulo, et al.
Published: (2025)
by: Mendes, Paulo, et al.
Published: (2025)
Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues
by: Chang, Zhiyuan, et al.
Published: (2024)
by: Chang, Zhiyuan, et al.
Published: (2024)
Towards Secure AI-driven Industrial Metaverse with NFT Digital Twins
by: Prakash, Ravi, et al.
Published: (2024)
by: Prakash, Ravi, et al.
Published: (2024)
Cyri: A Conversational AI-based Assistant for Supporting the Human User in Detecting and Responding to Phishing Attacks
by: La Torre, Antonio, et al.
Published: (2025)
by: La Torre, Antonio, et al.
Published: (2025)
Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents
by: Zhang, Zhiping, et al.
Published: (2025)
by: Zhang, Zhiping, et al.
Published: (2025)
Human-Centered Privacy Research in the Age of Large Language Models
by: Li, Tianshi, et al.
Published: (2024)
by: Li, Tianshi, et al.
Published: (2024)
JEEVHITAA -- An End-to-End HCAI System to Support Collective Care
by: Srinivasan, Shyama Sastha Krishnamoorthy, et al.
Published: (2025)
by: Srinivasan, Shyama Sastha Krishnamoorthy, et al.
Published: (2025)
Human-AI Collaboration in Cloud Security: Cognitive Hierarchy-Driven Deep Reinforcement Learning
by: Aref, Zahra, et al.
Published: (2025)
by: Aref, Zahra, et al.
Published: (2025)
AI-Assisted Adaptive Rendering for High-Frequency Security Telemetry in Web Interfaces
by: Rajhans, Mona
Published: (2026)
by: Rajhans, Mona
Published: (2026)
SECURE: Benchmarking Large Language Models for Cybersecurity
by: Bhusal, Dipkamal, et al.
Published: (2024)
by: Bhusal, Dipkamal, et al.
Published: (2024)
InjectLab: A Tactical Framework for Adversarial Threat Modeling Against Large Language Models
by: Howard, Austin
Published: (2025)
by: Howard, Austin
Published: (2025)
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
by: Dassanayake, Rishane, et al.
Published: (2025)
by: Dassanayake, Rishane, et al.
Published: (2025)
Human-Centered Explainability in AI-Enhanced UI Security Interfaces: Designing Trustworthy Copilots for Cybersecurity Analysts
by: Rajhans, Mona
Published: (2026)
by: Rajhans, Mona
Published: (2026)
From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
by: Wu, Liangxuan, et al.
Published: (2025)
by: Wu, Liangxuan, et al.
Published: (2025)
Personalised Feedback Framework for Online Education Programmes Using Generative AI
by: Kuzminykh, Ievgeniia, et al.
Published: (2024)
by: Kuzminykh, Ievgeniia, et al.
Published: (2024)
Hierarchical Concept-based Interpretable Models
by: Hill, Oscar, et al.
Published: (2026)
by: Hill, Oscar, et al.
Published: (2026)
Digging Deeper: Learning Multi-Level Concept Hierarchies
by: Hill, Oscar, et al.
Published: (2026)
by: Hill, Oscar, et al.
Published: (2026)
Foundations of Interpretable Models
by: Barbiero, Pietro, et al.
Published: (2025)
by: Barbiero, Pietro, et al.
Published: (2025)
"Shifting Access Control Left" using Asset and Goal Models
by: Faily, Shamal
Published: (2025)
by: Faily, Shamal
Published: (2025)
Generative AI-based closed-loop fMRI system
by: Kasahara, Mikihiro, et al.
Published: (2024)
by: Kasahara, Mikihiro, et al.
Published: (2024)
Hacc-Man: An Arcade Game for Jailbreaking LLMs
by: Valentim, Matheus, et al.
Published: (2024)
by: Valentim, Matheus, et al.
Published: (2024)
IFTT-PIN: A Self-Calibrating PIN-Entry Method
by: McConkey, Kathryn, et al.
Published: (2024)
by: McConkey, Kathryn, et al.
Published: (2024)
PerOS: Personalized Self-Adapting Operating Systems in the Cloud
by: Hè, Hongyu
Published: (2024)
by: Hè, Hongyu
Published: (2024)
Identify As A Human Does: A Pathfinder of Next-Generation Anti-Cheat Framework for First-Person Shooter Games
by: Zhang, Jiayi, et al.
Published: (2024)
by: Zhang, Jiayi, et al.
Published: (2024)
When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs
by: Chen, Baiyu, et al.
Published: (2025)
by: Chen, Baiyu, et al.
Published: (2025)
Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
by: Zhan, Xiao, et al.
Published: (2025)
by: Zhan, Xiao, et al.
Published: (2025)
Similar Items
-
Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent
by: Zhang, Zhiping, et al.
Published: (2024) -
Learning to Receive Help: Intervention-Aware Concept Embedding Models
by: Zarlenga, Mateo Espinosa, et al.
Published: (2023) -
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
by: Wang, Jiongxiao, et al.
Published: (2023) -
An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments
by: Yang, Hongjang, et al.
Published: (2026) -
BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos
by: Lin, Lehao, et al.
Published: (2025)