The Silicon Psyche: Anthropomorphic Vulnerabilities in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Canale, Giuseppe, Thimmaraju, Kashyap |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Stop Testing Attacks, Start Diagnosing Defenses: The Four-Checkpoint Framework Reveals Where LLM Safety Breaks
par: Dhabhi, Hayfa, et autres
Publié: (2026)
par: Dhabhi, Hayfa, et autres
Publié: (2026)
Before the Vicious Cycle Starts: Preventing Burnout Across SOC Roles Through Flow-Aligned Design
par: Thimmaraju, Kashyap, et autres
Publié: (2026)
par: Thimmaraju, Kashyap, et autres
Publié: (2026)
OpenAI's Approach to External Red Teaming for AI Models and Systems
par: Ahmad, Lama, et autres
Publié: (2025)
par: Ahmad, Lama, et autres
Publié: (2025)
SECURE: Benchmarking Large Language Models for Cybersecurity
par: Bhusal, Dipkamal, et autres
Publié: (2024)
par: Bhusal, Dipkamal, et autres
Publié: (2024)
When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs
par: Chen, Baiyu, et autres
Publié: (2025)
par: Chen, Baiyu, et autres
Publié: (2025)
Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
par: Zhan, Xiao, et autres
Publié: (2025)
par: Zhan, Xiao, et autres
Publié: (2025)
Personhood Credentials: Human-Centered Design Recommendation Balancing Security, Usability, and Trust
par: Ide, Ayae, et autres
Publié: (2025)
par: Ide, Ayae, et autres
Publié: (2025)
"Think First, Verify Always": Training Humans to Face AI Risks
par: Aydin, Yuksel
Publié: (2025)
par: Aydin, Yuksel
Publié: (2025)
To Patch or Not to Patch: Motivations, Challenges, and Implications for Cybersecurity
par: Nurse, Jason R. C.
Publié: (2025)
par: Nurse, Jason R. C.
Publié: (2025)
PRISM: A Personalized, Rapid, and Immersive Skill Mastery framework for personalizing experiential learning through Generative AI
par: Lin, Yu-Zheng, et autres
Publié: (2024)
par: Lin, Yu-Zheng, et autres
Publié: (2024)
What Security and Privacy Transparency Users Need from Consumer-Facing Generative AI
par: Cao, Jiaxun, et autres
Publié: (2026)
par: Cao, Jiaxun, et autres
Publié: (2026)
Human-Centered Privacy Research in the Age of Large Language Models
par: Li, Tianshi, et autres
Publié: (2024)
par: Li, Tianshi, et autres
Publié: (2024)
InjectLab: A Tactical Framework for Adversarial Threat Modeling Against Large Language Models
par: Howard, Austin
Publié: (2025)
par: Howard, Austin
Publié: (2025)
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis
par: Esposito, Matteo, et autres
Publié: (2024)
par: Esposito, Matteo, et autres
Publié: (2024)
On the Suitability of LLM-Driven Agents for Dark Pattern Audits
par: Sun, Chen, et autres
Publié: (2026)
par: Sun, Chen, et autres
Publié: (2026)
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
par: Zhang, Chen Bo Calvin, et autres
Publié: (2026)
par: Zhang, Chen Bo Calvin, et autres
Publié: (2026)
Assessing LLM Response Quality in the Context of Technology-Facilitated Abuse
par: Prakash, Vijay, et autres
Publié: (2026)
par: Prakash, Vijay, et autres
Publié: (2026)
NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey
par: Goswami, Dhiman, et autres
Publié: (2026)
par: Goswami, Dhiman, et autres
Publié: (2026)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
par: Wang, Jiongxiao, et autres
Publié: (2023)
par: Wang, Jiongxiao, et autres
Publié: (2023)
Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent
par: Zhang, Zhiping, et autres
Publié: (2024)
par: Zhang, Zhiping, et autres
Publié: (2024)
Intolerable Risk Threshold Recommendations for Artificial Intelligence
par: Raman, Deepika, et autres
Publié: (2025)
par: Raman, Deepika, et autres
Publié: (2025)
"Is it always watching? Is it always listening?" Exploring Contextual Privacy and Security Concerns Toward Domestic Social Robots
par: Bell, Henry, et autres
Publié: (2025)
par: Bell, Henry, et autres
Publié: (2025)
Enabling Cyber Security Education through Digital Twins and Generative AI
par: Barletta, Vita Santa, et autres
Publié: (2025)
par: Barletta, Vita Santa, et autres
Publié: (2025)
Exploring User Security and Privacy Attitudes and Concerns Toward the Use of General-Purpose LLM Chatbots for Mental Health
par: Kwesi, Jabari, et autres
Publié: (2025)
par: Kwesi, Jabari, et autres
Publié: (2025)
V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data
par: Sharma, Tanusree, et autres
Publié: (2026)
par: Sharma, Tanusree, et autres
Publié: (2026)
Impacts of Anthropomorphizing Large Language Models in Learning Environments
par: Schaaff, Kristina, et autres
Publié: (2024)
par: Schaaff, Kristina, et autres
Publié: (2024)
"Impressively Scary:" Exploring User Perceptions and Reactions to Unraveling Machine Learning Models in Social Media Applications
par: West, Jack, et autres
Publié: (2025)
par: West, Jack, et autres
Publié: (2025)
"These cameras are just like the Eye of Sauron": A Sociotechnical Threat Model for AI-Driven Smart Home Devices as Perceived by UK-Based Domestic Workers
par: He, Shijing, et autres
Publié: (2026)
par: He, Shijing, et autres
Publié: (2026)
BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos
par: Lin, Lehao, et autres
Publié: (2025)
par: Lin, Lehao, et autres
Publié: (2025)
PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy
par: Kundu, Ripan Kumar, et autres
Publié: (2025)
par: Kundu, Ripan Kumar, et autres
Publié: (2025)
Adversarial VR: An Open-Source Testbed for Evaluating Adversarial Robustness of VR Cybersickness Detection and Mitigation
par: Ahmed, Istiak, et autres
Publié: (2025)
par: Ahmed, Istiak, et autres
Publié: (2025)
MeAJOR Corpus: A Multi-Source Dataset for Phishing Email Detection
par: Mendes, Paulo, et autres
Publié: (2025)
par: Mendes, Paulo, et autres
Publié: (2025)
Cyri: A Conversational AI-based Assistant for Supporting the Human User in Detecting and Responding to Phishing Attacks
par: La Torre, Antonio, et autres
Publié: (2025)
par: La Torre, Antonio, et autres
Publié: (2025)
Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents
par: Zhang, Zhiping, et autres
Publié: (2025)
par: Zhang, Zhiping, et autres
Publié: (2025)
JEEVHITAA -- An End-to-End HCAI System to Support Collective Care
par: Srinivasan, Shyama Sastha Krishnamoorthy, et autres
Publié: (2025)
par: Srinivasan, Shyama Sastha Krishnamoorthy, et autres
Publié: (2025)
Human-AI Collaboration in Cloud Security: Cognitive Hierarchy-Driven Deep Reinforcement Learning
par: Aref, Zahra, et autres
Publié: (2025)
par: Aref, Zahra, et autres
Publié: (2025)
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
par: Dassanayake, Rishane, et autres
Publié: (2025)
par: Dassanayake, Rishane, et autres
Publié: (2025)
From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
par: Wu, Liangxuan, et autres
Publié: (2025)
par: Wu, Liangxuan, et autres
Publié: (2025)
Agentic AI and the Industrialization of Cyber Offense: Forecast, Consequences, and Defensive Priorities for Enterprises and the Mittelstand
par: Koch, Christopher
Publié: (2026)
par: Koch, Christopher
Publié: (2026)
Decision-Aware Trust Signal Alignment for SOC Alert Triage
par: Chowdhury, Israt Jahan, et autres
Publié: (2026)
par: Chowdhury, Israt Jahan, et autres
Publié: (2026)
Documents similaires
-
Stop Testing Attacks, Start Diagnosing Defenses: The Four-Checkpoint Framework Reveals Where LLM Safety Breaks
par: Dhabhi, Hayfa, et autres
Publié: (2026) -
Before the Vicious Cycle Starts: Preventing Burnout Across SOC Roles Through Flow-Aligned Design
par: Thimmaraju, Kashyap, et autres
Publié: (2026) -
OpenAI's Approach to External Red Teaming for AI Models and Systems
par: Ahmad, Lama, et autres
Publié: (2025) -
SECURE: Benchmarking Large Language Models for Cybersecurity
par: Bhusal, Dipkamal, et autres
Publié: (2024) -
When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs
par: Chen, Baiyu, et autres
Publié: (2025)