"Think First, Verify Always": Training Humans to Face AI Risks
Fuente:
arXiv
Saved in:
| Main Author: | Aydin, Yuksel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Security and Privacy Transparency Users Need from Consumer-Facing Generative AI
by: Cao, Jiaxun, et al.
Published: (2026)
by: Cao, Jiaxun, et al.
Published: (2026)
OpenAI's Approach to External Red Teaming for AI Models and Systems
by: Ahmad, Lama, et al.
Published: (2025)
by: Ahmad, Lama, et al.
Published: (2025)
When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs
by: Chen, Baiyu, et al.
Published: (2025)
by: Chen, Baiyu, et al.
Published: (2025)
Personhood Credentials: Human-Centered Design Recommendation Balancing Security, Usability, and Trust
by: Ide, Ayae, et al.
Published: (2025)
by: Ide, Ayae, et al.
Published: (2025)
Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
by: Zhan, Xiao, et al.
Published: (2025)
by: Zhan, Xiao, et al.
Published: (2025)
PRISM: A Personalized, Rapid, and Immersive Skill Mastery framework for personalizing experiential learning through Generative AI
by: Lin, Yu-Zheng, et al.
Published: (2024)
by: Lin, Yu-Zheng, et al.
Published: (2024)
NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey
by: Goswami, Dhiman, et al.
Published: (2026)
by: Goswami, Dhiman, et al.
Published: (2026)
Current state of LLM Risks and AI Guardrails
by: Ayyamperumal, Suriya Ganesh, et al.
Published: (2024)
by: Ayyamperumal, Suriya Ganesh, et al.
Published: (2024)
The Silicon Psyche: Anthropomorphic Vulnerabilities in Large Language Models
by: Canale, Giuseppe, et al.
Published: (2025)
by: Canale, Giuseppe, et al.
Published: (2025)
To Patch or Not to Patch: Motivations, Challenges, and Implications for Cybersecurity
by: Nurse, Jason R. C.
Published: (2025)
by: Nurse, Jason R. C.
Published: (2025)
Intolerable Risk Threshold Recommendations for Artificial Intelligence
by: Raman, Deepika, et al.
Published: (2025)
by: Raman, Deepika, et al.
Published: (2025)
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
by: Dassanayake, Rishane, et al.
Published: (2025)
by: Dassanayake, Rishane, et al.
Published: (2025)
Benchmarking and Understanding Safety Risks in AI Character Platforms
by: Wei, Yiluo, et al.
Published: (2025)
by: Wei, Yiluo, et al.
Published: (2025)
Human-AI Collaboration in Cloud Security: Cognitive Hierarchy-Driven Deep Reinforcement Learning
by: Aref, Zahra, et al.
Published: (2025)
by: Aref, Zahra, et al.
Published: (2025)
Enabling Cyber Security Education through Digital Twins and Generative AI
by: Barletta, Vita Santa, et al.
Published: (2025)
by: Barletta, Vita Santa, et al.
Published: (2025)
TRIP: Coercion-resistant Registration for E-Voting with Verifiability and Usability in Votegral
by: Merino, Louis-Henri, et al.
Published: (2022)
by: Merino, Louis-Henri, et al.
Published: (2022)
V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data
by: Sharma, Tanusree, et al.
Published: (2026)
by: Sharma, Tanusree, et al.
Published: (2026)
Cyri: A Conversational AI-based Assistant for Supporting the Human User in Detecting and Responding to Phishing Attacks
by: La Torre, Antonio, et al.
Published: (2025)
by: La Torre, Antonio, et al.
Published: (2025)
Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent
by: Zhang, Zhiping, et al.
Published: (2024)
by: Zhang, Zhiping, et al.
Published: (2024)
Human-Centered Explainability in AI-Enhanced UI Security Interfaces: Designing Trustworthy Copilots for Cybersecurity Analysts
by: Rajhans, Mona
Published: (2026)
by: Rajhans, Mona
Published: (2026)
Exploring the Privacy and Security Challenges Faced by Migrant Domestic Workers in Chinese Smart Homes
by: He, Shijing, et al.
Published: (2025)
by: He, Shijing, et al.
Published: (2025)
On the Suitability of LLM-Driven Agents for Dark Pattern Audits
by: Sun, Chen, et al.
Published: (2026)
by: Sun, Chen, et al.
Published: (2026)
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
by: Zhang, Chen Bo Calvin, et al.
Published: (2026)
by: Zhang, Chen Bo Calvin, et al.
Published: (2026)
Assessing LLM Response Quality in the Context of Technology-Facilitated Abuse
by: Prakash, Vijay, et al.
Published: (2026)
by: Prakash, Vijay, et al.
Published: (2026)
From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
by: Wu, Liangxuan, et al.
Published: (2025)
by: Wu, Liangxuan, et al.
Published: (2025)
Human-Centered Privacy Research in the Age of Large Language Models
by: Li, Tianshi, et al.
Published: (2024)
by: Li, Tianshi, et al.
Published: (2024)
Towards Secure AI-driven Industrial Metaverse with NFT Digital Twins
by: Prakash, Ravi, et al.
Published: (2024)
by: Prakash, Ravi, et al.
Published: (2024)
Personalised Feedback Framework for Online Education Programmes Using Generative AI
by: Kuzminykh, Ievgeniia, et al.
Published: (2024)
by: Kuzminykh, Ievgeniia, et al.
Published: (2024)
AI-Assisted Adaptive Rendering for High-Frequency Security Telemetry in Web Interfaces
by: Rajhans, Mona
Published: (2026)
by: Rajhans, Mona
Published: (2026)
"It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents
by: Zhang, Zhiping, et al.
Published: (2023)
by: Zhang, Zhiping, et al.
Published: (2023)
"Is it always watching? Is it always listening?" Exploring Contextual Privacy and Security Concerns Toward Domestic Social Robots
by: Bell, Henry, et al.
Published: (2025)
by: Bell, Henry, et al.
Published: (2025)
Exploring User Security and Privacy Attitudes and Concerns Toward the Use of General-Purpose LLM Chatbots for Mental Health
by: Kwesi, Jabari, et al.
Published: (2025)
by: Kwesi, Jabari, et al.
Published: (2025)
Stop Testing Attacks, Start Diagnosing Defenses: The Four-Checkpoint Framework Reveals Where LLM Safety Breaks
by: Dhabhi, Hayfa, et al.
Published: (2026)
by: Dhabhi, Hayfa, et al.
Published: (2026)
Agentic AI and the Industrialization of Cyber Offense: Forecast, Consequences, and Defensive Priorities for Enterprises and the Mittelstand
by: Koch, Christopher
Published: (2026)
by: Koch, Christopher
Published: (2026)
BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos
by: Lin, Lehao, et al.
Published: (2025)
by: Lin, Lehao, et al.
Published: (2025)
PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy
by: Kundu, Ripan Kumar, et al.
Published: (2025)
by: Kundu, Ripan Kumar, et al.
Published: (2025)
Chatting with Confidants or Corporations? Privacy Management with AI Companions
by: Chiu, Hsuen-Chi, et al.
Published: (2026)
by: Chiu, Hsuen-Chi, et al.
Published: (2026)
From Chat Control to Robot Control: Implications of the Chat Control Proposal for Human-Robot Interaction
by: Akalin, Neziha, et al.
Published: (2026)
by: Akalin, Neziha, et al.
Published: (2026)
How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape
by: Kelley, Patrick Gage, et al.
Published: (2025)
by: Kelley, Patrick Gage, et al.
Published: (2025)
Secure human oversight of AI: Threat modeling in a socio-technical context
by: Ditz, Jonas C., et al.
Published: (2025)
by: Ditz, Jonas C., et al.
Published: (2025)
Similar Items
-
What Security and Privacy Transparency Users Need from Consumer-Facing Generative AI
by: Cao, Jiaxun, et al.
Published: (2026) -
OpenAI's Approach to External Red Teaming for AI Models and Systems
by: Ahmad, Lama, et al.
Published: (2025) -
When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs
by: Chen, Baiyu, et al.
Published: (2025) -
Personhood Credentials: Human-Centered Design Recommendation Balancing Security, Usability, and Trust
by: Ide, Ayae, et al.
Published: (2025) -
Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
by: Zhan, Xiao, et al.
Published: (2025)