Current state of LLM Risks and AI Guardrails
Fuente:
arXiv
Guardado en:
| Autores principales: | Ayyamperumal, Suriya Ganesh, Ge, Limin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
por: Wu, Liangxuan, et al.
Publicado: (2025)
por: Wu, Liangxuan, et al.
Publicado: (2025)
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
por: Dassanayake, Rishane, et al.
Publicado: (2025)
por: Dassanayake, Rishane, et al.
Publicado: (2025)
"It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents
por: Zhang, Zhiping, et al.
Publicado: (2023)
por: Zhang, Zhiping, et al.
Publicado: (2023)
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots
por: Zhou, Jijie, et al.
Publicado: (2024)
por: Zhou, Jijie, et al.
Publicado: (2024)
"Think First, Verify Always": Training Humans to Face AI Risks
por: Aydin, Yuksel
Publicado: (2025)
por: Aydin, Yuksel
Publicado: (2025)
Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues
por: Chang, Zhiyuan, et al.
Publicado: (2024)
por: Chang, Zhiyuan, et al.
Publicado: (2024)
Empowering Users in Digital Privacy Management through Interactive LLM-Based Agents
por: Sun, Bolun, et al.
Publicado: (2024)
por: Sun, Bolun, et al.
Publicado: (2024)
Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents
por: Zhang, Zhiping, et al.
Publicado: (2025)
por: Zhang, Zhiping, et al.
Publicado: (2025)
Towards Secure AI-driven Industrial Metaverse with NFT Digital Twins
por: Prakash, Ravi, et al.
Publicado: (2024)
por: Prakash, Ravi, et al.
Publicado: (2024)
Personalised Feedback Framework for Online Education Programmes Using Generative AI
por: Kuzminykh, Ievgeniia, et al.
Publicado: (2024)
por: Kuzminykh, Ievgeniia, et al.
Publicado: (2024)
AI-Assisted Adaptive Rendering for High-Frequency Security Telemetry in Web Interfaces
por: Rajhans, Mona
Publicado: (2026)
por: Rajhans, Mona
Publicado: (2026)
Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
por: Zhan, Xiao, et al.
Publicado: (2025)
por: Zhan, Xiao, et al.
Publicado: (2025)
Agentic AI and the Industrialization of Cyber Offense: Forecast, Consequences, and Defensive Priorities for Enterprises and the Mittelstand
por: Koch, Christopher
Publicado: (2026)
por: Koch, Christopher
Publicado: (2026)
Human-AI Collaboration in Cloud Security: Cognitive Hierarchy-Driven Deep Reinforcement Learning
por: Aref, Zahra, et al.
Publicado: (2025)
por: Aref, Zahra, et al.
Publicado: (2025)
BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos
por: Lin, Lehao, et al.
Publicado: (2025)
por: Lin, Lehao, et al.
Publicado: (2025)
Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent
por: Zhang, Zhiping, et al.
Publicado: (2024)
por: Zhang, Zhiping, et al.
Publicado: (2024)
PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy
por: Kundu, Ripan Kumar, et al.
Publicado: (2025)
por: Kundu, Ripan Kumar, et al.
Publicado: (2025)
Cyri: A Conversational AI-based Assistant for Supporting the Human User in Detecting and Responding to Phishing Attacks
por: La Torre, Antonio, et al.
Publicado: (2025)
por: La Torre, Antonio, et al.
Publicado: (2025)
Human-Centered Explainability in AI-Enhanced UI Security Interfaces: Designing Trustworthy Copilots for Cybersecurity Analysts
por: Rajhans, Mona
Publicado: (2026)
por: Rajhans, Mona
Publicado: (2026)
Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions
por: Prakash, Vijay, et al.
Publicado: (2024)
por: Prakash, Vijay, et al.
Publicado: (2024)
OpenAI's Approach to External Red Teaming for AI Models and Systems
por: Ahmad, Lama, et al.
Publicado: (2025)
por: Ahmad, Lama, et al.
Publicado: (2025)
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis
por: Esposito, Matteo, et al.
Publicado: (2024)
por: Esposito, Matteo, et al.
Publicado: (2024)
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
por: Lin, Weiran, et al.
Publicado: (2024)
por: Lin, Weiran, et al.
Publicado: (2024)
When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs
por: Chen, Baiyu, et al.
Publicado: (2025)
por: Chen, Baiyu, et al.
Publicado: (2025)
Human-Centered Privacy Research in the Age of Large Language Models
por: Li, Tianshi, et al.
Publicado: (2024)
por: Li, Tianshi, et al.
Publicado: (2024)
SECURE: Benchmarking Large Language Models for Cybersecurity
por: Bhusal, Dipkamal, et al.
Publicado: (2024)
por: Bhusal, Dipkamal, et al.
Publicado: (2024)
Adversarial VR: An Open-Source Testbed for Evaluating Adversarial Robustness of VR Cybersickness Detection and Mitigation
por: Ahmed, Istiak, et al.
Publicado: (2025)
por: Ahmed, Istiak, et al.
Publicado: (2025)
Decision-Aware Trust Signal Alignment for SOC Alert Triage
por: Chowdhury, Israt Jahan, et al.
Publicado: (2026)
por: Chowdhury, Israt Jahan, et al.
Publicado: (2026)
"Impressively Scary:" Exploring User Perceptions and Reactions to Unraveling Machine Learning Models in Social Media Applications
por: West, Jack, et al.
Publicado: (2025)
por: West, Jack, et al.
Publicado: (2025)
MeAJOR Corpus: A Multi-Source Dataset for Phishing Email Detection
por: Mendes, Paulo, et al.
Publicado: (2025)
por: Mendes, Paulo, et al.
Publicado: (2025)
JEEVHITAA -- An End-to-End HCAI System to Support Collective Care
por: Srinivasan, Shyama Sastha Krishnamoorthy, et al.
Publicado: (2025)
por: Srinivasan, Shyama Sastha Krishnamoorthy, et al.
Publicado: (2025)
InjectLab: A Tactical Framework for Adversarial Threat Modeling Against Large Language Models
por: Howard, Austin
Publicado: (2025)
por: Howard, Austin
Publicado: (2025)
What Security and Privacy Transparency Users Need from Consumer-Facing Generative AI
por: Cao, Jiaxun, et al.
Publicado: (2026)
por: Cao, Jiaxun, et al.
Publicado: (2026)
PRISM: A Personalized, Rapid, and Immersive Skill Mastery framework for personalizing experiential learning through Generative AI
por: Lin, Yu-Zheng, et al.
Publicado: (2024)
por: Lin, Yu-Zheng, et al.
Publicado: (2024)
Securing the AI Supply Chain: What Can We Learn From Developer-Reported Security Issues and Solutions of AI Projects?
por: Nguyen, The Anh, et al.
Publicado: (2025)
por: Nguyen, The Anh, et al.
Publicado: (2025)
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
por: Lugoloobi, William, et al.
Publicado: (2026)
por: Lugoloobi, William, et al.
Publicado: (2026)
A Human-Centered Privacy Approach (HCP) to AI
por: Sun, Luyi, et al.
Publicado: (2026)
por: Sun, Luyi, et al.
Publicado: (2026)
Towards Automating Data Access Permissions in AI Agents
por: Wu, Yuhao, et al.
Publicado: (2025)
por: Wu, Yuhao, et al.
Publicado: (2025)
Generative AI-based closed-loop fMRI system
por: Kasahara, Mikihiro, et al.
Publicado: (2024)
por: Kasahara, Mikihiro, et al.
Publicado: (2024)
Securing Virtual Reality Experiences: Unveiling and Tackling Cybersickness Attacks with Explainable AI
por: Kundu, Ripan Kumar, et al.
Publicado: (2025)
por: Kundu, Ripan Kumar, et al.
Publicado: (2025)
Ejemplares similares
-
From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
por: Wu, Liangxuan, et al.
Publicado: (2025) -
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
por: Dassanayake, Rishane, et al.
Publicado: (2025) -
"It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents
por: Zhang, Zhiping, et al.
Publicado: (2023) -
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots
por: Zhou, Jijie, et al.
Publicado: (2024) -
"Think First, Verify Always": Training Humans to Face AI Risks
por: Aydin, Yuksel
Publicado: (2025)