Guardado en:
| Autores principales: | Verma, Rakesh M., Dershowitz, Nachum, Zeng, Victor, Boumber, Dainis, Liu, Xuting |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2402.01019 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
por: Boumber, Dainis, et al.
Publicado: (2024)
por: Boumber, Dainis, et al.
Publicado: (2024)
The Pitfalls of Publishing in the Age of LLMs: Strange and Surprising Adventures with a High-Impact NLP Journal
por: Verma, Rakesh M., et al.
Publicado: (2024)
por: Verma, Rakesh M., et al.
Publicado: (2024)
Fake News, Disinformation, and Deepfakes: Leveraging Distributed Ledger Technologies and Blockchain to Combat Digital Deception and Counterfeit Reality
por: Fraga-Lamas, Paula, et al.
Publicado: (2019)
por: Fraga-Lamas, Paula, et al.
Publicado: (2019)
Let's Measure the Elephant in the Room: Facilitating Personalized Automated Analysis of Privacy Policies at Scale
por: Zhao, Rui, et al.
Publicado: (2025)
por: Zhao, Rui, et al.
Publicado: (2025)
Homograph Attacks on Maghreb Sentiment Analyzers
por: Qachfar, Fatima Zahra, et al.
Publicado: (2024)
por: Qachfar, Fatima Zahra, et al.
Publicado: (2024)
Access Over Deception: Fighting Deceptive Patterns through Accessibility
por: Pellkvist, Tobias, et al.
Publicado: (2026)
por: Pellkvist, Tobias, et al.
Publicado: (2026)
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
por: Liu, Houjun, et al.
Publicado: (2026)
por: Liu, Houjun, et al.
Publicado: (2026)
Privacy Computing Meets Metaverse: Necessity, Taxonomy and Challenges
por: Chen, Chuan, et al.
Publicado: (2023)
por: Chen, Chuan, et al.
Publicado: (2023)
Chameleon Channels: Measuring YouTube Accounts Repurposed for Deception and Profit
por: Cuevas, Alejandro, et al.
Publicado: (2025)
por: Cuevas, Alejandro, et al.
Publicado: (2025)
Provably Secure Disambiguating Neural Linguistic Steganography
por: Qi, Yuang, et al.
Publicado: (2024)
por: Qi, Yuang, et al.
Publicado: (2024)
Honeyquest: Rapidly Measuring the Enticingness of Cyber Deception Techniques with Code-based Questionnaires
por: Kahlhofer, Mario, et al.
Publicado: (2024)
por: Kahlhofer, Mario, et al.
Publicado: (2024)
An Investigation into Misuse of Java Security APIs by Large Language Models
por: Mousavi, Zahra, et al.
Publicado: (2024)
por: Mousavi, Zahra, et al.
Publicado: (2024)
Data Defenses Against Large Language Models
por: Agnew, William, et al.
Publicado: (2024)
por: Agnew, William, et al.
Publicado: (2024)
How Susceptible are Large Language Models to Ideological Manipulation?
por: Chen, Kai, et al.
Publicado: (2024)
por: Chen, Kai, et al.
Publicado: (2024)
Get my drift? Catching LLM Task Drift with Activation Deltas
por: Abdelnabi, Sahar, et al.
Publicado: (2024)
por: Abdelnabi, Sahar, et al.
Publicado: (2024)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
por: Gomaa, Amr, et al.
Publicado: (2025)
por: Gomaa, Amr, et al.
Publicado: (2025)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
por: Cyberey, Hannah, et al.
Publicado: (2025)
por: Cyberey, Hannah, et al.
Publicado: (2025)
AI Agents May Always Fall for Prompt Injections
por: Abdelnabi, Sahar, et al.
Publicado: (2026)
por: Abdelnabi, Sahar, et al.
Publicado: (2026)
A Content-Preserving Secure Linguistic Steganography
por: Xiang, Lingyun, et al.
Publicado: (2025)
por: Xiang, Lingyun, et al.
Publicado: (2025)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
por: Wang, Peiran, et al.
Publicado: (2026)
por: Wang, Peiran, et al.
Publicado: (2026)
RealHarm: A Collection of Real-World Language Model Application Failures
por: Jeune, Pierre Le, et al.
Publicado: (2025)
por: Jeune, Pierre Le, et al.
Publicado: (2025)
Hardware-Level Governance of AI Compute: A Feasibility Taxonomy for Regulatory Compliance and Treaty Verification
por: Ansari, Samar
Publicado: (2026)
por: Ansari, Samar
Publicado: (2026)
Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing
por: Schmitt, Marc, et al.
Publicado: (2023)
por: Schmitt, Marc, et al.
Publicado: (2023)
Beyond Context: Large Language Models' Failure to Grasp Users' Intent
por: Hussain, Ahmed M., et al.
Publicado: (2025)
por: Hussain, Ahmed M., et al.
Publicado: (2025)
TombRaider: Entering the Vault of History to Jailbreak Large Language Models
por: Ding, Junchen, et al.
Publicado: (2025)
por: Ding, Junchen, et al.
Publicado: (2025)
Phare: A Safety Probe for Large Language Models
por: Jeune, Pierre Le, et al.
Publicado: (2025)
por: Jeune, Pierre Le, et al.
Publicado: (2025)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
por: Rassul, Yassin H., et al.
Publicado: (2026)
por: Rassul, Yassin H., et al.
Publicado: (2026)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
por: D'addario, Andrew Maranhão Ventura
Publicado: (2025)
por: D'addario, Andrew Maranhão Ventura
Publicado: (2025)
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
por: Wang, Huandong, et al.
Publicado: (2025)
por: Wang, Huandong, et al.
Publicado: (2025)
k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text
por: Hou, Abe Bohan, et al.
Publicado: (2024)
por: Hou, Abe Bohan, et al.
Publicado: (2024)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
por: Djuhera, Aladin, et al.
Publicado: (2025)
por: Djuhera, Aladin, et al.
Publicado: (2025)
SimMark: A Robust Sentence-Level Similarity-Based Watermarking Algorithm for Large Language Models
por: Dabiriaghdam, Amirhossein, et al.
Publicado: (2025)
por: Dabiriaghdam, Amirhossein, et al.
Publicado: (2025)
The TCF doesn't really A(A)ID -- Automatic Privacy Analysis and Legal Compliance of TCF-based Android Applications
por: Morel, Victor, et al.
Publicado: (2026)
por: Morel, Victor, et al.
Publicado: (2026)
Zero-shot Generative Linguistic Steganography
por: Lin, Ke, et al.
Publicado: (2024)
por: Lin, Ke, et al.
Publicado: (2024)
Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks
por: Iyer, Karthik Raghu, et al.
Publicado: (2026)
por: Iyer, Karthik Raghu, et al.
Publicado: (2026)
AuditGPT: Auditing Smart Contracts with ChatGPT
por: Xia, Shihao, et al.
Publicado: (2024)
por: Xia, Shihao, et al.
Publicado: (2024)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
por: Zhou, Zhenhong, et al.
Publicado: (2024)
por: Zhou, Zhenhong, et al.
Publicado: (2024)
Attacks on Third-Party APIs of Large Language Models
por: Zhao, Wanru, et al.
Publicado: (2024)
por: Zhao, Wanru, et al.
Publicado: (2024)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
por: Xue, Jiaqi, et al.
Publicado: (2024)
por: Xue, Jiaqi, et al.
Publicado: (2024)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
por: An, Bang, et al.
Publicado: (2024)
por: An, Bang, et al.
Publicado: (2024)
Ejemplares similares
-
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
por: Boumber, Dainis, et al.
Publicado: (2024) -
The Pitfalls of Publishing in the Age of LLMs: Strange and Surprising Adventures with a High-Impact NLP Journal
por: Verma, Rakesh M., et al.
Publicado: (2024) -
Fake News, Disinformation, and Deepfakes: Leveraging Distributed Ledger Technologies and Blockchain to Combat Digital Deception and Counterfeit Reality
por: Fraga-Lamas, Paula, et al.
Publicado: (2019) -
Let's Measure the Elephant in the Room: Facilitating Personalized Automated Analysis of Privacy Policies at Scale
por: Zhao, Rui, et al.
Publicado: (2025) -
Homograph Attacks on Maghreb Sentiment Analyzers
por: Qachfar, Fatima Zahra, et al.
Publicado: (2024)