Salvato in:
| Autori principali: | Verma, Rakesh M., Dershowitz, Nachum, Zeng, Victor, Boumber, Dainis, Liu, Xuting |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2402.01019 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
di: Boumber, Dainis, et al.
Pubblicazione: (2024)
di: Boumber, Dainis, et al.
Pubblicazione: (2024)
The Pitfalls of Publishing in the Age of LLMs: Strange and Surprising Adventures with a High-Impact NLP Journal
di: Verma, Rakesh M., et al.
Pubblicazione: (2024)
di: Verma, Rakesh M., et al.
Pubblicazione: (2024)
Fake News, Disinformation, and Deepfakes: Leveraging Distributed Ledger Technologies and Blockchain to Combat Digital Deception and Counterfeit Reality
di: Fraga-Lamas, Paula, et al.
Pubblicazione: (2019)
di: Fraga-Lamas, Paula, et al.
Pubblicazione: (2019)
Let's Measure the Elephant in the Room: Facilitating Personalized Automated Analysis of Privacy Policies at Scale
di: Zhao, Rui, et al.
Pubblicazione: (2025)
di: Zhao, Rui, et al.
Pubblicazione: (2025)
Homograph Attacks on Maghreb Sentiment Analyzers
di: Qachfar, Fatima Zahra, et al.
Pubblicazione: (2024)
di: Qachfar, Fatima Zahra, et al.
Pubblicazione: (2024)
Access Over Deception: Fighting Deceptive Patterns through Accessibility
di: Pellkvist, Tobias, et al.
Pubblicazione: (2026)
di: Pellkvist, Tobias, et al.
Pubblicazione: (2026)
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
di: Liu, Houjun, et al.
Pubblicazione: (2026)
di: Liu, Houjun, et al.
Pubblicazione: (2026)
Privacy Computing Meets Metaverse: Necessity, Taxonomy and Challenges
di: Chen, Chuan, et al.
Pubblicazione: (2023)
di: Chen, Chuan, et al.
Pubblicazione: (2023)
Chameleon Channels: Measuring YouTube Accounts Repurposed for Deception and Profit
di: Cuevas, Alejandro, et al.
Pubblicazione: (2025)
di: Cuevas, Alejandro, et al.
Pubblicazione: (2025)
Provably Secure Disambiguating Neural Linguistic Steganography
di: Qi, Yuang, et al.
Pubblicazione: (2024)
di: Qi, Yuang, et al.
Pubblicazione: (2024)
Honeyquest: Rapidly Measuring the Enticingness of Cyber Deception Techniques with Code-based Questionnaires
di: Kahlhofer, Mario, et al.
Pubblicazione: (2024)
di: Kahlhofer, Mario, et al.
Pubblicazione: (2024)
An Investigation into Misuse of Java Security APIs by Large Language Models
di: Mousavi, Zahra, et al.
Pubblicazione: (2024)
di: Mousavi, Zahra, et al.
Pubblicazione: (2024)
Data Defenses Against Large Language Models
di: Agnew, William, et al.
Pubblicazione: (2024)
di: Agnew, William, et al.
Pubblicazione: (2024)
How Susceptible are Large Language Models to Ideological Manipulation?
di: Chen, Kai, et al.
Pubblicazione: (2024)
di: Chen, Kai, et al.
Pubblicazione: (2024)
Get my drift? Catching LLM Task Drift with Activation Deltas
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2024)
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2024)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
di: Gomaa, Amr, et al.
Pubblicazione: (2025)
di: Gomaa, Amr, et al.
Pubblicazione: (2025)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
di: Cyberey, Hannah, et al.
Pubblicazione: (2025)
di: Cyberey, Hannah, et al.
Pubblicazione: (2025)
AI Agents May Always Fall for Prompt Injections
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2026)
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2026)
A Content-Preserving Secure Linguistic Steganography
di: Xiang, Lingyun, et al.
Pubblicazione: (2025)
di: Xiang, Lingyun, et al.
Pubblicazione: (2025)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
di: Wang, Peiran, et al.
Pubblicazione: (2026)
di: Wang, Peiran, et al.
Pubblicazione: (2026)
RealHarm: A Collection of Real-World Language Model Application Failures
di: Jeune, Pierre Le, et al.
Pubblicazione: (2025)
di: Jeune, Pierre Le, et al.
Pubblicazione: (2025)
Hardware-Level Governance of AI Compute: A Feasibility Taxonomy for Regulatory Compliance and Treaty Verification
di: Ansari, Samar
Pubblicazione: (2026)
di: Ansari, Samar
Pubblicazione: (2026)
Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing
di: Schmitt, Marc, et al.
Pubblicazione: (2023)
di: Schmitt, Marc, et al.
Pubblicazione: (2023)
Beyond Context: Large Language Models' Failure to Grasp Users' Intent
di: Hussain, Ahmed M., et al.
Pubblicazione: (2025)
di: Hussain, Ahmed M., et al.
Pubblicazione: (2025)
TombRaider: Entering the Vault of History to Jailbreak Large Language Models
di: Ding, Junchen, et al.
Pubblicazione: (2025)
di: Ding, Junchen, et al.
Pubblicazione: (2025)
Phare: A Safety Probe for Large Language Models
di: Jeune, Pierre Le, et al.
Pubblicazione: (2025)
di: Jeune, Pierre Le, et al.
Pubblicazione: (2025)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
di: Rassul, Yassin H., et al.
Pubblicazione: (2026)
di: Rassul, Yassin H., et al.
Pubblicazione: (2026)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
di: D'addario, Andrew Maranhão Ventura
Pubblicazione: (2025)
di: D'addario, Andrew Maranhão Ventura
Pubblicazione: (2025)
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
di: Wang, Huandong, et al.
Pubblicazione: (2025)
di: Wang, Huandong, et al.
Pubblicazione: (2025)
k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text
di: Hou, Abe Bohan, et al.
Pubblicazione: (2024)
di: Hou, Abe Bohan, et al.
Pubblicazione: (2024)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
di: Djuhera, Aladin, et al.
Pubblicazione: (2025)
di: Djuhera, Aladin, et al.
Pubblicazione: (2025)
SimMark: A Robust Sentence-Level Similarity-Based Watermarking Algorithm for Large Language Models
di: Dabiriaghdam, Amirhossein, et al.
Pubblicazione: (2025)
di: Dabiriaghdam, Amirhossein, et al.
Pubblicazione: (2025)
The TCF doesn't really A(A)ID -- Automatic Privacy Analysis and Legal Compliance of TCF-based Android Applications
di: Morel, Victor, et al.
Pubblicazione: (2026)
di: Morel, Victor, et al.
Pubblicazione: (2026)
Zero-shot Generative Linguistic Steganography
di: Lin, Ke, et al.
Pubblicazione: (2024)
di: Lin, Ke, et al.
Pubblicazione: (2024)
Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks
di: Iyer, Karthik Raghu, et al.
Pubblicazione: (2026)
di: Iyer, Karthik Raghu, et al.
Pubblicazione: (2026)
AuditGPT: Auditing Smart Contracts with ChatGPT
di: Xia, Shihao, et al.
Pubblicazione: (2024)
di: Xia, Shihao, et al.
Pubblicazione: (2024)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
Attacks on Third-Party APIs of Large Language Models
di: Zhao, Wanru, et al.
Pubblicazione: (2024)
di: Zhao, Wanru, et al.
Pubblicazione: (2024)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
di: Xue, Jiaqi, et al.
Pubblicazione: (2024)
di: Xue, Jiaqi, et al.
Pubblicazione: (2024)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
di: An, Bang, et al.
Pubblicazione: (2024)
di: An, Bang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
di: Boumber, Dainis, et al.
Pubblicazione: (2024) -
The Pitfalls of Publishing in the Age of LLMs: Strange and Surprising Adventures with a High-Impact NLP Journal
di: Verma, Rakesh M., et al.
Pubblicazione: (2024) -
Fake News, Disinformation, and Deepfakes: Leveraging Distributed Ledger Technologies and Blockchain to Combat Digital Deception and Counterfeit Reality
di: Fraga-Lamas, Paula, et al.
Pubblicazione: (2019) -
Let's Measure the Elephant in the Room: Facilitating Personalized Automated Analysis of Privacy Policies at Scale
di: Zhao, Rui, et al.
Pubblicazione: (2025) -
Homograph Attacks on Maghreb Sentiment Analyzers
di: Qachfar, Fatima Zahra, et al.
Pubblicazione: (2024)