Exploring the Adversarial Capabilities of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Struppek, Lukas, Le, Minh Hieu, Hintersdorf, Dominik, Kersting, Kristian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Finding NeMo: Localizing Neurons Responsible For Memorization in Diffusion Models
by: Hintersdorf, Dominik, et al.
Published: (2024)
by: Hintersdorf, Dominik, et al.
Published: (2024)
Be Careful What You Smooth For: Label Smoothing Can Be a Privacy Shield but Also a Catalyst for Model Inversion Attacks
by: Struppek, Lukas, et al.
Published: (2023)
by: Struppek, Lukas, et al.
Published: (2023)
Finding DoRI: Discovery of Retained Images in Diffusion Models
by: Kowalczuk, Antoni, et al.
Published: (2025)
by: Kowalczuk, Antoni, et al.
Published: (2025)
Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
by: Struppek, Lukas, et al.
Published: (2025)
by: Struppek, Lukas, et al.
Published: (2025)
Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
by: Struppek, Lukas, et al.
Published: (2022)
by: Struppek, Lukas, et al.
Published: (2022)
Learning to Break Deep Perceptual Hashing: The Use Case NeuralHash
by: Struppek, Lukas, et al.
Published: (2021)
by: Struppek, Lukas, et al.
Published: (2021)
Defending Our Privacy With Backdoors
by: Hintersdorf, Dominik, et al.
Published: (2023)
by: Hintersdorf, Dominik, et al.
Published: (2023)
Does CLIP Know My Face?
by: Hintersdorf, Dominik, et al.
Published: (2022)
by: Hintersdorf, Dominik, et al.
Published: (2022)
Navigating Shortcuts, Spurious Correlations, and Confounders: From Origins via Detection to Mitigation
by: Steinmann, David, et al.
Published: (2024)
by: Steinmann, David, et al.
Published: (2024)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026)
by: Struppek, Lukas, et al.
Published: (2026)
CollaFuse: Navigating Limited Resources and Privacy in Collaborative Generative AI
by: Zipperling, Domenique, et al.
Published: (2024)
by: Zipperling, Domenique, et al.
Published: (2024)
Deep Classifier Mimicry without Data Access
by: Braun, Steven, et al.
Published: (2023)
by: Braun, Steven, et al.
Published: (2023)
A Typology for Exploring the Mitigation of Shortcut Behavior
by: Friedrich, Felix, et al.
Published: (2022)
by: Friedrich, Felix, et al.
Published: (2022)
CollaFuse: Collaborative Diffusion Models
by: Allmendinger, Simeon, et al.
Published: (2024)
by: Allmendinger, Simeon, et al.
Published: (2024)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
by: Bohnet, Bernd, et al.
Published: (2024)
by: Bohnet, Bernd, et al.
Published: (2024)
SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems
by: Shindo, Hikaru, et al.
Published: (2026)
by: Shindo, Hikaru, et al.
Published: (2026)
HackAtari: Atari Learning Environments for Robust and Continual Reinforcement Learning
by: Delfosse, Quentin, et al.
Published: (2024)
by: Delfosse, Quentin, et al.
Published: (2024)
GRAIL: Autonomous Concept Grounding for Neuro-Symbolic Reinforcement Learning
by: Shindo, Hikaru, et al.
Published: (2026)
by: Shindo, Hikaru, et al.
Published: (2026)
Object Centric Concept Bottlenecks
by: Steinmann, David, et al.
Published: (2025)
by: Steinmann, David, et al.
Published: (2025)
Hyperparameter Optimization via Interacting with Probabilistic Circuits
by: Seng, Jonas, et al.
Published: (2025)
by: Seng, Jonas, et al.
Published: (2025)
Learning to Intervene on Concept Bottlenecks
by: Steinmann, David, et al.
Published: (2023)
by: Steinmann, David, et al.
Published: (2023)
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
by: Helff, Lukas, et al.
Published: (2024)
by: Helff, Lukas, et al.
Published: (2024)
BlendRL: A Framework for Merging Symbolic and Neural Policy Learning
by: Shindo, Hikaru, et al.
Published: (2024)
by: Shindo, Hikaru, et al.
Published: (2024)
Fodor and Pylyshyn's Legacy: Still No Human-like Systematic Compositionality in Neural Networks
by: Woydt, Tim, et al.
Published: (2025)
by: Woydt, Tim, et al.
Published: (2025)
R.I.P.: A Simple Black-box Attack on Continual Test-time Adaptation
by: Hoang, Trung-Hieu, et al.
Published: (2024)
by: Hoang, Trung-Hieu, et al.
Published: (2024)
Better Decisions through the Right Causal World Model
by: Dillies, Elisabeth, et al.
Published: (2025)
by: Dillies, Elisabeth, et al.
Published: (2025)
Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
by: Kohler, Hector, et al.
Published: (2024)
by: Kohler, Hector, et al.
Published: (2024)
Recursive Inference Machines for Neural Reasoning
by: Komisarczyk, Mieszko, et al.
Published: (2026)
by: Komisarczyk, Mieszko, et al.
Published: (2026)
Adversarial Reinforcement Learning for Large Language Model Agent Safety
by: Wang, Zizhao, et al.
Published: (2025)
by: Wang, Zizhao, et al.
Published: (2025)
Neural Concept Binder
by: Stammer, Wolfgang, et al.
Published: (2024)
by: Stammer, Wolfgang, et al.
Published: (2024)
Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?
by: Wüst, Antonia, et al.
Published: (2024)
by: Wüst, Antonia, et al.
Published: (2024)
Hierarchical Autoencoder-based Lossy Compression for Large-scale High-resolution Scientific Data
by: Le, Hieu, et al.
Published: (2023)
by: Le, Hieu, et al.
Published: (2023)
Scaling Capability in Token Space: An Analysis of Large Vision Language Model
by: Li, Tenghui, et al.
Published: (2024)
by: Li, Tenghui, et al.
Published: (2024)
V-LoL: A Diagnostic Dataset for Visual Logical Learning
by: Helff, Lukas, et al.
Published: (2023)
by: Helff, Lukas, et al.
Published: (2023)
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
by: Tong, Ziyi, et al.
Published: (2026)
by: Tong, Ziyi, et al.
Published: (2026)
Credibility-Aware Multi-Modal Fusion Using Probabilistic Circuits
by: Sidheekh, Sahil, et al.
Published: (2024)
by: Sidheekh, Sahil, et al.
Published: (2024)
Tagged for Direction: Pinning Down Causal Edge Directions with Precision
by: Busch, Florian Peter, et al.
Published: (2025)
by: Busch, Florian Peter, et al.
Published: (2025)
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
by: Deiseroth, Björn, et al.
Published: (2023)
by: Deiseroth, Björn, et al.
Published: (2023)
Tractable Representation Learning with Probabilistic Circuits
by: Braun, Steven, et al.
Published: (2025)
by: Braun, Steven, et al.
Published: (2025)
Deep Reinforcement Learning via Object-Centric Attention
by: Blüml, Jannis, et al.
Published: (2025)
by: Blüml, Jannis, et al.
Published: (2025)
Similar Items
-
Finding NeMo: Localizing Neurons Responsible For Memorization in Diffusion Models
by: Hintersdorf, Dominik, et al.
Published: (2024) -
Be Careful What You Smooth For: Label Smoothing Can Be a Privacy Shield but Also a Catalyst for Model Inversion Attacks
by: Struppek, Lukas, et al.
Published: (2023) -
Finding DoRI: Discovery of Retained Images in Diffusion Models
by: Kowalczuk, Antoni, et al.
Published: (2025) -
Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
by: Struppek, Lukas, et al.
Published: (2025) -
Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
by: Struppek, Lukas, et al.
Published: (2022)