Conditioned Activation Transport for T2I Safety Steering
Fuente:
arXiv
Saved in:
| Main Authors: | Chrabąszcz, Maciej, Szymczyk, Aleksander, Dubiński, Jan, Trzciński, Tomasz, Boenisch, Franziska, Dziedzic, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Radioactive Watermarks in Diffusion and Autoregressive Image Generative Models
by: Meintz, Michel, et al.
Published: (2025)
by: Meintz, Michel, et al.
Published: (2025)
Privacy Attacks on Image AutoRegressive Models
by: Kowalczuk, Antoni, et al.
Published: (2025)
by: Kowalczuk, Antoni, et al.
Published: (2025)
BitMark: Watermarking Bitwise Autoregressive Image Generative Models
by: Kerner, Louis, et al.
Published: (2025)
by: Kerner, Louis, et al.
Published: (2025)
Benchmarking Robust Self-Supervised Learning Across Diverse Downstream Tasks
by: Kowalczuk, Antoni, et al.
Published: (2024)
by: Kowalczuk, Antoni, et al.
Published: (2024)
ExpertSim: Fast Particle Detector Simulation Using Mixture-of-Generative-Experts
by: Będkowski, Patryk, et al.
Published: (2025)
by: Będkowski, Patryk, et al.
Published: (2025)
Demystifying Foreground-Background Memorization in Diffusion Models
by: Di, Jimmy Z., et al.
Published: (2025)
by: Di, Jimmy Z., et al.
Published: (2025)
Captured by Captions: On Memorization and its Mitigation in CLIP Models
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
Finding DoRI: Discovery of Retained Images in Diffusion Models
by: Kowalczuk, Antoni, et al.
Published: (2025)
by: Kowalczuk, Antoni, et al.
Published: (2025)
Beautiful Images, Toxic Words: Understanding and Addressing Offensive Text in Generated Images
by: Kumar, Aditya, et al.
Published: (2025)
by: Kumar, Aditya, et al.
Published: (2025)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
by: Chrabąszcz, Maciej, et al.
Published: (2025)
by: Chrabąszcz, Maciej, et al.
Published: (2025)
Localizing Memorization in SSL Vision Encoders
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Precise Parameter Localization for Textual Generation in Diffusion Models
by: Staniszewski, Łukasz, et al.
Published: (2025)
by: Staniszewski, Łukasz, et al.
Published: (2025)
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
by: Chrabąszcz, Maciej, et al.
Published: (2026)
by: Chrabąszcz, Maciej, et al.
Published: (2026)
SERUM: Simple, Efficient, Robust, and Unifying Marking for Diffusion-based Image Generation
by: Kociszewski, Jan, et al.
Published: (2026)
by: Kociszewski, Jan, et al.
Published: (2026)
MagMax: Leveraging Model Merging for Seamless Continual Learning
by: Marczak, Daniel, et al.
Published: (2024)
by: Marczak, Daniel, et al.
Published: (2024)
AR-TTA: A Simple Method for Real-World Continual Test-Time Adaptation
by: Sójka, Damian, et al.
Published: (2023)
by: Sójka, Damian, et al.
Published: (2023)
Exploring the Stability Gap in Continual Learning: The Role of the Classification Head
by: Łapacz, Wojciech, et al.
Published: (2024)
by: Łapacz, Wojciech, et al.
Published: (2024)
Implementing Adaptations for Vision AutoRegressive Model
by: Shaikh, Kaif, et al.
Published: (2025)
by: Shaikh, Kaif, et al.
Published: (2025)
Let Me DeCode You: Decoder Conditioning with Tabular Data
by: Szczepański, Tomasz, et al.
Published: (2024)
by: Szczepański, Tomasz, et al.
Published: (2024)
MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
by: Pham, Trong-Thang, et al.
Published: (2026)
by: Pham, Trong-Thang, et al.
Published: (2026)
Deep Generative Models for Proton Zero Degree Calorimeter Simulations in ALICE, CERN
by: Będkowski, Patryk, et al.
Published: (2024)
by: Będkowski, Patryk, et al.
Published: (2024)
Jailbreaking Vision-Language Models Through the Visual Modality
by: Azulay, Aharon, et al.
Published: (2026)
by: Azulay, Aharon, et al.
Published: (2026)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
by: Yin, Jianghao, et al.
Published: (2026)
by: Yin, Jianghao, et al.
Published: (2026)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
by: Yang, Jiaxi, et al.
Published: (2026)
by: Yang, Jiaxi, et al.
Published: (2026)
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
by: Wu, Sihao, et al.
Published: (2025)
by: Wu, Sihao, et al.
Published: (2025)
Modeling 3D Surface Manifolds with a Locally Conditioned Atlas
by: Spurek, Przemysław, et al.
Published: (2021)
by: Spurek, Przemysław, et al.
Published: (2021)
AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
by: Pardyl, Adam, et al.
Published: (2024)
by: Pardyl, Adam, et al.
Published: (2024)
MedShift: Implicit Conditional Transport for X-Ray Domain Adaptation
by: Caetano, Francisco, et al.
Published: (2025)
by: Caetano, Francisco, et al.
Published: (2025)
Beyond Grids: Exploring Elastic Input Sampling for Vision Transformers
by: Pardyl, Adam, et al.
Published: (2023)
by: Pardyl, Adam, et al.
Published: (2023)
SafeEditor: Unified MLLM for Efficient Post-hoc T2I Safety Editing
by: Zhang, Ruiyang, et al.
Published: (2025)
by: Zhang, Ruiyang, et al.
Published: (2025)
Inline Critic Steers Image Editing
by: Kang, Weitai, et al.
Published: (2026)
by: Kang, Weitai, et al.
Published: (2026)
Decoding Vision Transformers: the Diffusion Steering Lens
by: Takatsuki, Ryota, et al.
Published: (2025)
by: Takatsuki, Ryota, et al.
Published: (2025)
Causally Steered Diffusion for Automated Video Counterfactual Generation
by: Spyrou, Nikos, et al.
Published: (2025)
by: Spyrou, Nikos, et al.
Published: (2025)
Swin SMT: Global Sequential Modeling in 3D Medical Image Segmentation
by: Płotka, Szymon, et al.
Published: (2024)
by: Płotka, Szymon, et al.
Published: (2024)
GEPAR3D: Geometry Prior-Assisted Learning for 3D Tooth Segmentation
by: Szczepański, Tomasz, et al.
Published: (2025)
by: Szczepański, Tomasz, et al.
Published: (2025)
Language Models Can Explain Visual Features via Steering
by: Ferrando, Javier, et al.
Published: (2026)
by: Ferrando, Javier, et al.
Published: (2026)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026)
by: Liu, Zeyu, et al.
Published: (2026)
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
by: Hinojosa, Carlos, et al.
Published: (2026)
by: Hinojosa, Carlos, et al.
Published: (2026)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
by: Zhang, Yuanhong, et al.
Published: (2026)
by: Zhang, Yuanhong, et al.
Published: (2026)
FlowSteer: Guiding Few-Step Image Synthesis with Authentic Trajectories
by: Ke, Lei, et al.
Published: (2025)
by: Ke, Lei, et al.
Published: (2025)
Similar Items
-
Radioactive Watermarks in Diffusion and Autoregressive Image Generative Models
by: Meintz, Michel, et al.
Published: (2025) -
Privacy Attacks on Image AutoRegressive Models
by: Kowalczuk, Antoni, et al.
Published: (2025) -
BitMark: Watermarking Bitwise Autoregressive Image Generative Models
by: Kerner, Louis, et al.
Published: (2025) -
Benchmarking Robust Self-Supervised Learning Across Diverse Downstream Tasks
by: Kowalczuk, Antoni, et al.
Published: (2024) -
ExpertSim: Fast Particle Detector Simulation Using Mixture-of-Generative-Experts
by: Będkowski, Patryk, et al.
Published: (2025)