Latent Guard: a Safety Framework for Text-to-image Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Runtao, Khakzar, Ashkan, Gu, Jindong, Chen, Qifeng, Torr, Philip, Pizzati, Fabio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
por: Liu, Runtao, et al.
Publicado: (2024)
por: Liu, Runtao, et al.
Publicado: (2024)
Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
por: Iakovleva, Ekaterina, et al.
Publicado: (2024)
por: Iakovleva, Ekaterina, et al.
Publicado: (2024)
Video Motion Transfer with Diffusion Transformers
por: Pondaven, Alexander, et al.
Publicado: (2024)
por: Pondaven, Alexander, et al.
Publicado: (2024)
ActionParty: Multi-Subject Action Binding in Generative Video Games
por: Pondaven, Alexander, et al.
Publicado: (2026)
por: Pondaven, Alexander, et al.
Publicado: (2026)
Multimodal Pragmatic Jailbreak on Text-to-image Models
por: Liu, Tong, et al.
Publicado: (2024)
por: Liu, Tong, et al.
Publicado: (2024)
MatchDiffusion: Training-free Generation of Match-cuts
por: Pardo, Alejandro, et al.
Publicado: (2024)
por: Pardo, Alejandro, et al.
Publicado: (2024)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
por: Liu, Runtao, et al.
Publicado: (2024)
por: Liu, Runtao, et al.
Publicado: (2024)
On Pretraining Data Diversity for Self-Supervised Learning
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2024)
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2024)
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2024)
Fake it till You Make it: Reward Modeling as Discriminative Prediction
por: Liu, Runtao, et al.
Publicado: (2025)
por: Liu, Runtao, et al.
Publicado: (2025)
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
por: Rao, Zhefan, et al.
Publicado: (2024)
por: Rao, Zhefan, et al.
Publicado: (2024)
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
por: Khalifi, Omar El, et al.
Publicado: (2026)
por: Khalifi, Omar El, et al.
Publicado: (2026)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
por: Rezaei, Razieh, et al.
Publicado: (2024)
por: Rezaei, Razieh, et al.
Publicado: (2024)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
por: Venhoff, Constantin, et al.
Publicado: (2025)
por: Venhoff, Constantin, et al.
Publicado: (2025)
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
por: Deb, Oishi, et al.
Publicado: (2025)
por: Deb, Oishi, et al.
Publicado: (2025)
Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
por: Chen, Shuo, et al.
Publicado: (2023)
por: Chen, Shuo, et al.
Publicado: (2023)
Learning to Generate Rigid Body Interactions with Video Diffusion Models
por: Romero, David, et al.
Publicado: (2025)
por: Romero, David, et al.
Publicado: (2025)
LongVideoAgent: Multi-Agent Reasoning with Long Videos
por: Liu, Runtao, et al.
Publicado: (2025)
por: Liu, Runtao, et al.
Publicado: (2025)
Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
por: Li, Hang, et al.
Publicado: (2023)
por: Li, Hang, et al.
Publicado: (2023)
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
por: Yuan, Jianhao, et al.
Publicado: (2025)
por: Yuan, Jianhao, et al.
Publicado: (2025)
Minimalist Concept Erasure in Generative Models
por: Zhang, Yang, et al.
Publicado: (2025)
por: Zhang, Yang, et al.
Publicado: (2025)
True Multimodal In-Context Learning Needs Attention to the Visual Context
por: Chen, Shuo, et al.
Publicado: (2025)
por: Chen, Shuo, et al.
Publicado: (2025)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
por: Yoon, Jaehong, et al.
Publicado: (2024)
por: Yoon, Jaehong, et al.
Publicado: (2024)
Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
por: Gao, Kuofeng, et al.
Publicado: (2024)
por: Gao, Kuofeng, et al.
Publicado: (2024)
DiffGuard: Text-Based Safety Checker for Diffusion Models
por: Khader, Massine El, et al.
Publicado: (2024)
por: Khader, Massine El, et al.
Publicado: (2024)
FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding
por: Yang, Jinghan, et al.
Publicado: (2026)
por: Yang, Jinghan, et al.
Publicado: (2026)
DreamPolisher: Towards High-Quality Text-to-3D Generation via Geometric Diffusion
por: Lin, Yuanze, et al.
Publicado: (2024)
por: Lin, Yuanze, et al.
Publicado: (2024)
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
por: Wang, Zefeng, et al.
Publicado: (2024)
por: Wang, Zefeng, et al.
Publicado: (2024)
ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection
por: Ma, Ruize, et al.
Publicado: (2025)
por: Ma, Ruize, et al.
Publicado: (2025)
PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
por: Vu, Tung, et al.
Publicado: (2025)
por: Vu, Tung, et al.
Publicado: (2025)
Fair Generation without Unfair Distortions: Debiasing Text-to-Image Generation with Entanglement-Free Attention
por: Park, Jeonghoon, et al.
Publicado: (2025)
por: Park, Jeonghoon, et al.
Publicado: (2025)
Towards Understanding Multimodal Fine-Tuning: Spatial Features
por: Naghashyar, Lachin, et al.
Publicado: (2026)
por: Naghashyar, Lachin, et al.
Publicado: (2026)
A Survey on Responsible Generative AI: What to Generate and What Not
por: Gu, Jindong
Publicado: (2024)
por: Gu, Jindong
Publicado: (2024)
Self-Corrected Image Generation with Explainable Latent Rewards
por: Luo, Yinyi, et al.
Publicado: (2026)
por: Luo, Yinyi, et al.
Publicado: (2026)
Localizing Events in Videos with Multimodal Queries
por: Zhang, Gengyuan, et al.
Publicado: (2024)
por: Zhang, Gengyuan, et al.
Publicado: (2024)
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
por: Helff, Lukas, et al.
Publicado: (2024)
por: Helff, Lukas, et al.
Publicado: (2024)
Regularization by Texts for Latent Diffusion Inverse Solvers
por: Kim, Jeongsol, et al.
Publicado: (2023)
por: Kim, Jeongsol, et al.
Publicado: (2023)
Text-to-image Diffusion Models in Generative AI: A Survey
por: Zhang, Chenshuang, et al.
Publicado: (2023)
por: Zhang, Chenshuang, et al.
Publicado: (2023)
Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
por: Tang, Jiaqi, et al.
Publicado: (2025)
por: Tang, Jiaqi, et al.
Publicado: (2025)
Temporal Pair Consistency for Variance-Reduced Flow Matching
por: Maduabuchi, Chika, et al.
Publicado: (2026)
por: Maduabuchi, Chika, et al.
Publicado: (2026)
Ejemplares similares
-
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
por: Liu, Runtao, et al.
Publicado: (2024) -
Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
por: Iakovleva, Ekaterina, et al.
Publicado: (2024) -
Video Motion Transfer with Diffusion Transformers
por: Pondaven, Alexander, et al.
Publicado: (2024) -
ActionParty: Multi-Subject Action Binding in Generative Video Games
por: Pondaven, Alexander, et al.
Publicado: (2026) -
Multimodal Pragmatic Jailbreak on Text-to-image Models
por: Liu, Tong, et al.
Publicado: (2024)