SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yoon, Jaehong, Yu, Shoubin, Patil, Vaidehi, Yao, Huaxiu, Bansal, Mohit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
von: Yu, Shoubin, et al.
Veröffentlicht: (2024)
von: Yu, Shoubin, et al.
Veröffentlicht: (2024)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
Hierarchy-Aware Multimodal Unlearning for Medical AI
von: Wu, Fengli, et al.
Veröffentlicht: (2025)
von: Wu, Fengli, et al.
Veröffentlicht: (2025)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
von: Lee, Daeun, et al.
Veröffentlicht: (2026)
von: Lee, Daeun, et al.
Veröffentlicht: (2026)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
von: Wang, Zun, et al.
Veröffentlicht: (2026)
von: Wang, Zun, et al.
Veröffentlicht: (2026)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
Multimodal Representation Learning by Alternating Unimodal Adaptation
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2023)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
von: Wang, Zun, et al.
Veröffentlicht: (2025)
von: Wang, Zun, et al.
Veröffentlicht: (2025)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
von: Huang, Yidong, et al.
Veröffentlicht: (2026)
von: Huang, Yidong, et al.
Veröffentlicht: (2026)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
von: Deng, Andong, et al.
Veröffentlicht: (2025)
von: Deng, Andong, et al.
Veröffentlicht: (2025)
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
von: Huang, Yidong, et al.
Veröffentlicht: (2025)
von: Huang, Yidong, et al.
Veröffentlicht: (2025)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
von: Jang, Sangwon, et al.
Veröffentlicht: (2025)
von: Jang, Sangwon, et al.
Veröffentlicht: (2025)
Are Video Reasoning Models Ready to Go Outside?
von: He, Yangfan, et al.
Veröffentlicht: (2026)
von: He, Yangfan, et al.
Veröffentlicht: (2026)
Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering
von: Frahm, Noah, et al.
Veröffentlicht: (2025)
von: Frahm, Noah, et al.
Veröffentlicht: (2025)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
von: Deng, Andong, et al.
Veröffentlicht: (2024)
von: Deng, Andong, et al.
Veröffentlicht: (2024)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection
von: Ma, Ruize, et al.
Veröffentlicht: (2025)
von: Ma, Ruize, et al.
Veröffentlicht: (2025)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2025)
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2025)
RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image Alignment
von: Jiang, Liyao, et al.
Veröffentlicht: (2026)
von: Jiang, Liyao, et al.
Veröffentlicht: (2026)
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
von: Lakhanpal, Sanyam, et al.
Veröffentlicht: (2024)
von: Lakhanpal, Sanyam, et al.
Veröffentlicht: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
Training-Free Consistent Text-to-Image Generation
von: Tewel, Yoad, et al.
Veröffentlicht: (2024)
von: Tewel, Yoad, et al.
Veröffentlicht: (2024)
Capsule Vision Challenge 2024: Multi-Class Abnormality Classification for Video Capsule Endoscopy
von: Bansal, Aakarsh, et al.
Veröffentlicht: (2024)
von: Bansal, Aakarsh, et al.
Veröffentlicht: (2024)
VideoGuard: Protecting Video Content from Unauthorized Editing
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
Progressive Image Restoration via Text-Conditioned Video Generation
von: Kang, Peng, et al.
Veröffentlicht: (2025)
von: Kang, Peng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024) -
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
von: Yu, Shoubin, et al.
Veröffentlicht: (2024) -
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025) -
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026) -
Hierarchy-Aware Multimodal Unlearning for Medical AI
von: Wu, Fengli, et al.
Veröffentlicht: (2025)