Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Xingqi, Yi, Xiaoyuan, Xie, Xing, Jia, Jia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025)
di: Wu, Yi, et al.
Pubblicazione: (2025)
Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework
di: Arkhipkin, Vladimir, et al.
Pubblicazione: (2024)
di: Arkhipkin, Vladimir, et al.
Pubblicazione: (2024)
TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
di: Shi, Chuancheng, et al.
Pubblicazione: (2026)
di: Shi, Chuancheng, et al.
Pubblicazione: (2026)
Concept Conductor: Orchestrating Multiple Personalized Concepts in Text-to-Image Synthesis
di: Yao, Zebin, et al.
Pubblicazione: (2024)
di: Yao, Zebin, et al.
Pubblicazione: (2024)
ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer
di: Chen, Hongruixuan, et al.
Pubblicazione: (2023)
di: Chen, Hongruixuan, et al.
Pubblicazione: (2023)
Audio Visual Segmentation Through Text Embeddings
di: Lee, Kyungbok, et al.
Pubblicazione: (2025)
di: Lee, Kyungbok, et al.
Pubblicazione: (2025)
A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
di: Hei, Nailei, et al.
Pubblicazione: (2024)
di: Hei, Nailei, et al.
Pubblicazione: (2024)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
di: Han, Jiaming, et al.
Pubblicazione: (2025)
di: Han, Jiaming, et al.
Pubblicazione: (2025)
Image Conductor: Precision Control for Interactive Video Synthesis
di: Li, Yaowei, et al.
Pubblicazione: (2024)
di: Li, Yaowei, et al.
Pubblicazione: (2024)
Can LLMs Create Legally Relevant Summaries and Analyses of Videos?
di: Hoeben-Kuil, Lyra, et al.
Pubblicazione: (2025)
di: Hoeben-Kuil, Lyra, et al.
Pubblicazione: (2025)
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
di: Jo, Claire Wonjeong, et al.
Pubblicazione: (2024)
di: Jo, Claire Wonjeong, et al.
Pubblicazione: (2024)
KI-Bilder und die Widerständigkeit der Medienkonvergenz: Von primärer zu sekundärer Intermedialität?
di: Wilde, Lukas R. A.
Pubblicazione: (2024)
di: Wilde, Lukas R. A.
Pubblicazione: (2024)
Towards nation-wide analytical healthcare infrastructures: A privacy-preserving augmented knee rehabilitation case study
di: Bačić, Boris, et al.
Pubblicazione: (2024)
di: Bačić, Boris, et al.
Pubblicazione: (2024)
AI-based System for Transforming text and sound to Educational Videos
di: ElAlami, M. E., et al.
Pubblicazione: (2026)
di: ElAlami, M. E., et al.
Pubblicazione: (2026)
Text-Only Data Synthesis for Vision Language Model Training
di: Yu, Xiaomin, et al.
Pubblicazione: (2025)
di: Yu, Xiaomin, et al.
Pubblicazione: (2025)
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
Efficient Low-Resolution Face Recognition via Bridge Distillation
di: Ge, Shiming, et al.
Pubblicazione: (2024)
di: Ge, Shiming, et al.
Pubblicazione: (2024)
ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation
di: Ma, Zhiyuan, et al.
Pubblicazione: (2024)
di: Ma, Zhiyuan, et al.
Pubblicazione: (2024)
Truth in the Few: High-Value Data Selection for Efficient Multi-Modal Reasoning
di: Li, Shenshen, et al.
Pubblicazione: (2025)
di: Li, Shenshen, et al.
Pubblicazione: (2025)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
One Size, Many Fits: Aligning Diverse Group-Wise Click Preferences in Large-Scale Advertising Image Generation
di: Lu, Shuo, et al.
Pubblicazione: (2026)
di: Lu, Shuo, et al.
Pubblicazione: (2026)
PFB-Diff: Progressive Feature Blending Diffusion for Text-driven Image Editing
di: Huang, Wenjing, et al.
Pubblicazione: (2023)
di: Huang, Wenjing, et al.
Pubblicazione: (2023)
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
di: S, Sridhar, et al.
Pubblicazione: (2025)
di: S, Sridhar, et al.
Pubblicazione: (2025)
Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval
di: Luo, Bingjun, et al.
Pubblicazione: (2025)
di: Luo, Bingjun, et al.
Pubblicazione: (2025)
End-to-End Optimized Image Compression with the Frequency-Oriented Transform
di: Zhang, Yuefeng, et al.
Pubblicazione: (2024)
di: Zhang, Yuefeng, et al.
Pubblicazione: (2024)
Discriminative Probing and Tuning for Text-to-Image Generation
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
di: Ling, Jun, et al.
Pubblicazione: (2024)
di: Ling, Jun, et al.
Pubblicazione: (2024)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
di: Moratelli, Nicholas, et al.
Pubblicazione: (2024)
di: Moratelli, Nicholas, et al.
Pubblicazione: (2024)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
di: Qi, Peng, et al.
Pubblicazione: (2024)
di: Qi, Peng, et al.
Pubblicazione: (2024)
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
di: Lin, Xiao, et al.
Pubblicazione: (2025)
di: Lin, Xiao, et al.
Pubblicazione: (2025)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
di: Lv, Zheqi, et al.
Pubblicazione: (2025)
di: Lv, Zheqi, et al.
Pubblicazione: (2025)
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
di: Wu, Xiongwei, et al.
Pubblicazione: (2024)
di: Wu, Xiongwei, et al.
Pubblicazione: (2024)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
di: Jiang, Chen, et al.
Pubblicazione: (2023)
di: Jiang, Chen, et al.
Pubblicazione: (2023)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2024)
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2024)
High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model
di: Zhong, Weizhi, et al.
Pubblicazione: (2024)
di: Zhong, Weizhi, et al.
Pubblicazione: (2024)
FIGURA: A Modular Prompt Engineering Method for Artistic Figure Photography in Safety-Filtered Text-to-Image Models
di: Cazzaniga, Luca
Pubblicazione: (2026)
di: Cazzaniga, Luca
Pubblicazione: (2026)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
di: Yao, Lei, et al.
Pubblicazione: (2025)
di: Yao, Lei, et al.
Pubblicazione: (2025)
Composing Concepts from Images and Videos via Concept-prompt Binding
di: Kong, Xianghao, et al.
Pubblicazione: (2025)
di: Kong, Xianghao, et al.
Pubblicazione: (2025)
Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective
di: Zhu, Xiangru, et al.
Pubblicazione: (2024)
di: Zhu, Xiangru, et al.
Pubblicazione: (2024)
Documenti analoghi
-
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025) -
Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework
di: Arkhipkin, Vladimir, et al.
Pubblicazione: (2024) -
TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
di: Shi, Chuancheng, et al.
Pubblicazione: (2026) -
Concept Conductor: Orchestrating Multiple Personalized Concepts in Text-to-Image Synthesis
di: Yao, Zebin, et al.
Pubblicazione: (2024) -
ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer
di: Chen, Hongruixuan, et al.
Pubblicazione: (2023)