CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fu, Minghao, Wang, Guo-Hua, Cao, Liangfu, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
von: Cui, Tianyu, et al.
Veröffentlicht: (2025)
von: Cui, Tianyu, et al.
Veröffentlicht: (2025)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
von: Zhao, Shanshan, et al.
Veröffentlicht: (2025)
von: Zhao, Shanshan, et al.
Veröffentlicht: (2025)
Ovis-Image Technical Report
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
von: Duan, Lunhao, et al.
Veröffentlicht: (2024)
von: Duan, Lunhao, et al.
Veröffentlicht: (2024)
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
von: Pang, Lianyu, et al.
Veröffentlicht: (2024)
von: Pang, Lianyu, et al.
Veröffentlicht: (2024)
Ovis-U1 Technical Report
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)
Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
von: Yang, Wenhao, et al.
Veröffentlicht: (2025)
von: Yang, Wenhao, et al.
Veröffentlicht: (2025)
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
von: Yang, Wenhao, et al.
Veröffentlicht: (2026)
von: Yang, Wenhao, et al.
Veröffentlicht: (2026)
MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
von: Sun, Hui, et al.
Veröffentlicht: (2025)
von: Sun, Hui, et al.
Veröffentlicht: (2025)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
von: Lu, Shiyin, et al.
Veröffentlicht: (2024)
von: Lu, Shiyin, et al.
Veröffentlicht: (2024)
Test-Time Preference Optimization for Image Restoration
von: Li, Bingchen, et al.
Veröffentlicht: (2025)
von: Li, Bingchen, et al.
Veröffentlicht: (2025)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
von: Gao, Sensen, et al.
Veröffentlicht: (2025)
von: Gao, Sensen, et al.
Veröffentlicht: (2025)
Training-Free Image Editing with Visual Context Integration and Concept Alignment
von: Song, Rui, et al.
Veröffentlicht: (2026)
von: Song, Rui, et al.
Veröffentlicht: (2026)
TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images
von: Koltsov, Kirill, et al.
Veröffentlicht: (2026)
von: Koltsov, Kirill, et al.
Veröffentlicht: (2026)
No Concept Left Behind: Test-Time Optimization for Compositional Text-to-Image Generation
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2025)
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2025)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
HyperAlign: Hypernetwork for Efficient Test-Time Alignment of Diffusion Models
von: Xie, Xin, et al.
Veröffentlicht: (2026)
von: Xie, Xin, et al.
Veröffentlicht: (2026)
A Unified Agentic Framework for Evaluating Conditional Image Generation
von: Wang, Jifang, et al.
Veröffentlicht: (2025)
von: Wang, Jifang, et al.
Veröffentlicht: (2025)
Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh Recovery
von: Nie, Yongwei, et al.
Veröffentlicht: (2024)
von: Nie, Yongwei, et al.
Veröffentlicht: (2024)
Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
von: Shao, Jie, et al.
Veröffentlicht: (2025)
von: Shao, Jie, et al.
Veröffentlicht: (2025)
Verify Claimed Text-to-Image Models via Boundary-Aware Prompt Optimization
von: Zhao, Zidong, et al.
Veröffentlicht: (2026)
von: Zhao, Zidong, et al.
Veröffentlicht: (2026)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation
von: Liang, Xinyue, et al.
Veröffentlicht: (2025)
von: Liang, Xinyue, et al.
Veröffentlicht: (2025)
EmoAttack: Emotion-to-Image Diffusion Models for Emotional Backdoor Generation
von: Wei, Tianyu, et al.
Veröffentlicht: (2024)
von: Wei, Tianyu, et al.
Veröffentlicht: (2024)
Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation
von: Wang, Xintong, et al.
Veröffentlicht: (2025)
von: Wang, Xintong, et al.
Veröffentlicht: (2025)
Controllable Generation with Text-to-Image Diffusion Models: A Survey
von: Cao, Pu, et al.
Veröffentlicht: (2024)
von: Cao, Pu, et al.
Veröffentlicht: (2024)
Filter & Align: Leveraging Human Knowledge to Curate Image-Text Data
von: Zhang, Lei, et al.
Veröffentlicht: (2023)
von: Zhang, Lei, et al.
Veröffentlicht: (2023)
Diff-Instruct++: Training One-step Text-to-image Generator Model to Align with Human Preferences
von: Luo, Weijian
Veröffentlicht: (2024)
von: Luo, Weijian
Veröffentlicht: (2024)
MMIF-AMIN: Adaptive Loss-Driven Multi-Scale Invertible Dense Network for Multimodal Medical Image Fusion
von: Luo, Tao, et al.
Veröffentlicht: (2025)
von: Luo, Tao, et al.
Veröffentlicht: (2025)
Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
von: Hu, Zijing, et al.
Veröffentlicht: (2025)
von: Hu, Zijing, et al.
Veröffentlicht: (2025)
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2026)
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2026)
Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation
von: Li, Baoteng, et al.
Veröffentlicht: (2026)
von: Li, Baoteng, et al.
Veröffentlicht: (2026)
Dynamic Prompt Optimizing for Text-to-Image Generation
von: Mo, Wenyi, et al.
Veröffentlicht: (2024)
von: Mo, Wenyi, et al.
Veröffentlicht: (2024)
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
von: Xu, Zhenran, et al.
Veröffentlicht: (2025)
von: Xu, Zhenran, et al.
Veröffentlicht: (2025)
ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation
von: Huang, Qing, et al.
Veröffentlicht: (2026)
von: Huang, Qing, et al.
Veröffentlicht: (2026)
Rich Human Feedback for Text-to-Image Generation
von: Liang, Youwei, et al.
Veröffentlicht: (2023)
von: Liang, Youwei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
von: Fu, Minghao, et al.
Veröffentlicht: (2025) -
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
von: Fu, Minghao, et al.
Veröffentlicht: (2025) -
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
von: Cui, Tianyu, et al.
Veröffentlicht: (2025) -
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
von: Zhao, Shanshan, et al.
Veröffentlicht: (2025) -
Ovis-Image Technical Report
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)