POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Yaohou, Wang, Qingzhong, Huang, Yongsong, Liu, Junyi, Miyazaki, Tomo, Omachi, Shinichiro |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPSMamba: A Global Phase and Spectral Prompt-guided Mamba for Infrared Image Super-Resolution
by: Huang, Yongsong, et al.
Published: (2025)
by: Huang, Yongsong, et al.
Published: (2025)
IRSRMamba: Infrared Image Super-Resolution via Mamba-based Wavelet Transform Feature Modulation Model
by: Huang, Yongsong, et al.
Published: (2024)
by: Huang, Yongsong, et al.
Published: (2024)
Infrared Image Super-Resolution: Systematic Review, and Future Trends
by: Huang, Yongsong, et al.
Published: (2022)
by: Huang, Yongsong, et al.
Published: (2022)
Texture and Noise Dual Adaptation for Infrared Image Super-Resolution
by: Huang, Yongsong, et al.
Published: (2023)
by: Huang, Yongsong, et al.
Published: (2023)
Towards Cross-Domain Multi-Targeted Adversarial Attacks
by: Gonçalves, Taïga, et al.
Published: (2025)
by: Gonçalves, Taïga, et al.
Published: (2025)
Controlling Rate, Distortion, and Realism: Towards a Single Comprehensive Neural Image Compression Model
by: Iwai, Shoma, et al.
Published: (2024)
by: Iwai, Shoma, et al.
Published: (2024)
Joint Low-level and High-level Textual Representation Learning with Multiple Masking Strategies
by: Tang, Zhengmi, et al.
Published: (2025)
by: Tang, Zhengmi, et al.
Published: (2025)
GTFMN: Guided Texture and Feature Modulation Network for Low-Light Image Enhancement and Super-Resolution
by: Huang, Yongsong, et al.
Published: (2026)
by: Huang, Yongsong, et al.
Published: (2026)
Class-agnostic 3D Segmentation by Granularity-Consistent Automatic 2D Mask Tracking
by: Wang, Juan, et al.
Published: (2025)
by: Wang, Juan, et al.
Published: (2025)
U-Harmony: Enhancing Joint Training for Segmentation Models with Universal Harmonization
by: Ma, Weiwei, et al.
Published: (2026)
by: Ma, Weiwei, et al.
Published: (2026)
Pareto-Guided Optimal Transport for Multi-Reward Alignment
by: Ba, Ying, et al.
Published: (2026)
by: Ba, Ying, et al.
Published: (2026)
ALPS: An Auto-Labeling and Pre-training Scheme for Remote Sensing Segmentation With Segment Anything Model
by: Zhang, Song, et al.
Published: (2024)
by: Zhang, Song, et al.
Published: (2024)
Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics
by: Wang, Yunlong, et al.
Published: (2026)
by: Wang, Yunlong, et al.
Published: (2026)
DiCTI: Diffusion-based Clothing Designer via Text-guided Input
by: Lampe, Ajda, et al.
Published: (2024)
by: Lampe, Ajda, et al.
Published: (2024)
Layout-Corrector: Alleviating Layout Sticking Phenomenon in Discrete Diffusion Model
by: Iwai, Shoma, et al.
Published: (2024)
by: Iwai, Shoma, et al.
Published: (2024)
Visual Text Generation in the Wild
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment
by: Hu, Junyi, et al.
Published: (2026)
by: Hu, Junyi, et al.
Published: (2026)
Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
by: Wang, Shijian, et al.
Published: (2025)
by: Wang, Shijian, et al.
Published: (2025)
Native Audio-Visual Alignment for Generation
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
Fast Prompt Alignment for Text-to-Image Generation
by: Mrini, Khalil, et al.
Published: (2024)
by: Mrini, Khalil, et al.
Published: (2024)
Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation
by: Lee, Seung Hyun, et al.
Published: (2024)
by: Lee, Seung Hyun, et al.
Published: (2024)
HOTS3D: Hyper-Spherical Optimal Transport for Semantic Alignment of Text-to-3D Generation
by: Li, Zezeng, et al.
Published: (2024)
by: Li, Zezeng, et al.
Published: (2024)
Harmonizing Visual Text Comprehension and Generation
by: Zhao, Zhen, et al.
Published: (2024)
by: Zhao, Zhen, et al.
Published: (2024)
Text-Animator: Controllable Visual Text Video Generation
by: Liu, Lin, et al.
Published: (2024)
by: Liu, Lin, et al.
Published: (2024)
TiC: Exploring Vision Transformer in Convolution
by: Zhang, Song, et al.
Published: (2023)
by: Zhang, Song, et al.
Published: (2023)
InstructEngine: Instruction-driven Text-to-Image Alignment
by: Lu, Xingyu, et al.
Published: (2025)
by: Lu, Xingyu, et al.
Published: (2025)
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
Exploring Motion-Language Alignment for Text-driven Motion Generation
by: Gu, Ruxi, et al.
Published: (2026)
by: Gu, Ruxi, et al.
Published: (2026)
Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content
by: Zhang, Zicheng, et al.
Published: (2025)
by: Zhang, Zicheng, et al.
Published: (2025)
Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment
by: Xie, Xing, et al.
Published: (2025)
by: Xie, Xing, et al.
Published: (2025)
Zero-Shot Skeleton-based Action Recognition with Dual Visual-Text Alignment
by: Kuang, Jidong, et al.
Published: (2024)
by: Kuang, Jidong, et al.
Published: (2024)
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
GVGEN: Text-to-3D Generation with Volumetric Representation
by: He, Xianglong, et al.
Published: (2024)
by: He, Xianglong, et al.
Published: (2024)
Learning Visual Generative Priors without Text
by: Ma, Shuailei, et al.
Published: (2024)
by: Ma, Shuailei, et al.
Published: (2024)
Interactive Visual Assessment for Text-to-Image Generation Models
by: Mi, Xiaoyue, et al.
Published: (2024)
by: Mi, Xiaoyue, et al.
Published: (2024)
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
by: Pan, Kaihang, et al.
Published: (2025)
by: Pan, Kaihang, et al.
Published: (2025)
SpatialV2A: Visual-Guided High-fidelity Spatial Audio Generation
by: Wang, Yanan, et al.
Published: (2026)
by: Wang, Yanan, et al.
Published: (2026)
SAN: Structure-Aware Network for Complex and Long-tailed Chinese Text Recognition
by: Zhang, Junyi, et al.
Published: (2024)
by: Zhang, Junyi, et al.
Published: (2024)
Can Text-to-Video Generation help Video-Language Alignment?
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
by: Huang, Jun, et al.
Published: (2025)
by: Huang, Jun, et al.
Published: (2025)
Similar Items
-
GPSMamba: A Global Phase and Spectral Prompt-guided Mamba for Infrared Image Super-Resolution
by: Huang, Yongsong, et al.
Published: (2025) -
IRSRMamba: Infrared Image Super-Resolution via Mamba-based Wavelet Transform Feature Modulation Model
by: Huang, Yongsong, et al.
Published: (2024) -
Infrared Image Super-Resolution: Systematic Review, and Future Trends
by: Huang, Yongsong, et al.
Published: (2022) -
Texture and Noise Dual Adaptation for Infrared Image Super-Resolution
by: Huang, Yongsong, et al.
Published: (2023) -
Towards Cross-Domain Multi-Targeted Adversarial Attacks
by: Gonçalves, Taïga, et al.
Published: (2025)