Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Alex Jinpeng, Li, Linjie, Yang, Zhengyuan, Wang, Lijuan, Li, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering
von: Mao, Dongxing, et al.
Veröffentlicht: (2026)
von: Mao, Dongxing, et al.
Veröffentlicht: (2026)
FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
von: Yi, Junchao, et al.
Veröffentlicht: (2026)
von: Yi, Junchao, et al.
Veröffentlicht: (2026)
TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025)
V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
von: Zheng, Xiangxi, et al.
Veröffentlicht: (2025)
von: Zheng, Xiangxi, et al.
Veröffentlicht: (2025)
Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
Glance: Accelerating Diffusion Models with 1 Sample
von: Dong, Zhuobai, et al.
Veröffentlicht: (2025)
von: Dong, Zhuobai, et al.
Veröffentlicht: (2025)
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents
von: Lin, Yiqi, et al.
Veröffentlicht: (2025)
von: Lin, Yiqi, et al.
Veröffentlicht: (2025)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
von: Yang, Zhengyuan, et al.
Veröffentlicht: (2023)
von: Yang, Zhengyuan, et al.
Veröffentlicht: (2023)
Conditional Text-to-Image Generation with Reference Guidance
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
Bring Metric Functions into Diffusion Models
von: An, Jie, et al.
Veröffentlicht: (2024)
von: An, Jie, et al.
Veröffentlicht: (2024)
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
LiVOS: Light Video Object Segmentation with Gated Linear Matching
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
von: Gui, Rui, et al.
Veröffentlicht: (2025)
von: Gui, Rui, et al.
Veröffentlicht: (2025)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
von: Zhang, Jihai, et al.
Veröffentlicht: (2025)
von: Zhang, Jihai, et al.
Veröffentlicht: (2025)
Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition
von: Qiu, Jielin, et al.
Veröffentlicht: (2024)
von: Qiu, Jielin, et al.
Veröffentlicht: (2024)
Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification
von: Lin, Rifen, et al.
Veröffentlicht: (2025)
von: Lin, Rifen, et al.
Veröffentlicht: (2025)
Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024)
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024)
Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters
von: Wang, Weizhi, et al.
Veröffentlicht: (2024)
von: Wang, Weizhi, et al.
Veröffentlicht: (2024)
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2024)
von: Yu, Weihao, et al.
Veröffentlicht: (2024)
Autoregressive Image Generation with Randomized Parallel Decoding
von: Li, Haopeng, et al.
Veröffentlicht: (2025)
von: Li, Haopeng, et al.
Veröffentlicht: (2025)
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
Scalable Autoregressive Image Generation with Mamba
von: Li, Haopeng, et al.
Veröffentlicht: (2024)
von: Li, Haopeng, et al.
Veröffentlicht: (2024)
IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024)
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024)
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
Fast Prompt Alignment for Text-to-Image Generation
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
DisCo: Disentangled Control for Realistic Human Dance Generation
von: Wang, Tan, et al.
Veröffentlicht: (2023)
von: Wang, Tan, et al.
Veröffentlicht: (2023)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
von: Zheng, Zirui, et al.
Veröffentlicht: (2025)
von: Zheng, Zirui, et al.
Veröffentlicht: (2025)
SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
von: Wang, Xiyao, et al.
Veröffentlicht: (2025)
von: Wang, Xiyao, et al.
Veröffentlicht: (2025)
Planning with the Views via Scene Self-Exploration
von: Wang, Kangrui, et al.
Veröffentlicht: (2026)
von: Wang, Kangrui, et al.
Veröffentlicht: (2026)
TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
von: Wang, Juntong, et al.
Veröffentlicht: (2025)
von: Wang, Juntong, et al.
Veröffentlicht: (2025)
List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
von: Yan, An, et al.
Veröffentlicht: (2024)
von: Yan, An, et al.
Veröffentlicht: (2024)
GenXD: Generating Any 3D and 4D Scenes
von: Zhao, Yuyang, et al.
Veröffentlicht: (2024)
von: Zhao, Yuyang, et al.
Veröffentlicht: (2024)
End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
von: Chu, Wenda, et al.
Veröffentlicht: (2026)
von: Chu, Wenda, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering
von: Mao, Dongxing, et al.
Veröffentlicht: (2026) -
FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
von: Yi, Junchao, et al.
Veröffentlicht: (2026) -
TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025) -
V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
von: Zheng, Xiangxi, et al.
Veröffentlicht: (2025) -
Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)