Scalable Pre-training of Large Autoregressive Image Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | El-Nouby, Alaaeldin, Klein, Michal, Zhai, Shuangfei, Bautista, Miguel Angel, Toshev, Alexander, Shankar, Vaishaal, Susskind, Joshua M, Joulin, Armand |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Autoregressive Pre-training of Large Vision Encoders
von: Fini, Enrico, et al.
Veröffentlicht: (2024)
von: Fini, Enrico, et al.
Veröffentlicht: (2024)
World-consistent Video Diffusion with Explicit 3D Modeling
von: Zhang, Qihang, et al.
Veröffentlicht: (2024)
von: Zhang, Qihang, et al.
Veröffentlicht: (2024)
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2023)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2023)
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
von: Gu, Jiatao, et al.
Veröffentlicht: (2024)
von: Gu, Jiatao, et al.
Veröffentlicht: (2024)
Scaling Laws for Native Multimodal Models
von: Shukor, Mustafa, et al.
Veröffentlicht: (2025)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2025)
Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
von: Gu, Jiatao, et al.
Veröffentlicht: (2024)
von: Gu, Jiatao, et al.
Veröffentlicht: (2024)
Many-to-many Image Generation with Auto-regressive Diffusion Models
von: Shen, Ying, et al.
Veröffentlicht: (2024)
von: Shen, Ying, et al.
Veröffentlicht: (2024)
Normalizing Flows with Iterative Denoising
von: Chen, Tianrong, et al.
Veröffentlicht: (2026)
von: Chen, Tianrong, et al.
Veröffentlicht: (2026)
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
von: Shen, Ying, et al.
Veröffentlicht: (2026)
von: Shen, Ying, et al.
Veröffentlicht: (2026)
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
von: Gu, Jiatao, et al.
Veröffentlicht: (2025)
von: Gu, Jiatao, et al.
Veröffentlicht: (2025)
MobileCLIP2: Improving Multi-Modal Reinforced Training
von: Faghri, Fartash, et al.
Veröffentlicht: (2025)
von: Faghri, Fartash, et al.
Veröffentlicht: (2025)
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
von: Gu, Jiatao, et al.
Veröffentlicht: (2025)
von: Gu, Jiatao, et al.
Veröffentlicht: (2025)
Matryoshka Diffusion Models
von: Gu, Jiatao, et al.
Veröffentlicht: (2023)
von: Gu, Jiatao, et al.
Veröffentlicht: (2023)
The Coupling Within: Flow Matching via Distilled Normalizing Flows
von: Berthelot, David, et al.
Veröffentlicht: (2026)
von: Berthelot, David, et al.
Veröffentlicht: (2026)
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
von: Bachmann, Roman, et al.
Veröffentlicht: (2025)
von: Bachmann, Roman, et al.
Veröffentlicht: (2025)
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
von: Zhang, Ruixiang, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixiang, et al.
Veröffentlicht: (2025)
Improving GFlowNets for Text-to-Image Diffusion Alignment
von: Zhang, Dinghuai, et al.
Veröffentlicht: (2024)
von: Zhang, Dinghuai, et al.
Veröffentlicht: (2024)
Normalizing Flows are Capable Generative Models
von: Zhai, Shuangfei, et al.
Veröffentlicht: (2024)
von: Zhai, Shuangfei, et al.
Veröffentlicht: (2024)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
von: Sampaio, Georgia Gabriela, et al.
Veröffentlicht: (2024)
von: Sampaio, Georgia Gabriela, et al.
Veröffentlicht: (2024)
Normalizing Trajectory Models
von: Gu, Jiatao, et al.
Veröffentlicht: (2026)
von: Gu, Jiatao, et al.
Veröffentlicht: (2026)
Pseudo-Generalized Dynamic View Synthesis from a Video
von: Zhao, Xiaoming, et al.
Veröffentlicht: (2023)
von: Zhao, Xiaoming, et al.
Veröffentlicht: (2023)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
von: Crabbé, Jonathan, et al.
Veröffentlicht: (2023)
von: Crabbé, Jonathan, et al.
Veröffentlicht: (2023)
CTRLorALTer: Conditional LoRAdapter for Efficient 0-Shot Control & Altering of T2I Models
von: Stracke, Nick, et al.
Veröffentlicht: (2024)
von: Stracke, Nick, et al.
Veröffentlicht: (2024)
How Far Are We from Intelligent Visual Deductive Reasoning?
von: Zhang, Yizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yizhe, et al.
Veröffentlicht: (2024)
Datasets, Documents, and Repetitions: The Practicalities of Unequal Data Quality
von: Fang, Alex, et al.
Veröffentlicht: (2025)
von: Fang, Alex, et al.
Veröffentlicht: (2025)
Generative Pre-trained Autoregressive Diffusion Transformer
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
Towards Scalable Language-Image Pre-training for 3D Medical Imaging
von: Zhao, Chenhui, et al.
Veröffentlicht: (2025)
von: Zhao, Chenhui, et al.
Veröffentlicht: (2025)
An Empirical Study of Autoregressive Pre-training from Videos
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
Scalable Autoregressive Image Generation with Mamba
von: Li, Haopeng, et al.
Veröffentlicht: (2024)
von: Li, Haopeng, et al.
Veröffentlicht: (2024)
Swallowing the Bitter Pill: Simplified Scalable Conformer Generation
von: Wang, Yuyang, et al.
Veröffentlicht: (2023)
von: Wang, Yuyang, et al.
Veröffentlicht: (2023)
Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration
von: Dudhane, Akshay, et al.
Veröffentlicht: (2024)
von: Dudhane, Akshay, et al.
Veröffentlicht: (2024)
Learning Long-term Motion Embeddings for Efficient Kinematics Generation
von: Stracke, Nick, et al.
Veröffentlicht: (2026)
von: Stracke, Nick, et al.
Veröffentlicht: (2026)
From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs
von: Li, Mingxiao, et al.
Veröffentlicht: (2025)
von: Li, Mingxiao, et al.
Veröffentlicht: (2025)
Towards Scalable Pre-training of Visual Tokenizers for Generation
von: Yao, Jingfeng, et al.
Veröffentlicht: (2025)
von: Yao, Jingfeng, et al.
Veröffentlicht: (2025)
Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
von: Gui, Ming, et al.
Veröffentlicht: (2025)
von: Gui, Ming, et al.
Veröffentlicht: (2025)
Self-supervised Pre-training of Text Recognizers
von: Kišš, Martin, et al.
Veröffentlicht: (2024)
von: Kišš, Martin, et al.
Veröffentlicht: (2024)
DINOv2: Learning Robust Visual Features without Supervision
von: Oquab, Maxime, et al.
Veröffentlicht: (2023)
von: Oquab, Maxime, et al.
Veröffentlicht: (2023)
Self-Supervised Learning for Pre-training Capsule Networks: Overcoming Medical Imaging Dataset Challenges
von: El-Shimy, Heba, et al.
Veröffentlicht: (2025)
von: El-Shimy, Heba, et al.
Veröffentlicht: (2025)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
von: Sun, Peize, et al.
Veröffentlicht: (2024)
von: Sun, Peize, et al.
Veröffentlicht: (2024)
Generative Modeling with Phase Stochastic Bridges
von: Chen, Tianrong, et al.
Veröffentlicht: (2023)
von: Chen, Tianrong, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Multimodal Autoregressive Pre-training of Large Vision Encoders
von: Fini, Enrico, et al.
Veröffentlicht: (2024) -
World-consistent Video Diffusion with Explicit 3D Modeling
von: Zhang, Qihang, et al.
Veröffentlicht: (2024) -
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2023) -
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
von: Gu, Jiatao, et al.
Veröffentlicht: (2024) -
Scaling Laws for Native Multimodal Models
von: Shukor, Mustafa, et al.
Veröffentlicht: (2025)