Scalable Pre-training of Large Autoregressive Image Models
Fuente:
arXiv
Guardado en:
| Autores principales: | El-Nouby, Alaaeldin, Klein, Michal, Zhai, Shuangfei, Bautista, Miguel Angel, Toshev, Alexander, Shankar, Vaishaal, Susskind, Joshua M, Joulin, Armand |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multimodal Autoregressive Pre-training of Large Vision Encoders
por: Fini, Enrico, et al.
Publicado: (2024)
por: Fini, Enrico, et al.
Publicado: (2024)
World-consistent Video Diffusion with Explicit 3D Modeling
por: Zhang, Qihang, et al.
Publicado: (2024)
por: Zhang, Qihang, et al.
Publicado: (2024)
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
por: Zhang, Yuhui, et al.
Publicado: (2023)
por: Zhang, Yuhui, et al.
Publicado: (2023)
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
por: Gu, Jiatao, et al.
Publicado: (2024)
por: Gu, Jiatao, et al.
Publicado: (2024)
Scaling Laws for Native Multimodal Models
por: Shukor, Mustafa, et al.
Publicado: (2025)
por: Shukor, Mustafa, et al.
Publicado: (2025)
Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
por: Gu, Jiatao, et al.
Publicado: (2024)
por: Gu, Jiatao, et al.
Publicado: (2024)
Many-to-many Image Generation with Auto-regressive Diffusion Models
por: Shen, Ying, et al.
Publicado: (2024)
por: Shen, Ying, et al.
Publicado: (2024)
Normalizing Flows with Iterative Denoising
por: Chen, Tianrong, et al.
Publicado: (2026)
por: Chen, Tianrong, et al.
Publicado: (2026)
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
por: Shen, Ying, et al.
Publicado: (2026)
por: Shen, Ying, et al.
Publicado: (2026)
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
por: Gu, Jiatao, et al.
Publicado: (2025)
por: Gu, Jiatao, et al.
Publicado: (2025)
MobileCLIP2: Improving Multi-Modal Reinforced Training
por: Faghri, Fartash, et al.
Publicado: (2025)
por: Faghri, Fartash, et al.
Publicado: (2025)
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
por: Gu, Jiatao, et al.
Publicado: (2025)
por: Gu, Jiatao, et al.
Publicado: (2025)
Matryoshka Diffusion Models
por: Gu, Jiatao, et al.
Publicado: (2023)
por: Gu, Jiatao, et al.
Publicado: (2023)
The Coupling Within: Flow Matching via Distilled Normalizing Flows
por: Berthelot, David, et al.
Publicado: (2026)
por: Berthelot, David, et al.
Publicado: (2026)
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
por: Bachmann, Roman, et al.
Publicado: (2025)
por: Bachmann, Roman, et al.
Publicado: (2025)
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
por: Zhang, Ruixiang, et al.
Publicado: (2025)
por: Zhang, Ruixiang, et al.
Publicado: (2025)
Improving GFlowNets for Text-to-Image Diffusion Alignment
por: Zhang, Dinghuai, et al.
Publicado: (2024)
por: Zhang, Dinghuai, et al.
Publicado: (2024)
Normalizing Flows are Capable Generative Models
por: Zhai, Shuangfei, et al.
Publicado: (2024)
por: Zhai, Shuangfei, et al.
Publicado: (2024)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
por: Sampaio, Georgia Gabriela, et al.
Publicado: (2024)
por: Sampaio, Georgia Gabriela, et al.
Publicado: (2024)
Normalizing Trajectory Models
por: Gu, Jiatao, et al.
Publicado: (2026)
por: Gu, Jiatao, et al.
Publicado: (2026)
Pseudo-Generalized Dynamic View Synthesis from a Video
por: Zhao, Xiaoming, et al.
Publicado: (2023)
por: Zhao, Xiaoming, et al.
Publicado: (2023)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
por: Crabbé, Jonathan, et al.
Publicado: (2023)
por: Crabbé, Jonathan, et al.
Publicado: (2023)
CTRLorALTer: Conditional LoRAdapter for Efficient 0-Shot Control & Altering of T2I Models
por: Stracke, Nick, et al.
Publicado: (2024)
por: Stracke, Nick, et al.
Publicado: (2024)
How Far Are We from Intelligent Visual Deductive Reasoning?
por: Zhang, Yizhe, et al.
Publicado: (2024)
por: Zhang, Yizhe, et al.
Publicado: (2024)
Datasets, Documents, and Repetitions: The Practicalities of Unequal Data Quality
por: Fang, Alex, et al.
Publicado: (2025)
por: Fang, Alex, et al.
Publicado: (2025)
Generative Pre-trained Autoregressive Diffusion Transformer
por: Zhang, Yuan, et al.
Publicado: (2025)
por: Zhang, Yuan, et al.
Publicado: (2025)
Towards Scalable Language-Image Pre-training for 3D Medical Imaging
por: Zhao, Chenhui, et al.
Publicado: (2025)
por: Zhao, Chenhui, et al.
Publicado: (2025)
An Empirical Study of Autoregressive Pre-training from Videos
por: Rajasegaran, Jathushan, et al.
Publicado: (2025)
por: Rajasegaran, Jathushan, et al.
Publicado: (2025)
Scalable Autoregressive Image Generation with Mamba
por: Li, Haopeng, et al.
Publicado: (2024)
por: Li, Haopeng, et al.
Publicado: (2024)
Swallowing the Bitter Pill: Simplified Scalable Conformer Generation
por: Wang, Yuyang, et al.
Publicado: (2023)
por: Wang, Yuyang, et al.
Publicado: (2023)
Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration
por: Dudhane, Akshay, et al.
Publicado: (2024)
por: Dudhane, Akshay, et al.
Publicado: (2024)
Learning Long-term Motion Embeddings for Efficient Kinematics Generation
por: Stracke, Nick, et al.
Publicado: (2026)
por: Stracke, Nick, et al.
Publicado: (2026)
From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs
por: Li, Mingxiao, et al.
Publicado: (2025)
por: Li, Mingxiao, et al.
Publicado: (2025)
Towards Scalable Pre-training of Visual Tokenizers for Generation
por: Yao, Jingfeng, et al.
Publicado: (2025)
por: Yao, Jingfeng, et al.
Publicado: (2025)
Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
por: Gui, Ming, et al.
Publicado: (2025)
por: Gui, Ming, et al.
Publicado: (2025)
Self-supervised Pre-training of Text Recognizers
por: Kišš, Martin, et al.
Publicado: (2024)
por: Kišš, Martin, et al.
Publicado: (2024)
DINOv2: Learning Robust Visual Features without Supervision
por: Oquab, Maxime, et al.
Publicado: (2023)
por: Oquab, Maxime, et al.
Publicado: (2023)
Self-Supervised Learning for Pre-training Capsule Networks: Overcoming Medical Imaging Dataset Challenges
por: El-Shimy, Heba, et al.
Publicado: (2025)
por: El-Shimy, Heba, et al.
Publicado: (2025)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
por: Sun, Peize, et al.
Publicado: (2024)
por: Sun, Peize, et al.
Publicado: (2024)
Generative Modeling with Phase Stochastic Bridges
por: Chen, Tianrong, et al.
Publicado: (2023)
por: Chen, Tianrong, et al.
Publicado: (2023)
Ejemplares similares
-
Multimodal Autoregressive Pre-training of Large Vision Encoders
por: Fini, Enrico, et al.
Publicado: (2024) -
World-consistent Video Diffusion with Explicit 3D Modeling
por: Zhang, Qihang, et al.
Publicado: (2024) -
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
por: Zhang, Yuhui, et al.
Publicado: (2023) -
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
por: Gu, Jiatao, et al.
Publicado: (2024) -
Scaling Laws for Native Multimodal Models
por: Shukor, Mustafa, et al.
Publicado: (2025)