Rethinking Training Dynamics in Scale-wise Autoregressive Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Gengze, Ge, Chongjian, Tan, Hao, Liu, Feng, Hong, Yicong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
CapsFusion: Rethinking Image-Text Data at Scale
di: Yu, Qiying, et al.
Pubblicazione: (2023)
di: Yu, Qiying, et al.
Pubblicazione: (2023)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
di: Yu, Hao, et al.
Pubblicazione: (2025)
di: Yu, Hao, et al.
Pubblicazione: (2025)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
di: Li, Zerui, et al.
Pubblicazione: (2025)
di: Li, Zerui, et al.
Pubblicazione: (2025)
MoVE: Mixture of Value Embeddings -- A New Axis for Scaling Parametric Memory in Autoregressive Models
di: Li, Yangyan
Pubblicazione: (2026)
di: Li, Yangyan
Pubblicazione: (2026)
DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
di: Hong, Susung, et al.
Pubblicazione: (2025)
di: Hong, Susung, et al.
Pubblicazione: (2025)
Universal Approximation of Visual Autoregressive Transformers
di: Chen, Yifang, et al.
Pubblicazione: (2025)
di: Chen, Yifang, et al.
Pubblicazione: (2025)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
Rethinking Machine Unlearning in Image Generation Models
di: Liu, Renyang, et al.
Pubblicazione: (2025)
di: Liu, Renyang, et al.
Pubblicazione: (2025)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
di: Kou, Siqi, et al.
Pubblicazione: (2024)
di: Kou, Siqi, et al.
Pubblicazione: (2024)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
di: Lokegaonkar, Vaibhavi, et al.
Pubblicazione: (2026)
di: Lokegaonkar, Vaibhavi, et al.
Pubblicazione: (2026)
Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
di: Lin, Shanchuan, et al.
Pubblicazione: (2025)
di: Lin, Shanchuan, et al.
Pubblicazione: (2025)
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
di: Zhang, Jingyi, et al.
Pubblicazione: (2025)
di: Zhang, Jingyi, et al.
Pubblicazione: (2025)
Diving into Self-Evolving Training for Multimodal Reasoning
di: Liu, Wei, et al.
Pubblicazione: (2024)
di: Liu, Wei, et al.
Pubblicazione: (2024)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
di: Xiong, Weimin, et al.
Pubblicazione: (2026)
di: Xiong, Weimin, et al.
Pubblicazione: (2026)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
di: Chern, Ethan, et al.
Pubblicazione: (2024)
di: Chern, Ethan, et al.
Pubblicazione: (2024)
Rethinking Chain-of-Thought Reasoning for Videos
di: Zhong, Yiwu, et al.
Pubblicazione: (2025)
di: Zhong, Yiwu, et al.
Pubblicazione: (2025)
Progressive Autoregressive Video Diffusion Models
di: Xie, Desai, et al.
Pubblicazione: (2024)
di: Xie, Desai, et al.
Pubblicazione: (2024)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
di: Wang, Andrew Z., et al.
Pubblicazione: (2025)
di: Wang, Andrew Z., et al.
Pubblicazione: (2025)
Exploring Curriculum Learning for Vision-Language Tasks: A Study on Small-Scale Multimodal Training
di: Saha, Rohan, et al.
Pubblicazione: (2024)
di: Saha, Rohan, et al.
Pubblicazione: (2024)
AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation
di: Zhu, Jiayin, et al.
Pubblicazione: (2025)
di: Zhu, Jiayin, et al.
Pubblicazione: (2025)
Rethinking Genomic Modeling Through Optical Character Recognition
di: Xiang, Hongxin, et al.
Pubblicazione: (2026)
di: Xiang, Hongxin, et al.
Pubblicazione: (2026)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
di: Zhao, Haozhe, et al.
Pubblicazione: (2025)
di: Zhao, Haozhe, et al.
Pubblicazione: (2025)
When Diffusion Breaks Constraints: Sequential Autoregressive Generation with RL and MCTS
di: Zhao, Zirui, et al.
Pubblicazione: (2025)
di: Zhao, Zirui, et al.
Pubblicazione: (2025)
A General and Efficient Training for Transformer via Token Expansion
di: Huang, Wenxuan, et al.
Pubblicazione: (2024)
di: Huang, Wenxuan, et al.
Pubblicazione: (2024)
The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
di: Yin, Aoxiong, et al.
Pubblicazione: (2025)
di: Yin, Aoxiong, et al.
Pubblicazione: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
di: Zhang, Jun, et al.
Pubblicazione: (2025)
di: Zhang, Jun, et al.
Pubblicazione: (2025)
REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
di: Zhang, Yitian, et al.
Pubblicazione: (2025)
di: Zhang, Yitian, et al.
Pubblicazione: (2025)
Test-Time Training Done Right
di: Zhang, Tianyuan, et al.
Pubblicazione: (2025)
di: Zhang, Tianyuan, et al.
Pubblicazione: (2025)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
di: Zheng, Naishan, et al.
Pubblicazione: (2025)
di: Zheng, Naishan, et al.
Pubblicazione: (2025)
Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge
di: Lin, Yuanze, et al.
Pubblicazione: (2024)
di: Lin, Yuanze, et al.
Pubblicazione: (2024)
Can World Models Benefit VLMs for World Dynamics?
di: Zhang, Kevin, et al.
Pubblicazione: (2025)
di: Zhang, Kevin, et al.
Pubblicazione: (2025)
LRM: Large Reconstruction Model for Single Image to 3D
di: Hong, Yicong, et al.
Pubblicazione: (2023)
di: Hong, Yicong, et al.
Pubblicazione: (2023)
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
di: Huang, Chengyue, et al.
Pubblicazione: (2025)
di: Huang, Chengyue, et al.
Pubblicazione: (2025)
Astra: General Interactive World Model with Autoregressive Denoising
di: Zhu, Yixuan, et al.
Pubblicazione: (2025)
di: Zhu, Yixuan, et al.
Pubblicazione: (2025)
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
di: Xiao, Zhongyu, et al.
Pubblicazione: (2026)
di: Xiao, Zhongyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
di: Zhou, Gengze, et al.
Pubblicazione: (2024) -
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
di: Zhou, Gengze, et al.
Pubblicazione: (2024) -
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
di: Yu, Shoubin, et al.
Pubblicazione: (2025) -
CapsFusion: Rethinking Image-Text Data at Scale
di: Yu, Qiying, et al.
Pubblicazione: (2023) -
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
di: Yu, Hao, et al.
Pubblicazione: (2025)