Transition Models: Rethinking the Generative Learning Objective
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zidong, Zhang, Yiyuan, Yue, Xiaoyu, Yue, Xiangyu, Li, Yangguang, Ouyang, Wanli, Bai, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Native-Resolution Image Synthesis
by: Wang, Zidong, et al.
Published: (2025)
by: Wang, Zidong, et al.
Published: (2025)
EMR-Merging: Tuning-Free High-Performance Model Merging
by: Huang, Chenyu, et al.
Published: (2024)
by: Huang, Chenyu, et al.
Published: (2024)
Diffusion Models Need Visual Priors for Image Generation
by: Yue, Xiaoyu, et al.
Published: (2024)
by: Yue, Xiaoyu, et al.
Published: (2024)
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
by: Yue, Xiaoyu, et al.
Published: (2025)
by: Yue, Xiaoyu, et al.
Published: (2025)
Scaling Up Your Kernels: Large Kernel Design in ConvNets towards Universal Representations
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
by: Zhang, Mingyuan, et al.
Published: (2025)
by: Zhang, Mingyuan, et al.
Published: (2025)
Explore the Limits of Omni-modal Pretraining at Scale
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
by: Zhang, Zhixin, et al.
Published: (2024)
by: Zhang, Zhixin, et al.
Published: (2024)
Multimodal Long Video Modeling Based on Temporal Dynamic Context
by: Hao, Haoran, et al.
Published: (2025)
by: Hao, Haoran, et al.
Published: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Dynamic Base model Shift for Delta Compression
by: Huang, Chenyu, et al.
Published: (2025)
by: Huang, Chenyu, et al.
Published: (2025)
Rethinking Multi-domain Generalization with A General Learning Objective
by: Tan, Zhaorui, et al.
Published: (2024)
by: Tan, Zhaorui, et al.
Published: (2024)
Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
by: Zhu, Haoyi, et al.
Published: (2024)
by: Zhu, Haoyi, et al.
Published: (2024)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
PredBench: Benchmarking Spatio-Temporal Prediction across Diverse Disciplines
by: Wang, ZiDong, et al.
Published: (2024)
by: Wang, ZiDong, et al.
Published: (2024)
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
by: Ding, Xiaohan, et al.
Published: (2023)
by: Ding, Xiaohan, et al.
Published: (2023)
Approximating Signed Distance Fields With Sparse Ellipsoidal Radial Basis Function Networks: A Dynamic Multi-Objective Optimization Strategy
by: Lian, Bobo, et al.
Published: (2025)
by: Lian, Bobo, et al.
Published: (2025)
Rethinking the Evaluation Protocol of Domain Generalization
by: Yu, Han, et al.
Published: (2023)
by: Yu, Han, et al.
Published: (2023)
SupMAE: Supervised Masked Autoencoders Are Efficient Vision Learners
by: Liang, Feng, et al.
Published: (2022)
by: Liang, Feng, et al.
Published: (2022)
WeatherGFM: Learning A Weather Generalist Foundation Model via In-context Learning
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
FiT: Flexible Vision Transformer for Diffusion Model
by: Lu, Zeyu, et al.
Published: (2024)
by: Lu, Zeyu, et al.
Published: (2024)
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
by: Bai, Andrew, et al.
Published: (2025)
by: Bai, Andrew, et al.
Published: (2025)
Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression
by: Wang, Xiaohui, et al.
Published: (2025)
by: Wang, Xiaohui, et al.
Published: (2025)
Rethink Arbitrary Style Transfer with Transformer and Contrastive Learning
by: Zhang, Zhanjie, et al.
Published: (2024)
by: Zhang, Zhanjie, et al.
Published: (2024)
AsyCo: An Asymmetric Dual-task Co-training Model for Partial-label Learning
by: Li, Beibei, et al.
Published: (2024)
by: Li, Beibei, et al.
Published: (2024)
Transitive Vision-Language Prompt Learning for Domain Generalization
by: Wang, Liyuan, et al.
Published: (2024)
by: Wang, Liyuan, et al.
Published: (2024)
Divide-and-Conquer for Enhancing Unlabeled Learning, Stability, and Plasticity in Semi-supervised Continual Learning
by: Duan, Yue, et al.
Published: (2025)
by: Duan, Yue, et al.
Published: (2025)
GUPNet++: Geometry Uncertainty Propagation Network for Monocular 3D Object Detection
by: Lu, Yan, et al.
Published: (2023)
by: Lu, Yan, et al.
Published: (2023)
Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning
by: Yang, Suorong, et al.
Published: (2025)
by: Yang, Suorong, et al.
Published: (2025)
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
by: Li, Lingxiao, et al.
Published: (2025)
by: Li, Lingxiao, et al.
Published: (2025)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
An Adaptor for Triggering Semi-Supervised Learning to Out-of-Box Serve Deep Image Clustering
by: Duan, Yue, et al.
Published: (2025)
by: Duan, Yue, et al.
Published: (2025)
Growing Visual Generative Capacity for Pre-Trained MLLMs
by: Wang, Hanyu, et al.
Published: (2025)
by: Wang, Hanyu, et al.
Published: (2025)
Block Flow: Learning Straight Flow on Data Blocks
by: Wang, Zibin, et al.
Published: (2025)
by: Wang, Zibin, et al.
Published: (2025)
Bayesian Diffusion Models for 3D Shape Reconstruction
by: Xu, Haiyang, et al.
Published: (2024)
by: Xu, Haiyang, et al.
Published: (2024)
UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
by: Tang, Chen, et al.
Published: (2025)
by: Tang, Chen, et al.
Published: (2025)
ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
by: Li, Junxian, et al.
Published: (2024)
by: Li, Junxian, et al.
Published: (2024)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
Rethinking Meta-Learning from a Learning Lens
by: Wang, Jingyao, et al.
Published: (2024)
by: Wang, Jingyao, et al.
Published: (2024)
Similar Items
-
Native-Resolution Image Synthesis
by: Wang, Zidong, et al.
Published: (2025) -
EMR-Merging: Tuning-Free High-Performance Model Merging
by: Huang, Chenyu, et al.
Published: (2024) -
Diffusion Models Need Visual Priors for Image Generation
by: Yue, Xiaoyu, et al.
Published: (2024) -
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
by: Yue, Xiaoyu, et al.
Published: (2025) -
Scaling Up Your Kernels: Large Kernel Design in ConvNets towards Universal Representations
by: Zhang, Yiyuan, et al.
Published: (2024)