Emerging Properties in Unified Multimodal Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Chaorui, Zhu, Deyao, Li, Kunchang, Gou, Chenhui, Li, Feng, Wang, Zeyu, Zhong, Shu, Yu, Weihao, Nie, Xiaonan, Song, Ziang, Shi, Guang, Fan, Haoqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
Causal Diffusion Transformers for Generative Modeling
von: Deng, Chaorui, et al.
Veröffentlicht: (2024)
von: Deng, Chaorui, et al.
Veröffentlicht: (2024)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
Understanding and Harnessing Sparsity in Unified Multimodal Models
von: He, Shwai, et al.
Veröffentlicht: (2025)
von: He, Shwai, et al.
Veröffentlicht: (2025)
Context Unrolling in Omni Models
von: Yang, Ceyuan, et al.
Veröffentlicht: (2026)
von: Yang, Ceyuan, et al.
Veröffentlicht: (2026)
Strong and Controllable Blind Image Decomposition
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
How Well Can Vision Language Models See Image Details?
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
Real‐Time Energy Management Strategy for Fuel Cell/Battery Plug‐In Hybrid Electric Buses Based on Deep Reinforcement Learning and State of Charge Descent Curve Trajectory Control
von: Jing Lian, et al.
Veröffentlicht: (2024)
von: Jing Lian, et al.
Veröffentlicht: (2024)
Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model
von: Chen, Xiaolin, et al.
Veröffentlicht: (2022)
von: Chen, Xiaolin, et al.
Veröffentlicht: (2022)
Harvest Video Foundation Models via Efficient Post-Pretraining
von: Li, Yizhuo, et al.
Veröffentlicht: (2023)
von: Li, Yizhuo, et al.
Veröffentlicht: (2023)
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
von: Zhang, Yabo, et al.
Veröffentlicht: (2026)
von: Zhang, Yabo, et al.
Veröffentlicht: (2026)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
von: Li, Kunchang, et al.
Veröffentlicht: (2022)
von: Li, Kunchang, et al.
Veröffentlicht: (2022)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
UMBRAE: Unified Multimodal Brain Decoding
von: Xia, Weihao, et al.
Veröffentlicht: (2024)
von: Xia, Weihao, et al.
Veröffentlicht: (2024)
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
von: Chen, Feng, et al.
Veröffentlicht: (2024)
von: Chen, Feng, et al.
Veröffentlicht: (2024)
Primitive elements in Ringel-Hall algebras of tame hereditary algebras
von: Deng, Bangming, et al.
Veröffentlicht: (2026)
von: Deng, Bangming, et al.
Veröffentlicht: (2026)
TokenUnify: Scaling Up Autoregressive Pretraining for Neuron Segmentation
von: Chen, Yinda, et al.
Veröffentlicht: (2024)
von: Chen, Yinda, et al.
Veröffentlicht: (2024)
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
von: Xie, Wulin, et al.
Veröffentlicht: (2025)
von: Xie, Wulin, et al.
Veröffentlicht: (2025)
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
von: Niu, Yuwei, et al.
Veröffentlicht: (2025)
von: Niu, Yuwei, et al.
Veröffentlicht: (2025)
ID-NeRF: Indirect Diffusion-guided Neural Radiance Fields for Generalizable View Synthesis
von: Li, Yaokun, et al.
Veröffentlicht: (2024)
von: Li, Yaokun, et al.
Veröffentlicht: (2024)
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
von: Chen, Boyu, et al.
Veröffentlicht: (2024)
von: Chen, Boyu, et al.
Veröffentlicht: (2024)
Node Level Graph Autoencoder: Unified Pretraining for Textual Graph Learning
von: Hu, Wenbin, et al.
Veröffentlicht: (2024)
von: Hu, Wenbin, et al.
Veröffentlicht: (2024)
Representation Forcing for Bottleneck-Free Unified Multimodal Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2026)
von: Wang, Yuqing, et al.
Veröffentlicht: (2026)
TUNI: Real-time RGB-T Semantic Segmentation with Unified Multi-Modal Feature Extraction and Cross-Modal Feature Fusion
von: Guo, Xiaodong, et al.
Veröffentlicht: (2025)
von: Guo, Xiaodong, et al.
Veröffentlicht: (2025)
Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
von: Feng, Yicheng, et al.
Veröffentlicht: (2025)
von: Feng, Yicheng, et al.
Veröffentlicht: (2025)
How Does Controllability Emerge In Language Models During Pretraining?
von: She, Jianshu, et al.
Veröffentlicht: (2025)
von: She, Jianshu, et al.
Veröffentlicht: (2025)
MTPNet: Multi-Grained Target Perception for Unified Activity Cliff Prediction
von: Shu, Zishan, et al.
Veröffentlicht: (2025)
von: Shu, Zishan, et al.
Veröffentlicht: (2025)
Emu: Generative Pretraining in Multimodality
von: Sun, Quan, et al.
Veröffentlicht: (2023)
von: Sun, Quan, et al.
Veröffentlicht: (2023)
Task-Centric Policy Optimization from Misaligned Motion Priors
von: Zheng, Ziang, et al.
Veröffentlicht: (2026)
von: Zheng, Ziang, et al.
Veröffentlicht: (2026)
UniDiff: A Unified Diffusion Framework for Multimodal Time Series Forecasting
von: Zhang, Da, et al.
Veröffentlicht: (2025)
von: Zhang, Da, et al.
Veröffentlicht: (2025)
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
von: Luo, Hao, et al.
Veröffentlicht: (2026)
von: Luo, Hao, et al.
Veröffentlicht: (2026)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
Valuation of variable annuities under the Volterra mortality and rough Heston models
von: Li, Wenyuan, et al.
Veröffentlicht: (2026)
von: Li, Wenyuan, et al.
Veröffentlicht: (2026)
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
DrVideo: Document Retrieval Based Long Video Understanding
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
von: Wang, Zeyu, et al.
Veröffentlicht: (2025) -
Causal Diffusion Transformers for Generative Modeling
von: Deng, Chaorui, et al.
Veröffentlicht: (2024) -
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
von: Gou, Chenhui, et al.
Veröffentlicht: (2025) -
Understanding and Harnessing Sparsity in Unified Multimodal Models
von: He, Shwai, et al.
Veröffentlicht: (2025) -
Context Unrolling in Omni Models
von: Yang, Ceyuan, et al.
Veröffentlicht: (2026)