Emu3.5: Native Multimodal Models are World Learners
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Yufeng, Chen, Honghao, Deng, Haoge, Huang, Xu, Li, Xinghang, Liu, Jirong, Liu, Yang, Luo, Zhuoyan, Wang, Jinsheng, Wang, Wenxuan, Wang, Yueze, Wang, Chengyuan, Zhang, Fan, Zhao, Yingli, Pan, Ting, Li, Xianduo, Hao, Zecheng, Ma, Wenxuan, Chen, Zhuo, Ao, Yulong, Huang, Tiejun, Wang, Zhongyuan, Wang, Xinlong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emu: Generative Pretraining in Multimodality
by: Sun, Quan, et al.
Published: (2023)
by: Sun, Quan, et al.
Published: (2023)
Emu3: Next-Token Prediction is All You Need
by: Wang, Xinlong, et al.
Published: (2024)
by: Wang, Xinlong, et al.
Published: (2024)
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2025)
by: Diao, Haiwen, et al.
Published: (2025)
Generative Multimodal Models are In-Context Learners
by: Sun, Quan, et al.
Published: (2023)
by: Sun, Quan, et al.
Published: (2023)
Uniform Discrete Diffusion with Metric Path for Video Generation
by: Deng, Haoge, et al.
Published: (2025)
by: Deng, Haoge, et al.
Published: (2025)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
Ge$^\text{2}$mS-T: Multi-Dimensional Grouping for Ultra-High Energy Efficiency in Spiking Transformer
by: Hao, Zecheng, et al.
Published: (2026)
by: Hao, Zecheng, et al.
Published: (2026)
Scaling World Model for Hierarchical Manipulation Policies
by: Long, Qian, et al.
Published: (2026)
by: Long, Qian, et al.
Published: (2026)
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
by: Ma, Baorui, et al.
Published: (2024)
by: Ma, Baorui, et al.
Published: (2024)
Unveiling Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
Rethinking SNN Online Training and Deployment: Gradient-Coherent Learning via Hybrid-Driven LIF Model
by: Hao, Zecheng, et al.
Published: (2024)
by: Hao, Zecheng, et al.
Published: (2024)
On the monotonicity of affine quermassintegrals
by: Chen, Shibing, et al.
Published: (2026)
by: Chen, Shibing, et al.
Published: (2026)
Uniqueness of Blow-ups for the Superconductivity Free Boundary Problem
by: Chen, Shibing, et al.
Published: (2026)
by: Chen, Shibing, et al.
Published: (2026)
QGHNN: A quantum graph Hamiltonian neural network
by: Wang, Wenxuan
Published: (2025)
by: Wang, Wenxuan
Published: (2025)
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
by: Wang, Wenxuan
Published: (2024)
by: Wang, Wenxuan
Published: (2024)
Noise-resistant adaptive Hamiltonian learning
by: Wang, Wenxuan
Published: (2025)
by: Wang, Wenxuan
Published: (2025)
Uncertainty-Aware Token Importance Estimation in Spiking Transformers
by: Liu, Wenxuan, et al.
Published: (2026)
by: Liu, Wenxuan, et al.
Published: (2026)
Diffusion Feedback Helps CLIP See Better
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
by: Chen, Honghao, et al.
Published: (2025)
by: Chen, Honghao, et al.
Published: (2025)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
by: Wu, Chenyuan, et al.
Published: (2025)
by: Wu, Chenyuan, et al.
Published: (2025)
Confidential Databases Without Cryptographic Mappings
by: Huang, Wenxuan, et al.
Published: (2026)
by: Huang, Wenxuan, et al.
Published: (2026)
MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis
by: Chen, Wenting, et al.
Published: (2026)
by: Chen, Wenting, et al.
Published: (2026)
Differential Coding for Training-Free ANN-to-SNN Conversion
by: Huang, Zihan, et al.
Published: (2025)
by: Huang, Zihan, et al.
Published: (2025)
Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Image Difference Grounding with Natural Language
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Efficient Multimodal Learning from Data-centric Perspective
by: He, Muyang, et al.
Published: (2024)
by: He, Muyang, et al.
Published: (2024)
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
by: Sun, Quan, et al.
Published: (2024)
by: Sun, Quan, et al.
Published: (2024)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
SOTA: Spike-Navigated Optimal TrAnsport Saliency Region Detection in Composite-bias Videos
by: Liu, Wenxuan, et al.
Published: (2025)
by: Liu, Wenxuan, et al.
Published: (2025)
Agentic AIs Are the Missing Paradigm for Out-of-Distribution Generalization in Foundation Models
by: Wang, Xin, et al.
Published: (2026)
by: Wang, Xin, et al.
Published: (2026)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
by: Wang, Cheng, et al.
Published: (2026)
by: Wang, Cheng, et al.
Published: (2026)
Strong holomorphic Morse inequalities on non-compact complex manifolds with optimal fundamental estimate
by: Liu, Manli, et al.
Published: (2024)
by: Liu, Manli, et al.
Published: (2024)
Dynamic Textual Prompt For Rehearsal-free Lifelong Person Re-identification
by: Chen, Hongyu, et al.
Published: (2024)
by: Chen, Hongyu, et al.
Published: (2024)
EVA-02: A Visual Representation for Neon Genesis
by: Fang, Yuxin, et al.
Published: (2023)
by: Fang, Yuxin, et al.
Published: (2023)
SECURE: Stable Early Collision Understanding via Robust Embeddings in Autonomous Driving
by: Wang, Wenjing, et al.
Published: (2026)
by: Wang, Wenjing, et al.
Published: (2026)
Dynamic Dual-level Defense Routing for Continual Adversarial Training
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
What Matters in Building Vision-Language-Action Models for Generalist Robots
by: Li, Xinghang, et al.
Published: (2024)
by: Li, Xinghang, et al.
Published: (2024)
Similar Items
-
Emu: Generative Pretraining in Multimodality
by: Sun, Quan, et al.
Published: (2023) -
Emu3: Next-Token Prediction is All You Need
by: Wang, Xinlong, et al.
Published: (2024) -
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2025) -
Generative Multimodal Models are In-Context Learners
by: Sun, Quan, et al.
Published: (2023) -
Uniform Discrete Diffusion with Metric Path for Video Generation
by: Deng, Haoge, et al.
Published: (2025)