HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Yicheng, Song, Lin, Yang, Rui, Cheng, Cheng, Xu, Zunnan, Zhang, Zhaoyang, Ge, Yixiao, Li, Xiu, Shan, Ying |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
GrootVL: Tree Topology is All You Need in State Space Model
by: Xiao, Yicheng, et al.
Published: (2024)
by: Xiao, Yicheng, et al.
Published: (2024)
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025)
by: Tan, Zhiyu, et al.
Published: (2025)
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
by: Ge, Yuying, et al.
Published: (2024)
by: Ge, Yuying, et al.
Published: (2024)
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
by: Ge, Yuying, et al.
Published: (2024)
by: Ge, Yuying, et al.
Published: (2024)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
by: Cheng, Junhao, et al.
Published: (2025)
by: Cheng, Junhao, et al.
Published: (2025)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
by: Li, Lijiang, et al.
Published: (2026)
by: Li, Lijiang, et al.
Published: (2026)
OmniCam: Unified Multimodal Video Generation via Camera Control
by: Yang, Xiaoda, et al.
Published: (2025)
by: Yang, Xiaoda, et al.
Published: (2025)
Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
by: Liu, Jialun, et al.
Published: (2026)
by: Liu, Jialun, et al.
Published: (2026)
From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
by: Cheng, Cheng, et al.
Published: (2025)
by: Cheng, Cheng, et al.
Published: (2025)
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
by: Xie, Jinheng, et al.
Published: (2024)
by: Xie, Jinheng, et al.
Published: (2024)
OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
by: Xiao, Teng, et al.
Published: (2025)
by: Xiao, Teng, et al.
Published: (2025)
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
by: Li, Yizhuo, et al.
Published: (2024)
by: Li, Yizhuo, et al.
Published: (2024)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
SEED-Story: Multimodal Long Story Generation with Large Language Model
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026)
by: Zhou, Donghao, et al.
Published: (2026)
YOLO-World: Real-Time Open-Vocabulary Object Detection
by: Cheng, Tianheng, et al.
Published: (2024)
by: Cheng, Tianheng, et al.
Published: (2024)
Meta-Adapter: An Online Few-shot Learner for Vision-Language Model
by: Cheng, Cheng, et al.
Published: (2023)
by: Cheng, Cheng, et al.
Published: (2023)
Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and Generation
by: Zhou, Zhiwang, et al.
Published: (2025)
by: Zhou, Zhiwang, et al.
Published: (2025)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
by: Cheng, Junhao, et al.
Published: (2025)
by: Cheng, Junhao, et al.
Published: (2025)
SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning
by: Huang, Jiaqi, et al.
Published: (2025)
by: Huang, Jiaqi, et al.
Published: (2025)
Ming-Omni: A Unified Multimodal Model for Perception and Generation
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
by: Qiu, Lu, et al.
Published: (2025)
by: Qiu, Lu, et al.
Published: (2025)
ViT-Lens: Towards Omni-modal Representations
by: Lei, Weixian, et al.
Published: (2023)
by: Lei, Weixian, et al.
Published: (2023)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
by: Lin, Haokun, et al.
Published: (2025)
by: Lin, Haokun, et al.
Published: (2025)
Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning
by: Cheng, Dongjie, et al.
Published: (2026)
by: Cheng, Dongjie, et al.
Published: (2026)
OmniPSD: Layered PSD Generation with Diffusion Transformer
by: Liu, Cheng, et al.
Published: (2025)
by: Liu, Cheng, et al.
Published: (2025)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
by: Wang, Le, et al.
Published: (2025)
by: Wang, Le, et al.
Published: (2025)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights
by: Lei, Weixian, et al.
Published: (2023)
by: Lei, Weixian, et al.
Published: (2023)
OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
by: Zhang, Guohui, et al.
Published: (2026)
by: Zhang, Guohui, et al.
Published: (2026)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
by: Liu, Ruyang, et al.
Published: (2023)
by: Liu, Ruyang, et al.
Published: (2023)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
Unified Reward Model for Multimodal Understanding and Generation
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Similar Items
-
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
by: Yang, Rui, et al.
Published: (2025) -
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
by: Xiao, Yicheng, et al.
Published: (2025) -
GrootVL: Tree Topology is All You Need in State Space Model
by: Xiao, Yicheng, et al.
Published: (2024) -
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025) -
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
by: Ge, Yuying, et al.
Published: (2024)