MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Tao, Wan, Xin, Chen, Taicai, Zhang, Rui, Pan, Junwen, Lu, Dawei, Lei, Fanding, Lu, Zhilin, Yang, Yunfei, Cheng, Chen, She, Qi, Liu, Chang, Sun, Zhenbang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MammothModa: Multi-Modal Large Language Model
by: She, Qi, et al.
Published: (2024)
by: She, Qi, et al.
Published: (2024)
TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding
by: Pan, Junwen, et al.
Published: (2025)
by: Pan, Junwen, et al.
Published: (2025)
Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
by: Shi, Yangming, et al.
Published: (2026)
by: Shi, Yangming, et al.
Published: (2026)
TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
by: Pan, Junwen, et al.
Published: (2025)
by: Pan, Junwen, et al.
Published: (2025)
ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
Divide-and-Conquer for Enhancing Unlabeled Learning, Stability, and Plasticity in Semi-supervised Continual Learning
by: Duan, Yue, et al.
Published: (2025)
by: Duan, Yue, et al.
Published: (2025)
UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
by: Bai, Chengyu, et al.
Published: (2025)
by: Bai, Chengyu, et al.
Published: (2025)
Mammoths and Neanderthals in the Thames Valley
by: Scott, Katharine
Published: (2025)
by: Scott, Katharine
Published: (2025)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
by: Zhang, Junwen, et al.
Published: (2025)
by: Zhang, Junwen, et al.
Published: (2025)
On the Faithfulness of Visual Thinking: Measurement and Enhancement
by: Liu, Zujing, et al.
Published: (2025)
by: Liu, Zujing, et al.
Published: (2025)
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding
by: Jiang, Yibo, et al.
Published: (2026)
by: Jiang, Yibo, et al.
Published: (2026)
Potentiometric surface of the Mammoth Cave Area
by: Bosch, Rachel
Published: (2021)
by: Bosch, Rachel
Published: (2021)
Will Cataloging Go the Way of the Wooly Mammoth?
by: Crockford, Sue
Published: (1976)
by: Crockford, Sue
Published: (1976)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
by: Li, Lijiang, et al.
Published: (2026)
by: Li, Lijiang, et al.
Published: (2026)
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
by: Zhang, Qizhe, et al.
Published: (2025)
by: Zhang, Qizhe, et al.
Published: (2025)
AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
by: AI, Inclusion, et al.
Published: (2026)
by: AI, Inclusion, et al.
Published: (2026)
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction
by: Ren, Xiaoming, et al.
Published: (2026)
by: Ren, Xiaoming, et al.
Published: (2026)
Vidi2.5: Large Multimodal Models for Video Understanding and Creation
by: Vidi Team, et al.
Published: (2025)
by: Vidi Team, et al.
Published: (2025)
On Pre-training of Multimodal Language Models Customized for Chart Understanding
by: Fan, Wan-Cyuan, et al.
Published: (2024)
by: Fan, Wan-Cyuan, et al.
Published: (2024)
X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation
by: Ma, Jian, et al.
Published: (2025)
by: Ma, Jian, et al.
Published: (2025)
A Two-step Estimating Approach for Heavy-tailed AR Models with Non-zero Median GARCH-type Noises
by: She, Rui, et al.
Published: (2025)
by: She, Rui, et al.
Published: (2025)
ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition
by: Xie, Weidong, et al.
Published: (2024)
by: Xie, Weidong, et al.
Published: (2024)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
by: Tian, Rui, et al.
Published: (2025)
by: Tian, Rui, et al.
Published: (2025)
Understanding and Harnessing Sparsity in Unified Multimodal Models
by: He, Shwai, et al.
Published: (2025)
by: He, Shwai, et al.
Published: (2025)
Low-Rank Similarity Mining for Multimodal Dataset Distillation
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
by: Xu, Lu, et al.
Published: (2024)
by: Xu, Lu, et al.
Published: (2024)
Vidi: Large Multimodal Models for Video Understanding and Editing
by: Vidi Team, et al.
Published: (2025)
by: Vidi Team, et al.
Published: (2025)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
by: Chen, Xiaokang, et al.
Published: (2025)
by: Chen, Xiaokang, et al.
Published: (2025)
Evaluating Chinese Ambiguity Understanding in Large Language Models
by: Mo, Junwen, et al.
Published: (2026)
by: Mo, Junwen, et al.
Published: (2026)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
by: Wu, Chengyue, et al.
Published: (2024)
by: Wu, Chengyue, et al.
Published: (2024)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
by: Wang, Chenlong, et al.
Published: (2026)
by: Wang, Chenlong, et al.
Published: (2026)
Vacuum initial data with minimal decay and borderline decay
by: Shen, Dawei, et al.
Published: (2026)
by: Shen, Dawei, et al.
Published: (2026)
Formation of multiple Black Holes from Cauchy data
by: Shen, Dawei, et al.
Published: (2025)
by: Shen, Dawei, et al.
Published: (2025)
Cauchy Data for Formation of Multiple Black Holes with Prescribed ADM Parameters
by: Shen, Dawei, et al.
Published: (2026)
by: Shen, Dawei, et al.
Published: (2026)
Formation of trapped surfaces for the Einstein--Maxwell--charged scalar field system
by: Shen, Dawei, et al.
Published: (2025)
by: Shen, Dawei, et al.
Published: (2025)
Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
by: Lu, Yanzuo, et al.
Published: (2025)
by: Lu, Yanzuo, et al.
Published: (2025)
Similar Items
-
MammothModa: Multi-Modal Large Language Model
by: She, Qi, et al.
Published: (2024) -
TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding
by: Pan, Junwen, et al.
Published: (2025) -
Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
by: Shi, Yangming, et al.
Published: (2026) -
TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
by: Pan, Junwen, et al.
Published: (2025) -
ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
by: Zhang, Yuan, et al.
Published: (2025)