Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhiqian, Zhao, Xu, Xu, Xiaoqing, Liang, Guangdong, Wang, Weijia, Lv, Xiaolei, Li, Bo, Gao, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
von: Diesendruck, Maurice, et al.
Veröffentlicht: (2024)
von: Diesendruck, Maurice, et al.
Veröffentlicht: (2024)
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
von: Chen, Xiuwei, et al.
Veröffentlicht: (2025)
von: Chen, Xiuwei, et al.
Veröffentlicht: (2025)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
von: Xu, Run, et al.
Veröffentlicht: (2026)
von: Xu, Run, et al.
Veröffentlicht: (2026)
Intern-S1: A Scientific Multimodal Foundation Model
von: Bai, Lei, et al.
Veröffentlicht: (2025)
von: Bai, Lei, et al.
Veröffentlicht: (2025)
ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation
von: Wu, Mengyang, et al.
Veröffentlicht: (2024)
von: Wu, Mengyang, et al.
Veröffentlicht: (2024)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
von: Yang, Ruichao, et al.
Veröffentlicht: (2026)
von: Yang, Ruichao, et al.
Veröffentlicht: (2026)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
von: Gao, Sensen, et al.
Veröffentlicht: (2025)
von: Gao, Sensen, et al.
Veröffentlicht: (2025)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
Toward Robust Multimodal Learning using Multimodal Foundational Models
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
UEval: A Benchmark for Unified Multimodal Generation
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments
von: Li, Dinging, et al.
Veröffentlicht: (2026)
von: Li, Dinging, et al.
Veröffentlicht: (2026)
TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
von: Yang, Zheyuan, et al.
Veröffentlicht: (2026)
von: Yang, Zheyuan, et al.
Veröffentlicht: (2026)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
von: LASA Team, et al.
Veröffentlicht: (2025)
von: LASA Team, et al.
Veröffentlicht: (2025)
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
von: Zou, Yicheng, et al.
Veröffentlicht: (2026)
von: Zou, Yicheng, et al.
Veröffentlicht: (2026)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
von: Wu, Wei, et al.
Veröffentlicht: (2026)
von: Wu, Wei, et al.
Veröffentlicht: (2026)
DiSA: Diffusion Step Annealing in Autoregressive Image Generation
von: Zhao, Qinyu, et al.
Veröffentlicht: (2025)
von: Zhao, Qinyu, et al.
Veröffentlicht: (2025)
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
von: Hu, Yushi, et al.
Veröffentlicht: (2024)
von: Hu, Yushi, et al.
Veröffentlicht: (2024)
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2023)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2023)
How Far Are We from Generating Missing Modalities with Foundation Models?
von: Ke, Guanzhou, et al.
Veröffentlicht: (2025)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2025)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Breaking through Deterministic Barriers: Randomized Pruning Mask Generation and Selection
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
von: Xu, Hongshen, et al.
Veröffentlicht: (2024)
von: Xu, Hongshen, et al.
Veröffentlicht: (2024)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
Labeling Comic Mischief Content in Online Videos with a Multimodal Hierarchical-Cross-Attention Model
von: Baharlouei, Elaheh, et al.
Veröffentlicht: (2024)
von: Baharlouei, Elaheh, et al.
Veröffentlicht: (2024)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
von: Shangguan, Ziyao, et al.
Veröffentlicht: (2024)
von: Shangguan, Ziyao, et al.
Veröffentlicht: (2024)
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
von: Li, Kaican, et al.
Veröffentlicht: (2025)
von: Li, Kaican, et al.
Veröffentlicht: (2025)
Diving into Self-Evolving Training for Multimodal Reasoning
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
C${^2}$RL: Content and Context Representation Learning for Gloss-free Sign Language Translation and Retrieval
von: Chen, Zhigang, et al.
Veröffentlicht: (2024)
von: Chen, Zhigang, et al.
Veröffentlicht: (2024)
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification
von: Liu, Rui, et al.
Veröffentlicht: (2026)
von: Liu, Rui, et al.
Veröffentlicht: (2026)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection
von: Zhang, Shuguang, et al.
Veröffentlicht: (2026)
von: Zhang, Shuguang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
von: Diesendruck, Maurice, et al.
Veröffentlicht: (2024) -
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
von: Wen, Zichen, et al.
Veröffentlicht: (2025) -
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
von: Chen, Xiuwei, et al.
Veröffentlicht: (2025) -
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
von: Xu, Run, et al.
Veröffentlicht: (2026) -
Intern-S1: A Scientific Multimodal Foundation Model
von: Bai, Lei, et al.
Veröffentlicht: (2025)