UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhaowei, Wang, Wei, Cai, YiQing, Qi, Xu, Wang, Pengyu, Zhang, Dong, Song, Hang, Jiang, Botian, Huang, Zhida, Wang, Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
von: Wang, Wei, et al.
Veröffentlicht: (2024)
von: Wang, Wei, et al.
Veröffentlicht: (2024)
GroundingGPT:Language Enhanced Multi-modal Grounding Model
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
von: Li, Zhaowei, et al.
Veröffentlicht: (2024)
UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
von: Tan, Chenkun, et al.
Veröffentlicht: (2025)
von: Tan, Chenkun, et al.
Veröffentlicht: (2025)
M4SC: An MLLM-based Multi-modal, Multi-task and Multi-user Semantic Communication System
von: Jiang, Feibo, et al.
Veröffentlicht: (2025)
von: Jiang, Feibo, et al.
Veröffentlicht: (2025)
QCRD: Quality-guided Contrastive Rationale Distillation for Large Language Models
von: Wang, Wei, et al.
Veröffentlicht: (2024)
von: Wang, Wei, et al.
Veröffentlicht: (2024)
Liquid: Language Models are Scalable and Unified Multi-modal Generators
von: Wu, Junfeng, et al.
Veröffentlicht: (2024)
von: Wu, Junfeng, et al.
Veröffentlicht: (2024)
A Unified Representation Underlying the Judgment of Large Language Models
von: Lu, Yi-Long, et al.
Veröffentlicht: (2025)
von: Lu, Yi-Long, et al.
Veröffentlicht: (2025)
Attenuation of LHAASO PeVatrons by Interstellar Radiation Field and Cosmic Microwave Background Radiation
von: Zhang, Jianli, et al.
Veröffentlicht: (2024)
von: Zhang, Jianli, et al.
Veröffentlicht: (2024)
Multi-modal Relation Distillation for Unified 3D Representation Learning
von: Wang, Huiqun, et al.
Veröffentlicht: (2024)
von: Wang, Huiqun, et al.
Veröffentlicht: (2024)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)
Unified Scene Representation and Reconstruction for 3D Large Language Models
von: Chu, Tao, et al.
Veröffentlicht: (2024)
von: Chu, Tao, et al.
Veröffentlicht: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
Collaborative Multi-LoRA Experts with Achievement-based Multi-Tasks Loss for Unified Multimodal Information Extraction
von: Yuan, Li, et al.
Veröffentlicht: (2025)
von: Yuan, Li, et al.
Veröffentlicht: (2025)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
MMGen: Unified Multi-modal Image Generation and Understanding in One Go
von: Wang, Jiepeng, et al.
Veröffentlicht: (2025)
von: Wang, Jiepeng, et al.
Veröffentlicht: (2025)
SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
von: Huang, Junming, et al.
Veröffentlicht: (2026)
von: Huang, Junming, et al.
Veröffentlicht: (2026)
OmniColor: A Unified Framework for Multi-modal Lineart Colorization
von: Zhang, Xulu, et al.
Veröffentlicht: (2026)
von: Zhang, Xulu, et al.
Veröffentlicht: (2026)
Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework
von: Zhang, Yukun, et al.
Veröffentlicht: (2025)
von: Zhang, Yukun, et al.
Veröffentlicht: (2025)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
von: Wei, Cong, et al.
Veröffentlicht: (2024)
von: Wei, Cong, et al.
Veröffentlicht: (2024)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
von: Xu, Runsen, et al.
Veröffentlicht: (2025)
von: Xu, Runsen, et al.
Veröffentlicht: (2025)
SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
A Unified Graph Transformer for Overcoming Isolations in Multi-modal Recommendation
von: Yi, Zixuan, et al.
Veröffentlicht: (2024)
von: Yi, Zixuan, et al.
Veröffentlicht: (2024)
Empowering Segmentation Ability to Multi-modal Large Language Models
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
von: Qiu, Xuerui, et al.
Veröffentlicht: (2026)
von: Qiu, Xuerui, et al.
Veröffentlicht: (2026)
PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLM
von: Yang, Tao, et al.
Veröffentlicht: (2024)
von: Yang, Tao, et al.
Veröffentlicht: (2024)
UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark
von: Zhou, Zhaokun, et al.
Veröffentlicht: (2024)
von: Zhou, Zhaokun, et al.
Veröffentlicht: (2024)
$M^3EL$: A Multi-task Multi-topic Dataset for Multi-modal Entity Linking
von: Wang, Fang, et al.
Veröffentlicht: (2024)
von: Wang, Fang, et al.
Veröffentlicht: (2024)
Learning Unified Reference Representation for Unsupervised Multi-class Anomaly Detection
von: He, Liren, et al.
Veröffentlicht: (2024)
von: He, Liren, et al.
Veröffentlicht: (2024)
Bridge: A Unified Framework to Knowledge Graph Completion via Language Models and Knowledge Representation
von: Qiao, Qiao, et al.
Veröffentlicht: (2024)
von: Qiao, Qiao, et al.
Veröffentlicht: (2024)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)
von: Walimbe, Soham, et al.
Veröffentlicht: (2025)
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
von: Wang, Yukun, et al.
Veröffentlicht: (2026)
von: Wang, Yukun, et al.
Veröffentlicht: (2026)
DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
von: Chen, Yuntao, et al.
Veröffentlicht: (2024)
von: Chen, Yuntao, et al.
Veröffentlicht: (2024)
BrepCoder: A Unified Multimodal Large Language Model for Multi-task B-rep Reasoning
von: Kim, Mingi, et al.
Veröffentlicht: (2026)
von: Kim, Mingi, et al.
Veröffentlicht: (2026)
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
von: Huang, Runhui, et al.
Veröffentlicht: (2025)
von: Huang, Runhui, et al.
Veröffentlicht: (2025)
LiDARFormer: A Unified Transformer-based Multi-task Network for LiDAR Perception
von: Zhou, Zixiang, et al.
Veröffentlicht: (2023)
von: Zhou, Zixiang, et al.
Veröffentlicht: (2023)
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
von: Duan, Lunhao, et al.
Veröffentlicht: (2024)
von: Duan, Lunhao, et al.
Veröffentlicht: (2024)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
von: Wang, Wei, et al.
Veröffentlicht: (2024) -
GroundingGPT:Language Enhanced Multi-modal Grounding Model
von: Li, Zhaowei, et al.
Veröffentlicht: (2024) -
UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
von: Wang, Pengyu, et al.
Veröffentlicht: (2025) -
Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
von: Tan, Chenkun, et al.
Veröffentlicht: (2025) -
M4SC: An MLLM-based Multi-modal, Multi-task and Multi-user Semantic Communication System
von: Jiang, Feibo, et al.
Veröffentlicht: (2025)