LLM-I: LLMs are Naturally Interleaved Multimodal Creators
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Zirun, Zhang, Feng, Jia, Kai, Jin, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
Efficient Prompting for Continual Adaptation to Missing Modalities
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
Thinking with Programming Vision: Towards a Unified View for Thinking with Images
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
by: Hong, Minjie, et al.
Published: (2025)
by: Hong, Minjie, et al.
Published: (2025)
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
by: Li, Feng, et al.
Published: (2024)
by: Li, Feng, et al.
Published: (2024)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
by: Jin, Jiachun, et al.
Published: (2026)
by: Jin, Jiachun, et al.
Published: (2026)
Multimodal LLMs under Pairwise Modalities
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Farm-LightSeek: An Edge-centric Multimodal Agricultural IoT Data Analytics Framework with Lightweight LLMs
by: Jiang, Dawen, et al.
Published: (2025)
by: Jiang, Dawen, et al.
Published: (2025)
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
by: Huo, Simin, et al.
Published: (2025)
by: Huo, Simin, et al.
Published: (2025)
Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
by: Xu, Haochuan, et al.
Published: (2025)
by: Xu, Haochuan, et al.
Published: (2025)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
by: Cheng, Jintao, et al.
Published: (2025)
by: Cheng, Jintao, et al.
Published: (2025)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
CrossLLM-Mamba: Multimodal State Space Fusion of LLMs for RNA Interaction Prediction
by: Sadia, Rabeya Tus, et al.
Published: (2026)
by: Sadia, Rabeya Tus, et al.
Published: (2026)
Using Interleaved Ensemble Unlearning to Keep Backdoors at Bay for Finetuning Vision Transformers
by: Li, Zeyu Michael
Published: (2024)
by: Li, Zeyu Michael
Published: (2024)
Interleaved-Modal Chain-of-Thought
by: Gao, Jun, et al.
Published: (2024)
by: Gao, Jun, et al.
Published: (2024)
Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptation
by: He, Yongbo, et al.
Published: (2026)
by: He, Yongbo, et al.
Published: (2026)
PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
by: Li, Loka, et al.
Published: (2025)
by: Li, Loka, et al.
Published: (2025)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
by: Ye, Wenqian, et al.
Published: (2024)
by: Ye, Wenqian, et al.
Published: (2024)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
by: Daxberger, Erik, et al.
Published: (2025)
by: Daxberger, Erik, et al.
Published: (2025)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects
by: Zhou, Zirun, et al.
Published: (2025)
by: Zhou, Zirun, et al.
Published: (2025)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity
by: Tang, Jian, et al.
Published: (2026)
by: Tang, Jian, et al.
Published: (2026)
Interleaving Reasoning for Better Text-to-Image Generation
by: Huang, Wenxuan, et al.
Published: (2025)
by: Huang, Wenxuan, et al.
Published: (2025)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
by: Neuhaus, Yannic, et al.
Published: (2026)
by: Neuhaus, Yannic, et al.
Published: (2026)
Contrastive Learning for Multimodal Human Activity Recognition with Limited Labeled Data
by: Jing, Long, et al.
Published: (2026)
by: Jing, Long, et al.
Published: (2026)
CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning
by: Li, Wenjie, et al.
Published: (2025)
by: Li, Wenjie, et al.
Published: (2025)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
by: Chi, Xiaowei, et al.
Published: (2023)
by: Chi, Xiaowei, et al.
Published: (2023)
Toward Unified Multimodal Representation Learning for Autonomous Driving
by: Tao, Ximeng, et al.
Published: (2026)
by: Tao, Ximeng, et al.
Published: (2026)
DreamLLM: Synergistic Multimodal Comprehension and Creation
by: Dong, Runpei, et al.
Published: (2023)
by: Dong, Runpei, et al.
Published: (2023)
CrunchLLM: Multitask LLMs for Structured Business Reasoning and Outcome Prediction
by: Sadia, Rabeya Tus, et al.
Published: (2025)
by: Sadia, Rabeya Tus, et al.
Published: (2025)
Test-Time Multimodal Backdoor Detection by Contrastive Prompting
by: Niu, Yuwei, et al.
Published: (2024)
by: Niu, Yuwei, et al.
Published: (2024)
Similar Items
-
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
by: Guo, Zirun, et al.
Published: (2025) -
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
by: Guo, Zirun, et al.
Published: (2025) -
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
by: Guo, Zirun, et al.
Published: (2024) -
Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
by: Guo, Zirun, et al.
Published: (2024) -
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
by: Guo, Zirun, et al.
Published: (2025)