LLM-I: LLMs are Naturally Interleaved Multimodal Creators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Zirun, Zhang, Feng, Jia, Kai, Jin, Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
Efficient Prompting for Continual Adaptation to Missing Modalities
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
Thinking with Programming Vision: Towards a Unified View for Thinking with Images
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
von: Hong, Minjie, et al.
Veröffentlicht: (2025)
von: Hong, Minjie, et al.
Veröffentlicht: (2025)
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
von: Li, Feng, et al.
Veröffentlicht: (2024)
von: Li, Feng, et al.
Veröffentlicht: (2024)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
von: Jin, Jiachun, et al.
Veröffentlicht: (2026)
von: Jin, Jiachun, et al.
Veröffentlicht: (2026)
Multimodal LLMs under Pairwise Modalities
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
Farm-LightSeek: An Edge-centric Multimodal Agricultural IoT Data Analytics Framework with Lightweight LLMs
von: Jiang, Dawen, et al.
Veröffentlicht: (2025)
von: Jiang, Dawen, et al.
Veröffentlicht: (2025)
Skipping Computations in Multimodal LLMs
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
von: Huo, Simin, et al.
Veröffentlicht: (2025)
von: Huo, Simin, et al.
Veröffentlicht: (2025)
Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
von: Xu, Haochuan, et al.
Veröffentlicht: (2025)
von: Xu, Haochuan, et al.
Veröffentlicht: (2025)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
von: Cheng, Jintao, et al.
Veröffentlicht: (2025)
von: Cheng, Jintao, et al.
Veröffentlicht: (2025)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
CrossLLM-Mamba: Multimodal State Space Fusion of LLMs for RNA Interaction Prediction
von: Sadia, Rabeya Tus, et al.
Veröffentlicht: (2026)
von: Sadia, Rabeya Tus, et al.
Veröffentlicht: (2026)
Using Interleaved Ensemble Unlearning to Keep Backdoors at Bay for Finetuning Vision Transformers
von: Li, Zeyu Michael
Veröffentlicht: (2024)
von: Li, Zeyu Michael
Veröffentlicht: (2024)
Interleaved-Modal Chain-of-Thought
von: Gao, Jun, et al.
Veröffentlicht: (2024)
von: Gao, Jun, et al.
Veröffentlicht: (2024)
Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptation
von: He, Yongbo, et al.
Veröffentlicht: (2026)
von: He, Yongbo, et al.
Veröffentlicht: (2026)
PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
von: Li, Loka, et al.
Veröffentlicht: (2025)
von: Li, Loka, et al.
Veröffentlicht: (2025)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
von: Ye, Wenqian, et al.
Veröffentlicht: (2024)
von: Ye, Wenqian, et al.
Veröffentlicht: (2024)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
von: Daxberger, Erik, et al.
Veröffentlicht: (2025)
von: Daxberger, Erik, et al.
Veröffentlicht: (2025)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects
von: Zhou, Zirun, et al.
Veröffentlicht: (2025)
von: Zhou, Zirun, et al.
Veröffentlicht: (2025)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity
von: Tang, Jian, et al.
Veröffentlicht: (2026)
von: Tang, Jian, et al.
Veröffentlicht: (2026)
Interleaving Reasoning for Better Text-to-Image Generation
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
von: Neuhaus, Yannic, et al.
Veröffentlicht: (2026)
von: Neuhaus, Yannic, et al.
Veröffentlicht: (2026)
Contrastive Learning for Multimodal Human Activity Recognition with Limited Labeled Data
von: Jing, Long, et al.
Veröffentlicht: (2026)
von: Jing, Long, et al.
Veröffentlicht: (2026)
CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning
von: Li, Wenjie, et al.
Veröffentlicht: (2025)
von: Li, Wenjie, et al.
Veröffentlicht: (2025)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023)
Toward Unified Multimodal Representation Learning for Autonomous Driving
von: Tao, Ximeng, et al.
Veröffentlicht: (2026)
von: Tao, Ximeng, et al.
Veröffentlicht: (2026)
DreamLLM: Synergistic Multimodal Comprehension and Creation
von: Dong, Runpei, et al.
Veröffentlicht: (2023)
von: Dong, Runpei, et al.
Veröffentlicht: (2023)
CrunchLLM: Multitask LLMs for Structured Business Reasoning and Outcome Prediction
von: Sadia, Rabeya Tus, et al.
Veröffentlicht: (2025)
von: Sadia, Rabeya Tus, et al.
Veröffentlicht: (2025)
Test-Time Multimodal Backdoor Detection by Contrastive Prompting
von: Niu, Yuwei, et al.
Veröffentlicht: (2024)
von: Niu, Yuwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
von: Guo, Zirun, et al.
Veröffentlicht: (2025) -
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
von: Guo, Zirun, et al.
Veröffentlicht: (2025) -
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
von: Guo, Zirun, et al.
Veröffentlicht: (2024) -
Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
von: Guo, Zirun, et al.
Veröffentlicht: (2024) -
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
von: Guo, Zirun, et al.
Veröffentlicht: (2025)