ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chern, Ethan, Su, Jiadi, Ma, Yan, Liu, Pengfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Thinking with Generated Images
von: Chern, Ethan, et al.
Veröffentlicht: (2025)
von: Chern, Ethan, et al.
Veröffentlicht: (2025)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
Holistic Evaluation for Interleaved Text-and-Image Generation
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
von: Chern, Ethan, et al.
Veröffentlicht: (2025)
von: Chern, Ethan, et al.
Veröffentlicht: (2025)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
Interleaving Reasoning for Better Text-to-Image Generation
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
von: Zhang, Leixin, et al.
Veröffentlicht: (2024)
von: Zhang, Leixin, et al.
Veröffentlicht: (2024)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
von: Hu, Yushi, et al.
Veröffentlicht: (2025)
von: Hu, Yushi, et al.
Veröffentlicht: (2025)
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
von: Chen, Chao, et al.
Veröffentlicht: (2025)
von: Chen, Chao, et al.
Veröffentlicht: (2025)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
von: Ma, Yiyang, et al.
Veröffentlicht: (2024)
von: Ma, Yiyang, et al.
Veröffentlicht: (2024)
MANTIS: Interleaved Multi-Image Instruction Tuning
von: Jiang, Dongfu, et al.
Veröffentlicht: (2024)
von: Jiang, Dongfu, et al.
Veröffentlicht: (2024)
MULTI: Multimodal Understanding Leaderboard with Text and Images
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
von: Zhao, Haozhe, et al.
Veröffentlicht: (2025)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2025)
MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text
von: Zhang, Junzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2025)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
von: Li, Xiujun, et al.
Veröffentlicht: (2023)
von: Li, Xiujun, et al.
Veröffentlicht: (2023)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
von: Li, Zhang, et al.
Veröffentlicht: (2023)
von: Li, Zhang, et al.
Veröffentlicht: (2023)
IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web
von: Guo, Hongcheng, et al.
Veröffentlicht: (2024)
von: Guo, Hongcheng, et al.
Veröffentlicht: (2024)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
von: Gu, Zeqi, et al.
Veröffentlicht: (2025)
von: Gu, Zeqi, et al.
Veröffentlicht: (2025)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
von: Tang, Bohao, et al.
Veröffentlicht: (2025)
von: Tang, Bohao, et al.
Veröffentlicht: (2025)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
von: Ma, Yan, et al.
Veröffentlicht: (2025)
von: Ma, Yan, et al.
Veröffentlicht: (2025)
Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
von: Buckley, Thomas, et al.
Veröffentlicht: (2023)
von: Buckley, Thomas, et al.
Veröffentlicht: (2023)
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
von: Hong, Jixiang, et al.
Veröffentlicht: (2025)
von: Hong, Jixiang, et al.
Veröffentlicht: (2025)
Towards Text-Image Interleaved Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
von: Wang, Junjie, et al.
Veröffentlicht: (2024)
von: Wang, Junjie, et al.
Veröffentlicht: (2024)
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
von: Yamabe, Shojiro, et al.
Veröffentlicht: (2025)
von: Yamabe, Shojiro, et al.
Veröffentlicht: (2025)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
von: Yu, Hao, et al.
Veröffentlicht: (2025)
von: Yu, Hao, et al.
Veröffentlicht: (2025)
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
von: Xue, Le, et al.
Veröffentlicht: (2024)
von: Xue, Le, et al.
Veröffentlicht: (2024)
Model Composition for Multimodal Large Language Models
von: Chen, Chi, et al.
Veröffentlicht: (2024)
von: Chen, Chi, et al.
Veröffentlicht: (2024)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
von: S, Sridhar, et al.
Veröffentlicht: (2025)
von: S, Sridhar, et al.
Veröffentlicht: (2025)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
von: Wu, Yi, et al.
Veröffentlicht: (2025)
von: Wu, Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Thinking with Generated Images
von: Chern, Ethan, et al.
Veröffentlicht: (2025) -
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
von: Kou, Siqi, et al.
Veröffentlicht: (2024) -
Holistic Evaluation for Interleaved Text-and-Image Generation
von: Liu, Minqian, et al.
Veröffentlicht: (2024) -
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
von: Chern, Ethan, et al.
Veröffentlicht: (2025) -
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)