JPEG-LM: LLMs as Image Generators with Canonical Codec Representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Xiaochuang, Ghazvininejad, Marjan, Koh, Pang Wei, Tsvetkov, Yulia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation
di: Chen, Junhao, et al.
Pubblicazione: (2025)
di: Chen, Junhao, et al.
Pubblicazione: (2025)
David helps Goliath: Inference-Time Collaboration Between Small Specialized and Large General Diffusion LMs
di: Han, Xiaochuang, et al.
Pubblicazione: (2023)
di: Han, Xiaochuang, et al.
Pubblicazione: (2023)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
di: Li, Siting, et al.
Pubblicazione: (2024)
di: Li, Siting, et al.
Pubblicazione: (2024)
TV2TV: A Unified Framework for Interleaved Language and Video Generation
di: Han, Xiaochuang, et al.
Pubblicazione: (2025)
di: Han, Xiaochuang, et al.
Pubblicazione: (2025)
Dissecting Adversarial Robustness of Multimodal LM Agents
di: Wu, Chen Henry, et al.
Pubblicazione: (2024)
di: Wu, Chen Henry, et al.
Pubblicazione: (2024)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
di: Shi, Weijia, et al.
Pubblicazione: (2024)
di: Shi, Weijia, et al.
Pubblicazione: (2024)
ICONS: Influence Consensus for Vision-Language Data Selection
di: Wu, Xindi, et al.
Pubblicazione: (2024)
di: Wu, Xindi, et al.
Pubblicazione: (2024)
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
di: Hu, Yushi, et al.
Pubblicazione: (2025)
di: Hu, Yushi, et al.
Pubblicazione: (2025)
MatFormer: Nested Transformer for Elastic Inference
di: Devvrit, et al.
Pubblicazione: (2023)
di: Devvrit, et al.
Pubblicazione: (2023)
Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models
di: Yang, Xu, et al.
Pubblicazione: (2023)
di: Yang, Xu, et al.
Pubblicazione: (2023)
LLMs as Visual Explainers: Advancing Image Classification with Evolving Visual Descriptions
di: Han, Songhao, et al.
Pubblicazione: (2023)
di: Han, Songhao, et al.
Pubblicazione: (2023)
MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization
di: Chaubey, Ashutosh, et al.
Pubblicazione: (2026)
di: Chaubey, Ashutosh, et al.
Pubblicazione: (2026)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
di: Yang, Cheng-Fu, et al.
Pubblicazione: (2024)
di: Yang, Cheng-Fu, et al.
Pubblicazione: (2024)
MOFI: Learning Image Representations from Noisy Entity Annotated Images
di: Wu, Wentao, et al.
Pubblicazione: (2023)
di: Wu, Wentao, et al.
Pubblicazione: (2023)
AiGen-FoodReview: A Multimodal Dataset of Machine-Generated Restaurant Reviews and Images on Social Media
di: Gambetti, Alessandro, et al.
Pubblicazione: (2024)
di: Gambetti, Alessandro, et al.
Pubblicazione: (2024)
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
di: Cai, Zhenyang, et al.
Pubblicazione: (2024)
di: Cai, Zhenyang, et al.
Pubblicazione: (2024)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
di: Hsieh, Cheng-Yu, et al.
Pubblicazione: (2025)
di: Hsieh, Cheng-Yu, et al.
Pubblicazione: (2025)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
di: Shukor, Mustafa, et al.
Pubblicazione: (2024)
di: Shukor, Mustafa, et al.
Pubblicazione: (2024)
Fake or JPEG? Revealing Common Biases in Generated Image Detection Datasets
di: Grommelt, Patrick, et al.
Pubblicazione: (2024)
di: Grommelt, Patrick, et al.
Pubblicazione: (2024)
Dual-Process Image Generation
di: Luo, Grace, et al.
Pubblicazione: (2025)
di: Luo, Grace, et al.
Pubblicazione: (2025)
General Transform: A Unified Framework for Adaptive Transform to Enhance Representations
di: Budiutama, Gekko, et al.
Pubblicazione: (2025)
di: Budiutama, Gekko, et al.
Pubblicazione: (2025)
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
di: Wang, Andrew Z., et al.
Pubblicazione: (2025)
di: Wang, Andrew Z., et al.
Pubblicazione: (2025)
Multilingual Diversity Improves Vision-Language Representations
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
di: Lei, Jiayi, et al.
Pubblicazione: (2025)
di: Lei, Jiayi, et al.
Pubblicazione: (2025)
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
di: Zhu, Mengdan, et al.
Pubblicazione: (2024)
di: Zhu, Mengdan, et al.
Pubblicazione: (2024)
Unified Text-Image Generation with Weakness-Targeted Post-Training
di: Chen, Jiahui, et al.
Pubblicazione: (2026)
di: Chen, Jiahui, et al.
Pubblicazione: (2026)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
di: Yu, Eric Yang, et al.
Pubblicazione: (2024)
di: Yu, Eric Yang, et al.
Pubblicazione: (2024)
Text-to-Image Cross-Modal Generation: A Systematic Review
di: Żelaszczyk, Maciej, et al.
Pubblicazione: (2024)
di: Żelaszczyk, Maciej, et al.
Pubblicazione: (2024)
Canonical Latent Representations in Conditional Diffusion Models
di: Xu, Yitao, et al.
Pubblicazione: (2025)
di: Xu, Yitao, et al.
Pubblicazione: (2025)
Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
di: Han, Tianyang, et al.
Pubblicazione: (2024)
di: Han, Tianyang, et al.
Pubblicazione: (2024)
Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image Generation
di: Collins, Katherine M., et al.
Pubblicazione: (2024)
di: Collins, Katherine M., et al.
Pubblicazione: (2024)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
di: Meng, Debin, et al.
Pubblicazione: (2025)
di: Meng, Debin, et al.
Pubblicazione: (2025)
Hyperbolic Multimodal Representation Learning for Biological Taxonomies
di: Gong, ZeMing, et al.
Pubblicazione: (2025)
di: Gong, ZeMing, et al.
Pubblicazione: (2025)
Mixture of Group Experts for Learning Invariant Representations
di: Kang, Lei, et al.
Pubblicazione: (2025)
di: Kang, Lei, et al.
Pubblicazione: (2025)
RADAR: Relative Angular Divergence Across Representations
di: Cadet, Xavier, et al.
Pubblicazione: (2026)
di: Cadet, Xavier, et al.
Pubblicazione: (2026)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
di: Ghazanfari, Sara, et al.
Pubblicazione: (2024)
di: Ghazanfari, Sara, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation
di: Chen, Junhao, et al.
Pubblicazione: (2025) -
David helps Goliath: Inference-Time Collaboration Between Small Specialized and Large General Diffusion LMs
di: Han, Xiaochuang, et al.
Pubblicazione: (2023) -
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
di: Li, Siting, et al.
Pubblicazione: (2024) -
TV2TV: A Unified Framework for Interleaved Language and Video Generation
di: Han, Xiaochuang, et al.
Pubblicazione: (2025) -
Dissecting Adversarial Robustness of Multimodal LM Agents
di: Wu, Chen Henry, et al.
Pubblicazione: (2024)