Instruct-Imagen: Image Generation with Multi-modal Instruction
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Hexiang, Chan, Kelvin C. K., Su, Yu-Chuan, Chen, Wenhu, Li, Yandong, Sohn, Kihyuk, Zhao, Yang, Ben, Xue, Gong, Boqing, Cohen, William, Chang, Ming-Wei, Jia, Xuhui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions
por: Zhang, Kai, et al.
Publicado: (2024)
por: Zhang, Kai, et al.
Publicado: (2024)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities
por: Huang, Hsin-Ping, et al.
Publicado: (2024)
por: Huang, Hsin-Ping, et al.
Publicado: (2024)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
por: Jia, Yiming, et al.
Publicado: (2025)
por: Jia, Yiming, et al.
Publicado: (2025)
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion Models
por: Lee, Kyungmin, et al.
Publicado: (2024)
por: Lee, Kyungmin, et al.
Publicado: (2024)
Improving Subject-Driven Image Synthesis with Subject-Agnostic Guidance
por: Chan, Kelvin C. K., et al.
Publicado: (2024)
por: Chan, Kelvin C. K., et al.
Publicado: (2024)
Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
por: Ma, Nanye, et al.
Publicado: (2025)
por: Ma, Nanye, et al.
Publicado: (2025)
DreamFlow: High-Quality Text-to-3D Generation by Approximating Probability Flow
por: Lee, Kyungmin, et al.
Publicado: (2024)
por: Lee, Kyungmin, et al.
Publicado: (2024)
Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images!
por: Zou, Zihang, et al.
Publicado: (2026)
por: Zou, Zihang, et al.
Publicado: (2026)
Epsilon-VAE: Denoising as Visual Decoding
por: Zhao, Long, et al.
Publicado: (2024)
por: Zhao, Long, et al.
Publicado: (2024)
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
por: Zhang, Kai, et al.
Publicado: (2023)
por: Zhang, Kai, et al.
Publicado: (2023)
Culture in Action: Evaluating Text-to-Image Models through Social Activities
por: Malakouti, Sina, et al.
Publicado: (2025)
por: Malakouti, Sina, et al.
Publicado: (2025)
What's in a Name? Beyond Class Indices for Image Recognition
por: Han, Kai, et al.
Publicado: (2023)
por: Han, Kai, et al.
Publicado: (2023)
InstructEngine: Instruction-driven Text-to-Image Alignment
por: Lu, Xingyu, et al.
Publicado: (2025)
por: Lu, Xingyu, et al.
Publicado: (2025)
Where is the answer? Investigating Positional Bias in Language Model Knowledge Extraction
por: Saito, Kuniaki, et al.
Publicado: (2024)
por: Saito, Kuniaki, et al.
Publicado: (2024)
Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
por: Tan, Yuwen, et al.
Publicado: (2025)
por: Tan, Yuwen, et al.
Publicado: (2025)
InstructEdit: Instruction-based Knowledge Editing for Large Language Models
por: Zhang, Ningyu, et al.
Publicado: (2024)
por: Zhang, Ningyu, et al.
Publicado: (2024)
Cross-Model Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Across Three Large Language Models
por: Lee, Kihyuk
Publicado: (2026)
por: Lee, Kihyuk
Publicado: (2026)
InstructRestore: Region-Customized Image Restoration with Human Instructions
por: Liu, Shuaizheng, et al.
Publicado: (2025)
por: Liu, Shuaizheng, et al.
Publicado: (2025)
Learning to Instruct for Visual Instruction Tuning
por: Zhou, Zhihan, et al.
Publicado: (2025)
por: Zhou, Zhihan, et al.
Publicado: (2025)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
por: Sakai, Yuki, et al.
Publicado: (2025)
por: Sakai, Yuki, et al.
Publicado: (2025)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
por: Wei, Cong, et al.
Publicado: (2024)
por: Wei, Cong, et al.
Publicado: (2024)
HypDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-shot Image Generation
por: Li, Lingxiao, et al.
Publicado: (2024)
por: Li, Lingxiao, et al.
Publicado: (2024)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
por: Tan, Yuwen, et al.
Publicado: (2025)
por: Tan, Yuwen, et al.
Publicado: (2025)
InstructBooth: Instruction-following Personalized Text-to-Image Generation
por: Chae, Daewon, et al.
Publicado: (2023)
por: Chae, Daewon, et al.
Publicado: (2023)
PhotoFramer: Multi-modal Image Composition Instruction
por: You, Zhiyuan, et al.
Publicado: (2025)
por: You, Zhiyuan, et al.
Publicado: (2025)
Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoring
por: Xie, Qizhi, et al.
Publicado: (2025)
por: Xie, Qizhi, et al.
Publicado: (2025)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
por: Liu, Jihao, et al.
Publicado: (2024)
por: Liu, Jihao, et al.
Publicado: (2024)
MANTIS: Interleaved Multi-Image Instruction Tuning
por: Jiang, Dongfu, et al.
Publicado: (2024)
por: Jiang, Dongfu, et al.
Publicado: (2024)
ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in Instructions
por: He, Xingwei, et al.
Publicado: (2025)
por: He, Xingwei, et al.
Publicado: (2025)
InstructBrush: Learning Attention-based Instruction Optimization for Image Editing
por: Zhao, Ruoyu, et al.
Publicado: (2024)
por: Zhao, Ruoyu, et al.
Publicado: (2024)
InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
por: Gan, Yulu, et al.
Publicado: (2023)
por: Gan, Yulu, et al.
Publicado: (2023)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
por: Yu, Haoran, et al.
Publicado: (2025)
por: Yu, Haoran, et al.
Publicado: (2025)
ChartInstruct: Instruction Tuning for Chart Comprehension and Reasoning
por: Masry, Ahmed, et al.
Publicado: (2024)
por: Masry, Ahmed, et al.
Publicado: (2024)
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
por: Wan, Zifu, et al.
Publicado: (2025)
por: Wan, Zifu, et al.
Publicado: (2025)
InstructDET: Diversifying Referring Object Detection with Generalized Instructions
por: Dang, Ronghao, et al.
Publicado: (2023)
por: Dang, Ronghao, et al.
Publicado: (2023)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
por: Zhang, Ming, et al.
Publicado: (2024)
por: Zhang, Ming, et al.
Publicado: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
por: Ren, Huimin, et al.
Publicado: (2025)
por: Ren, Huimin, et al.
Publicado: (2025)
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models
por: Ou, Yixin, et al.
Publicado: (2024)
por: Ou, Yixin, et al.
Publicado: (2024)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
por: Wu, Keming, et al.
Publicado: (2025)
por: Wu, Keming, et al.
Publicado: (2025)
Ejemplares similares
-
MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions
por: Zhang, Kai, et al.
Publicado: (2024) -
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
por: Chen, Lichang, et al.
Publicado: (2024) -
KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities
por: Huang, Hsin-Ping, et al.
Publicado: (2024) -
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
por: Jia, Yiming, et al.
Publicado: (2025) -
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion Models
por: Lee, Kyungmin, et al.
Publicado: (2024)