PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Weifeng, Wei, Xinyu, Zhang, Renrui, Zhuo, Le, Zhao, Shitian, Huang, Siyuan, Teng, Huan, Xie, Junlin, Qiao, Yu, Gao, Peng, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
by: Liu, Dongyang, et al.
Published: (2024)
by: Liu, Dongyang, et al.
Published: (2024)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
by: Lei, Jiayi, et al.
Published: (2025)
by: Lei, Jiayi, et al.
Published: (2025)
Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models
by: Zhao, Shitian, et al.
Published: (2024)
by: Zhao, Shitian, et al.
Published: (2024)
ReasonPix2Pix: Instruction Reasoning Dataset for Advanced Image Editing
by: Jin, Ying, et al.
Published: (2024)
by: Jin, Ying, et al.
Published: (2024)
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
by: Zhuo, Le, et al.
Published: (2025)
by: Zhuo, Le, et al.
Published: (2025)
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
by: Lin, Weifeng, et al.
Published: (2025)
by: Lin, Weifeng, et al.
Published: (2025)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
by: Huang, Victor Shea-Jay, et al.
Published: (2025)
by: Huang, Victor Shea-Jay, et al.
Published: (2025)
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
Enhanced Pix2Pix GAN for Visual Defect Removal in UAV-Captured Images
by: Rizun, Volodymyr
Published: (2024)
by: Rizun, Volodymyr
Published: (2024)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
by: Liu, Dongyang, et al.
Published: (2024)
by: Liu, Dongyang, et al.
Published: (2024)
Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks
by: Kim, Jonghun, et al.
Published: (2026)
by: Kim, Jonghun, et al.
Published: (2026)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
by: Chen, Xinyan, et al.
Published: (2025)
by: Chen, Xinyan, et al.
Published: (2025)
Image-To-Image Translation: A Comprehensive Study on the Efficacy of Pix2Pix GAN in Producing High-Quality Visual Transformations
by: Dhakaa Mohsin Kareem
Published: (2025)
by: Dhakaa Mohsin Kareem
Published: (2025)
MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
by: Zhang, Renrui, et al.
Published: (2024)
by: Zhang, Renrui, et al.
Published: (2024)
Mapping New Realities: Ground Truth Image Creation with Pix2Pix Image-to-Image Translation
by: Li, Zhenglin, et al.
Published: (2024)
by: Li, Zhenglin, et al.
Published: (2024)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
by: Gao, Peng, et al.
Published: (2021)
by: Gao, Peng, et al.
Published: (2021)
MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions
by: Zhang, Kai, et al.
Published: (2024)
by: Zhang, Kai, et al.
Published: (2024)
Large Continual Instruction Assistant
by: Qiao, Jingyang, et al.
Published: (2024)
by: Qiao, Jingyang, et al.
Published: (2024)
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
by: Tong, Chengzhuo, et al.
Published: (2025)
by: Tong, Chengzhuo, et al.
Published: (2025)
Ambient-Pix2PixGAN for Translating Medical Images from Noisy Data
by: Chen, Wentao, et al.
Published: (2024)
by: Chen, Wentao, et al.
Published: (2024)
On LLM Wizards: Identifying Large Language Models' Behaviors for Wizard of Oz Experiments
by: Fang, Jingchao, et al.
Published: (2024)
by: Fang, Jingchao, et al.
Published: (2024)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
by: Zhao, Liangbing, et al.
Published: (2026)
by: Zhao, Liangbing, et al.
Published: (2026)
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
by: Zhang, Yabo, et al.
Published: (2026)
by: Zhang, Yabo, et al.
Published: (2026)
Novel Hybrid Integrated Pix2Pix and WGAN Model with Gradient Penalty for Binary Images Denoising
by: Tirel, Luca, et al.
Published: (2024)
by: Tirel, Luca, et al.
Published: (2024)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following
by: Li, Shufan, et al.
Published: (2023)
by: Li, Shufan, et al.
Published: (2023)
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
by: Zhang, Renrui, et al.
Published: (2023)
by: Zhang, Renrui, et al.
Published: (2023)
LLMs as Visual Explainers: Advancing Image Classification with Evolving Visual Descriptions
by: Han, Songhao, et al.
Published: (2023)
by: Han, Songhao, et al.
Published: (2023)
TalkPhoto: A Versatile Training-Free Conversational Assistant for Intelligent Image Editing
by: Hu, Yujie, et al.
Published: (2026)
by: Hu, Yujie, et al.
Published: (2026)
Empowering Visual Creativity: A Vision-Language Assistant to Image Editing Recommendations
by: Shen, Tiancheng, et al.
Published: (2024)
by: Shen, Tiancheng, et al.
Published: (2024)
Factuality Matters: When Image Generation and Editing Meet Structured Visuals
by: Zhuo, Le, et al.
Published: (2025)
by: Zhuo, Le, et al.
Published: (2025)
Pix2Key: Controllable Open-Vocabulary Retrieval with Semantic Decomposition and Self-Supervised Visual Dictionary Learning
by: Wei, Guoyizhe, et al.
Published: (2026)
by: Wei, Guoyizhe, et al.
Published: (2026)
VS-Assistant: Versatile Surgery Assistant on the Demand of Surgeons
by: Chen, Zhen, et al.
Published: (2024)
by: Chen, Zhen, et al.
Published: (2024)
Large Language Models in Cardiovascular Imaging: Current Applications and Future Prospects
by: Weifeng Yuan, et al.
Published: (2025)
by: Weifeng Yuan, et al.
Published: (2025)
Rethinking VLM Representation for VLA Initialization
by: Lin, Weifeng, et al.
Published: (2026)
by: Lin, Weifeng, et al.
Published: (2026)
SRU-Pix2Pix: A Fusion-Driven Generator Network for Medical Image Translation with Few-Shot Learning
by: Qiu, Xihe, et al.
Published: (2026)
by: Qiu, Xihe, et al.
Published: (2026)
AstroPix for the Barrel Imaging Calorimeter in ePIC experiment
by: Kim, Bobae
Published: (2024)
by: Kim, Bobae
Published: (2024)
Similar Items
-
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
by: Liu, Dongyang, et al.
Published: (2024) -
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
by: Lei, Jiayi, et al.
Published: (2025) -
Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models
by: Zhao, Shitian, et al.
Published: (2024) -
ReasonPix2Pix: Instruction Reasoning Dataset for Advanced Image Editing
by: Jin, Ying, et al.
Published: (2024) -
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
by: Zhuo, Le, et al.
Published: (2025)