InstructEngine: Instruction-driven Text-to-Image Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Xingyu, Hu, Yuhang, Zhang, YiFan, Jiang, Kaiyu, Liu, Changyi, Zhang, Tianke, Wang, Jinpeng, Yuan, Chun, Wen, Bin, Yang, Fan, Gao, Tingting, Zhang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ContextRL: Enhancing MLLM's Knowledge Discovery Efficiency with Context-Augmented RL
by: Lu, Xingyu, et al.
Published: (2026)
by: Lu, Xingyu, et al.
Published: (2026)
VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform
by: Lu, Xingyu, et al.
Published: (2025)
by: Lu, Xingyu, et al.
Published: (2025)
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
by: Hu, Xiao, et al.
Published: (2025)
by: Hu, Xiao, et al.
Published: (2025)
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
by: Lu, Xingyu, et al.
Published: (2026)
by: Lu, Xingyu, et al.
Published: (2026)
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
by: Chen, Jiankang, et al.
Published: (2025)
by: Chen, Jiankang, et al.
Published: (2025)
Thyme: Think Beyond Images
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
Kwai-STaR: Transform LLMs into State-Transition Reasoners
by: Lu, Xingyu, et al.
Published: (2024)
by: Lu, Xingyu, et al.
Published: (2024)
SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning
by: Long, Yancheng, et al.
Published: (2026)
by: Long, Yancheng, et al.
Published: (2026)
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
by: Yang, Yankai, et al.
Published: (2026)
by: Yang, Yankai, et al.
Published: (2026)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
by: Liu, Wenqi, et al.
Published: (2026)
by: Liu, Wenqi, et al.
Published: (2026)
UniRef-Image-Edit: Towards Scalable and Consistent Multi-Reference Image Editing
by: Wei, Hongyang, et al.
Published: (2026)
by: Wei, Hongyang, et al.
Published: (2026)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
by: Yu, Haoran, et al.
Published: (2025)
by: Yu, Haoran, et al.
Published: (2025)
InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing
by: Zhang, Zhedong, et al.
Published: (2025)
by: Zhang, Zhedong, et al.
Published: (2025)
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
by: Huang, Kexin, et al.
Published: (2025)
by: Huang, Kexin, et al.
Published: (2025)
InstructBooth: Instruction-following Personalized Text-to-Image Generation
by: Chae, Daewon, et al.
Published: (2023)
by: Chae, Daewon, et al.
Published: (2023)
InstructBrush: Learning Attention-based Instruction Optimization for Image Editing
by: Zhao, Ruoyu, et al.
Published: (2024)
by: Zhao, Ruoyu, et al.
Published: (2024)
EVLM: An Efficient Vision-Language Model for Visual Understanding
by: Chen, Kaibing, et al.
Published: (2024)
by: Chen, Kaibing, et al.
Published: (2024)
InstructRestore: Region-Customized Image Restoration with Human Instructions
by: Liu, Shuaizheng, et al.
Published: (2025)
by: Liu, Shuaizheng, et al.
Published: (2025)
PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning
by: Wang, Yunxiao, et al.
Published: (2025)
by: Wang, Yunxiao, et al.
Published: (2025)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
by: Lu, Jinda, et al.
Published: (2025)
by: Lu, Jinda, et al.
Published: (2025)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
by: Yue, Xinli, et al.
Published: (2025)
by: Yue, Xinli, et al.
Published: (2025)
SciInstruct: a Self-Reflective Instruction Annotated Dataset for Training Scientific Language Models
by: Zhang, Dan, et al.
Published: (2024)
by: Zhang, Dan, et al.
Published: (2024)
TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs
by: Wang, Yunxiao, et al.
Published: (2025)
by: Wang, Yunxiao, et al.
Published: (2025)
InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
by: Gan, Yulu, et al.
Published: (2023)
by: Gan, Yulu, et al.
Published: (2023)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
by: Ren, Yong, et al.
Published: (2026)
by: Ren, Yong, et al.
Published: (2026)
InstructVEdit: A Holistic Approach for Instructional Video Editing
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct
by: Wu, Yutong, et al.
Published: (2024)
by: Wu, Yutong, et al.
Published: (2024)
InstructOCR: Instruction Boosting Scene Text Spotting
by: Duan, Chen, et al.
Published: (2024)
by: Duan, Chen, et al.
Published: (2024)
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
TimeBrush : An Intelligent Expert System for Restoring Historical Images With Temporal and Stylistic Guidance
by: Kaiyu Zhang
Published: (2025)
by: Kaiyu Zhang
Published: (2025)
Point and Instruct: Enabling Precise Image Editing by Unifying Direct Manipulation and Text Instructions
by: Helbling, Alec, et al.
Published: (2024)
by: Helbling, Alec, et al.
Published: (2024)
Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
Learning to Instruct for Visual Instruction Tuning
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
Debiasing Multimodal Large Language Models via Penalization of Language Priors
by: Zhang, YiFan, et al.
Published: (2024)
by: Zhang, YiFan, et al.
Published: (2024)
InstructSAM: Segment Any Instance with Any Instructions
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
LogoRA: Local-Global Representation Alignment for Robust Time Series Classification
by: Zhang, Huanyu, et al.
Published: (2024)
by: Zhang, Huanyu, et al.
Published: (2024)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
EmoLat: Text-driven Image Sentiment Transfer via Emotion Latent Space
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
Similar Items
-
ContextRL: Enhancing MLLM's Knowledge Discovery Efficiency with Context-Augmented RL
by: Lu, Xingyu, et al.
Published: (2026) -
VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform
by: Lu, Xingyu, et al.
Published: (2025) -
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
by: Hu, Xiao, et al.
Published: (2025) -
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
by: Lu, Xingyu, et al.
Published: (2026) -
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
by: Chen, Jiankang, et al.
Published: (2025)