TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Xinyu, Zhang, Jinrui, Wang, Zeqing, Wei, Hongyang, Guo, Zhen, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
by: Wei, Xinyu, et al.
Published: (2025)
by: Wei, Xinyu, et al.
Published: (2025)
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
by: Zhou, Chao, et al.
Published: (2025)
by: Zhou, Chao, et al.
Published: (2025)
How to Design and Train Your Implicit Neural Representation for Video Compression
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
by: Feng, Kunyu, et al.
Published: (2025)
by: Feng, Kunyu, et al.
Published: (2025)
Follow-Your-Color: Multi-Instance Sketch Colorization
by: Zhang, Yinhan, et al.
Published: (2025)
by: Zhang, Yinhan, et al.
Published: (2025)
Let Your Video Listen to Your Music!
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
by: Peng, Haosong, et al.
Published: (2026)
by: Peng, Haosong, et al.
Published: (2026)
Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer
by: Li, Yuze, et al.
Published: (2026)
by: Li, Yuze, et al.
Published: (2026)
PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation
by: Chen, Qihua, et al.
Published: (2024)
by: Chen, Qihua, et al.
Published: (2024)
Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation
by: Ma, Yue, et al.
Published: (2024)
by: Ma, Yue, et al.
Published: (2024)
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
by: Zhang, Jinrui, et al.
Published: (2024)
by: Zhang, Jinrui, et al.
Published: (2024)
InstructionBench: An Instructional Video Understanding Benchmark
by: Wei, Haiwan, et al.
Published: (2025)
by: Wei, Haiwan, et al.
Published: (2025)
Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts
by: Ma, Yue, et al.
Published: (2024)
by: Ma, Yue, et al.
Published: (2024)
Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
Follow-Your-Preference: Towards Preference-Aligned Image Inpainting
by: Shen, Yutao, et al.
Published: (2025)
by: Shen, Yutao, et al.
Published: (2025)
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
by: Zhang, Huanyu, et al.
Published: (2026)
by: Zhang, Huanyu, et al.
Published: (2026)
Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
by: Long, Zeqian, et al.
Published: (2025)
by: Long, Zeqian, et al.
Published: (2025)
Attention! Your Vision Language Model Could Be Maliciously Manipulated
by: Wang, Xiaosen, et al.
Published: (2025)
by: Wang, Xiaosen, et al.
Published: (2025)
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
by: Wei, Hongyang, et al.
Published: (2025)
by: Wei, Hongyang, et al.
Published: (2025)
Puppeteer: Rig and Animate Your 3D Models
by: Song, Chaoyue, et al.
Published: (2025)
by: Song, Chaoyue, et al.
Published: (2025)
Video, How Do Your Tokens Merge?
by: Pollard, Sam, et al.
Published: (2025)
by: Pollard, Sam, et al.
Published: (2025)
How Video Meetings Change Your Expression
by: Sarin, Sumit, et al.
Published: (2024)
by: Sarin, Sumit, et al.
Published: (2024)
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
by: Wang, Chunwei, et al.
Published: (2024)
by: Wang, Chunwei, et al.
Published: (2024)
How I Met Your Bias: Investigating Bias Amplification in Diffusion Models
by: Roos, Nathan, et al.
Published: (2025)
by: Roos, Nathan, et al.
Published: (2025)
Follow Your Motion: A Generic Temporal Consistency Portrait Editing Framework with Trajectory Guidance
by: Yang, Haijie, et al.
Published: (2025)
by: Yang, Haijie, et al.
Published: (2025)
Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
by: Qu, Tianyuan, et al.
Published: (2025)
by: Qu, Tianyuan, et al.
Published: (2025)
How to Squeeze An Explanation Out of Your Model
by: Roxo, Tiago, et al.
Published: (2024)
by: Roxo, Tiago, et al.
Published: (2024)
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
by: Wu, Junfeng, et al.
Published: (2025)
by: Wu, Junfeng, et al.
Published: (2025)
PushupBench: Your VLM is not good at counting pushups
by: Li, Shengzhi, et al.
Published: (2026)
by: Li, Shengzhi, et al.
Published: (2026)
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
by: Yi, Hongwei, et al.
Published: (2025)
by: Yi, Hongwei, et al.
Published: (2025)
Watch Your Steps: Local Image and Scene Editing by Text Instructions
by: Mirzaei, Ashkan, et al.
Published: (2023)
by: Mirzaei, Ashkan, et al.
Published: (2023)
Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection
by: Lundqvist, Lars, et al.
Published: (2026)
by: Lundqvist, Lars, et al.
Published: (2026)
Bringing Your Portrait to 3D Presence
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
How Animals Dance (When You're Not Looking)
by: Wang, Xiaojuan, et al.
Published: (2025)
by: Wang, Xiaojuan, et al.
Published: (2025)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
Similar Items
-
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025) -
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
by: Wei, Xinyu, et al.
Published: (2025) -
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
by: Zhou, Chao, et al.
Published: (2025) -
How to Design and Train Your Implicit Neural Representation for Video Compression
by: Gwilliam, Matthew, et al.
Published: (2025) -
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
by: Feng, Kunyu, et al.
Published: (2025)