TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
Fuente:
arXiv
Guardado en:
| Autores principales: | Wei, Xinyu, Zhang, Jinrui, Wang, Zeqing, Wei, Hongyang, Guo, Zhen, Zhang, Lei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
por: Wang, Zeqing, et al.
Publicado: (2025)
por: Wang, Zeqing, et al.
Publicado: (2025)
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
por: Wei, Xinyu, et al.
Publicado: (2025)
por: Wei, Xinyu, et al.
Publicado: (2025)
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
por: Zhou, Chao, et al.
Publicado: (2025)
por: Zhou, Chao, et al.
Publicado: (2025)
How to Design and Train Your Implicit Neural Representation for Video Compression
por: Gwilliam, Matthew, et al.
Publicado: (2025)
por: Gwilliam, Matthew, et al.
Publicado: (2025)
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
por: Feng, Kunyu, et al.
Publicado: (2025)
por: Feng, Kunyu, et al.
Publicado: (2025)
Follow-Your-Color: Multi-Instance Sketch Colorization
por: Zhang, Yinhan, et al.
Publicado: (2025)
por: Zhang, Yinhan, et al.
Publicado: (2025)
Let Your Video Listen to Your Music!
por: Zhang, Xinyu, et al.
Publicado: (2025)
por: Zhang, Xinyu, et al.
Publicado: (2025)
SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
por: Peng, Haosong, et al.
Publicado: (2026)
por: Peng, Haosong, et al.
Publicado: (2026)
Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer
por: Li, Yuze, et al.
Publicado: (2026)
por: Li, Yuze, et al.
Publicado: (2026)
PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
por: Wang, Zeqing, et al.
Publicado: (2025)
por: Wang, Zeqing, et al.
Publicado: (2025)
Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation
por: Chen, Qihua, et al.
Publicado: (2024)
por: Chen, Qihua, et al.
Publicado: (2024)
Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
por: Ma, Yue, et al.
Publicado: (2025)
por: Ma, Yue, et al.
Publicado: (2025)
Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation
por: Ma, Yue, et al.
Publicado: (2024)
por: Ma, Yue, et al.
Publicado: (2024)
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
por: Zhang, Jinrui, et al.
Publicado: (2024)
por: Zhang, Jinrui, et al.
Publicado: (2024)
InstructionBench: An Instructional Video Understanding Benchmark
por: Wei, Haiwan, et al.
Publicado: (2025)
por: Wei, Haiwan, et al.
Publicado: (2025)
Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts
por: Ma, Yue, et al.
Publicado: (2024)
por: Ma, Yue, et al.
Publicado: (2024)
Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
por: Ma, Yue, et al.
Publicado: (2025)
por: Ma, Yue, et al.
Publicado: (2025)
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
por: Wang, Fengxiang, et al.
Publicado: (2025)
por: Wang, Fengxiang, et al.
Publicado: (2025)
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
por: Li, Yifei, et al.
Publicado: (2025)
por: Li, Yifei, et al.
Publicado: (2025)
Follow-Your-Preference: Towards Preference-Aligned Image Inpainting
por: Shen, Yutao, et al.
Publicado: (2025)
por: Shen, Yutao, et al.
Publicado: (2025)
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
por: Zhang, Huanyu, et al.
Publicado: (2026)
por: Zhang, Huanyu, et al.
Publicado: (2026)
Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
por: Long, Zeqian, et al.
Publicado: (2025)
por: Long, Zeqian, et al.
Publicado: (2025)
Attention! Your Vision Language Model Could Be Maliciously Manipulated
por: Wang, Xiaosen, et al.
Publicado: (2025)
por: Wang, Xiaosen, et al.
Publicado: (2025)
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
por: Wei, Hongyang, et al.
Publicado: (2025)
por: Wei, Hongyang, et al.
Publicado: (2025)
Puppeteer: Rig and Animate Your 3D Models
por: Song, Chaoyue, et al.
Publicado: (2025)
por: Song, Chaoyue, et al.
Publicado: (2025)
Video, How Do Your Tokens Merge?
por: Pollard, Sam, et al.
Publicado: (2025)
por: Pollard, Sam, et al.
Publicado: (2025)
How Video Meetings Change Your Expression
por: Sarin, Sumit, et al.
Publicado: (2024)
por: Sarin, Sumit, et al.
Publicado: (2024)
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
por: Wang, Chunwei, et al.
Publicado: (2024)
por: Wang, Chunwei, et al.
Publicado: (2024)
How I Met Your Bias: Investigating Bias Amplification in Diffusion Models
por: Roos, Nathan, et al.
Publicado: (2025)
por: Roos, Nathan, et al.
Publicado: (2025)
Follow Your Motion: A Generic Temporal Consistency Portrait Editing Framework with Trajectory Guidance
por: Yang, Haijie, et al.
Publicado: (2025)
por: Yang, Haijie, et al.
Publicado: (2025)
Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
por: Qu, Tianyuan, et al.
Publicado: (2025)
por: Qu, Tianyuan, et al.
Publicado: (2025)
How to Squeeze An Explanation Out of Your Model
por: Roxo, Tiago, et al.
Publicado: (2024)
por: Roxo, Tiago, et al.
Publicado: (2024)
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
por: Wu, Junfeng, et al.
Publicado: (2025)
por: Wu, Junfeng, et al.
Publicado: (2025)
PushupBench: Your VLM is not good at counting pushups
por: Li, Shengzhi, et al.
Publicado: (2026)
por: Li, Shengzhi, et al.
Publicado: (2026)
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
por: Yi, Hongwei, et al.
Publicado: (2025)
por: Yi, Hongwei, et al.
Publicado: (2025)
Watch Your Steps: Local Image and Scene Editing by Text Instructions
por: Mirzaei, Ashkan, et al.
Publicado: (2023)
por: Mirzaei, Ashkan, et al.
Publicado: (2023)
Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection
por: Lundqvist, Lars, et al.
Publicado: (2026)
por: Lundqvist, Lars, et al.
Publicado: (2026)
Bringing Your Portrait to 3D Presence
por: Zhang, Jiawei, et al.
Publicado: (2025)
por: Zhang, Jiawei, et al.
Publicado: (2025)
How Animals Dance (When You're Not Looking)
por: Wang, Xiaojuan, et al.
Publicado: (2025)
por: Wang, Xiaojuan, et al.
Publicado: (2025)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
por: Ma, Yue, et al.
Publicado: (2025)
por: Ma, Yue, et al.
Publicado: (2025)
Ejemplares similares
-
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
por: Wang, Zeqing, et al.
Publicado: (2025) -
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
por: Wei, Xinyu, et al.
Publicado: (2025) -
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
por: Zhou, Chao, et al.
Publicado: (2025) -
How to Design and Train Your Implicit Neural Representation for Video Compression
por: Gwilliam, Matthew, et al.
Publicado: (2025) -
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
por: Feng, Kunyu, et al.
Publicado: (2025)