FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Tong, Xu, Yinghao, Po, Ryan, Zhang, Mengchen, Yang, Guandao, Wang, Jiaqi, Liu, Ziwei, Lin, Dahua, Wetzstein, Gordon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video World Models with Long-term Spatial Memory
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Orthogonal Adaptation for Modular Customization of Diffusion Models
by: Po, Ryan, et al.
Published: (2023)
by: Po, Ryan, et al.
Published: (2023)
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
by: Zhang, Yuhan, et al.
Published: (2025)
by: Zhang, Yuhan, et al.
Published: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
by: Fang, Ye, et al.
Published: (2025)
by: Fang, Ye, et al.
Published: (2025)
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
by: Fang, Ye, et al.
Published: (2024)
by: Fang, Ye, et al.
Published: (2024)
Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
by: Ackermann, Jan, et al.
Published: (2025)
by: Ackermann, Jan, et al.
Published: (2025)
Omni6D: Large-Vocabulary 3D Object Dataset for Category-Level 6D Object Pose Estimation
by: Zhang, Mengchen, et al.
Published: (2024)
by: Zhang, Mengchen, et al.
Published: (2024)
BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
by: Po, Ryan, et al.
Published: (2025)
by: Po, Ryan, et al.
Published: (2025)
ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization
by: Guo, Yuanhe, et al.
Published: (2025)
by: Guo, Yuanhe, et al.
Published: (2025)
IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations
by: Li, Zhibing, et al.
Published: (2024)
by: Li, Zhibing, et al.
Published: (2024)
SS4D: Native 4D Generative Model via Structured Spacetime Latents
by: Li, Zhibing, et al.
Published: (2025)
by: Li, Zhibing, et al.
Published: (2025)
ChartLens: Fine-grained Visual Attribution in Charts
by: Suri, Manan, et al.
Published: (2025)
by: Suri, Manan, et al.
Published: (2025)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
by: He, Hao, et al.
Published: (2024)
by: He, Hao, et al.
Published: (2024)
MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines
by: Po, Ryan, et al.
Published: (2026)
by: Po, Ryan, et al.
Published: (2026)
Visual Analytics for Fine-grained Text Classification Models and Datasets
by: Battogtokh, Munkhtulga, et al.
Published: (2024)
by: Battogtokh, Munkhtulga, et al.
Published: (2024)
AttriStory: Fine-grained Attribute Realization for Visual Storytelling with Diffusion Models
by: Sreenivas, Manogna, et al.
Published: (2026)
by: Sreenivas, Manogna, et al.
Published: (2026)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
by: Kuang, Zhengfei, et al.
Published: (2024)
by: Kuang, Zhengfei, et al.
Published: (2024)
Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images
by: Qi, Zhangyang, et al.
Published: (2024)
by: Qi, Zhangyang, et al.
Published: (2024)
Spectral Progressive Diffusion for Efficient Image and Video Generation
by: Xiao, Howard, et al.
Published: (2026)
by: Xiao, Howard, et al.
Published: (2026)
Fine-grained Text to Image Synthesis
by: Ouyang, Xu, et al.
Published: (2024)
by: Ouyang, Xu, et al.
Published: (2024)
Foveated Diffusion: Efficient Spatially Adaptive Image and Video Generation
by: Chao, Brian, et al.
Published: (2026)
by: Chao, Brian, et al.
Published: (2026)
Neural Control Variates with Automatic Integration
by: Li, Zilu, et al.
Published: (2024)
by: Li, Zilu, et al.
Published: (2024)
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction
by: Deng, Youming, et al.
Published: (2025)
by: Deng, Youming, et al.
Published: (2025)
Detecting Dataset Abuse in Fine-Tuning Stable Diffusion Models for Text-to-Image Synthesis
by: Wang, Songrui, et al.
Published: (2024)
by: Wang, Songrui, et al.
Published: (2024)
Visual-RFT: Visual Reinforcement Fine-Tuning
by: Liu, Ziyu, et al.
Published: (2025)
by: Liu, Ziyu, et al.
Published: (2025)
Dual Ascent Diffusion for Inverse Problems
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
Robust Symmetry Detection via Riemannian Langevin Dynamics
by: Je, Jihyeon, et al.
Published: (2024)
by: Je, Jihyeon, et al.
Published: (2024)
Diffusion Self-Distillation for Zero-Shot Customized Image Generation
by: Cai, Shengqu, et al.
Published: (2024)
by: Cai, Shengqu, et al.
Published: (2024)
Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution
by: Wang, Ying, et al.
Published: (2023)
by: Wang, Ying, et al.
Published: (2023)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
by: Gu, Yuzhe, et al.
Published: (2025)
by: Gu, Yuzhe, et al.
Published: (2025)
Flow as the Cross-Domain Manipulation Interface
by: Xu, Mengda, et al.
Published: (2024)
by: Xu, Mengda, et al.
Published: (2024)
MegaScenes: Scene-Level View Synthesis at Scale
by: Tung, Joseph, et al.
Published: (2024)
by: Tung, Joseph, et al.
Published: (2024)
Detecting Origin Attribution for Text-to-Image Diffusion Models
by: Xu, Katherine, et al.
Published: (2024)
by: Xu, Katherine, et al.
Published: (2024)
Visual Agentic Reinforcement Fine-Tuning
by: Liu, Ziyu, et al.
Published: (2025)
by: Liu, Ziyu, et al.
Published: (2025)
VSC: Visual Search Compositional Text-to-Image Diffusion Model
by: Dat, Do Huu, et al.
Published: (2025)
by: Dat, Do Huu, et al.
Published: (2025)
DiLightNet: Fine-grained Lighting Control for Diffusion-based Image Generation
by: Zeng, Chong, et al.
Published: (2024)
by: Zeng, Chong, et al.
Published: (2024)
Similar Items
-
Video World Models with Long-term Spatial Memory
by: Wu, Tong, et al.
Published: (2025) -
Orthogonal Adaptation for Modular Customization of Diffusion Models
by: Po, Ryan, et al.
Published: (2023) -
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
by: Zhang, Mengchen, et al.
Published: (2025) -
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
by: Wu, Tong, et al.
Published: (2024) -
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
by: Zhang, Yuhan, et al.
Published: (2025)