KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Hsin-Ping, Wang, Xinyi, Bitton, Yonatan, Taitelbaum, Hagai, Tomar, Gaurav Singh, Chang, Ming-Wei, Jia, Xuhui, Chan, Kelvin C. K., Hu, Hexiang, Su, Yu-Chuan, Yang, Ming-Hsuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
by: Slobodkin, Aviv, et al.
Published: (2025)
by: Slobodkin, Aviv, et al.
Published: (2025)
Improving Subject-Driven Image Synthesis with Subject-Agnostic Guidance
by: Chan, Kelvin C. K., et al.
Published: (2024)
by: Chan, Kelvin C. K., et al.
Published: (2024)
CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts
by: Hsin-Ying, Lee, et al.
Published: (2025)
by: Hsin-Ying, Lee, et al.
Published: (2025)
Instruct-Imagen: Image Generation with Multi-modal Instruction
by: Hu, Hexiang, et al.
Published: (2024)
by: Hu, Hexiang, et al.
Published: (2024)
Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis
by: Ventura, Mor, et al.
Published: (2026)
by: Ventura, Mor, et al.
Published: (2026)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024)
by: Yanuka, Moran, et al.
Published: (2024)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
Investigating Content Planning for Navigating Trade-offs in Knowledge-Grounded Dialogue
by: Chawla, Kushal, et al.
Published: (2024)
by: Chawla, Kushal, et al.
Published: (2024)
On Reference (In-)Determinacy in Natural Language Inference
by: Chen, Sihao, et al.
Published: (2025)
by: Chen, Sihao, et al.
Published: (2025)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
by: Gordon, Brian, et al.
Published: (2023)
by: Gordon, Brian, et al.
Published: (2023)
Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
by: Hu, Qingguo, et al.
Published: (2025)
by: Hu, Qingguo, et al.
Published: (2025)
Centre mode instability of a dilute particle-laden swirling jet in a swirl flow combustor
by: Warrier, Srikumar, et al.
Published: (2024)
by: Warrier, Srikumar, et al.
Published: (2024)
Weakly nonlinear analysis of particle-laden Rayleigh-Bénard convection
by: Srinivas, Thota, et al.
Published: (2025)
by: Srinivas, Thota, et al.
Published: (2025)
ParallelPARC: A Scalable Pipeline for Generating Natural-Language Analogies
by: Sultan, Oren, et al.
Published: (2024)
by: Sultan, Oren, et al.
Published: (2024)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
by: Yosef, Ron, et al.
Published: (2025)
by: Yosef, Ron, et al.
Published: (2025)
From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute Transition
by: Lo, Ling, et al.
Published: (2025)
by: Lo, Ling, et al.
Published: (2025)
Distributed Indexing Schemes for k-Dominant Skyline Analytics on Uncertain Edge-IoT Data
by: Lai, Chuan-Chi, et al.
Published: (2023)
by: Lai, Chuan-Chi, et al.
Published: (2023)
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
On the deformation of a shear thinning viscoelastic drop in a steady electric field
by: Bangar, Sarika Shivaji, et al.
Published: (2026)
by: Bangar, Sarika Shivaji, et al.
Published: (2026)
On Large Deformations of Oldroyd-B Drops in a Steady Electric Field
by: Bangar, Sarika Shivaji, et al.
Published: (2026)
by: Bangar, Sarika Shivaji, et al.
Published: (2026)
Hegedus' Conjecture and Tighter Upper Bounds for Equidistant Codes in Hamming Spaces
by: Hu, Sihuang, et al.
Published: (2025)
by: Hu, Sihuang, et al.
Published: (2025)
Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
by: Ma, Nanye, et al.
Published: (2025)
by: Ma, Nanye, et al.
Published: (2025)
MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
by: Hsin-Ying, Lee, et al.
Published: (2026)
by: Hsin-Ying, Lee, et al.
Published: (2026)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Exploiting Diffusion Prior for Generalizable Dense Prediction
by: Lee, Hsin-Ying, et al.
Published: (2023)
by: Lee, Hsin-Ying, et al.
Published: (2023)
Weakly Supervised 3D Object Detection via Multi-Level Visual Guidance
by: Huang, Kuan-Chih, et al.
Published: (2023)
by: Huang, Kuan-Chih, et al.
Published: (2023)
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
by: Ramos, Vasco, et al.
Published: (2024)
by: Ramos, Vasco, et al.
Published: (2024)
Robust Guidance for Unsupervised Data Selection: Capturing Perplexing Named Entities for Domain-Specific Machine Translation
by: Ji, Seunghyun, et al.
Published: (2024)
by: Ji, Seunghyun, et al.
Published: (2024)
Holographic covering and the fortuity of black holes
by: Chang, Chi-Ming, et al.
Published: (2024)
by: Chang, Chi-Ming, et al.
Published: (2024)
Violation of S-duality in classical $Q$-cohomology
by: Chang, Chi-Ming, et al.
Published: (2025)
by: Chang, Chi-Ming, et al.
Published: (2025)
Entity-Aware Biaffine Attention Model for Improved Constituent Parsing with Reduced Entity Violations
by: Bai, Xinyi
Published: (2024)
by: Bai, Xinyi
Published: (2024)
Human-Centric Community Detection in Hybrid Metaverse Networks with Integrated AI Entities
by: Chiu, Shih-Hsuan, et al.
Published: (2025)
by: Chiu, Shih-Hsuan, et al.
Published: (2025)
WimPyC: an extension module of WimPyDD for the calculation of WIMP capture in celestial bodies
by: Kang, Sunghyun, et al.
Published: (2025)
by: Kang, Sunghyun, et al.
Published: (2025)
Low-mass constraints on WIMP effective models of inelastic scattering using the Migdal effect
by: Kang, Sunghyun, et al.
Published: (2024)
by: Kang, Sunghyun, et al.
Published: (2024)
Reemergence of Trampolining in a Leidenfrost Droplet
by: Agrawal, Pranjal, et al.
Published: (2024)
by: Agrawal, Pranjal, et al.
Published: (2024)
Probing Dark Matter Electromagnetic Properties in Direct Detection Experiments
by: Ibarra, Alejandro, et al.
Published: (2024)
by: Ibarra, Alejandro, et al.
Published: (2024)
Inviscid stability of compressible flows past compliant surfaces
by: Deka, Mandeep, et al.
Published: (2024)
by: Deka, Mandeep, et al.
Published: (2024)
Dual Associated Encoder for Face Restoration
by: Tsai, Yu-Ju, et al.
Published: (2023)
by: Tsai, Yu-Ju, et al.
Published: (2023)
Similar Items
-
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
by: Slobodkin, Aviv, et al.
Published: (2025) -
Improving Subject-Driven Image Synthesis with Subject-Agnostic Guidance
by: Chan, Kelvin C. K., et al.
Published: (2024) -
CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts
by: Hsin-Ying, Lee, et al.
Published: (2025) -
Instruct-Imagen: Image Generation with Multi-modal Instruction
by: Hu, Hexiang, et al.
Published: (2024) -
Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis
by: Ventura, Mor, et al.
Published: (2026)