WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Niu, Yuwei, Ning, Munan, Zheng, Mengren, Jin, Weiyang, Lin, Bin, Jin, Peng, Liao, Jiaqi, Feng, Chaoran, Ning, Kunpeng, Zhu, Bin, Yuan, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
by: Tourani, Ali, et al.
Published: (2023)
by: Tourani, Ali, et al.
Published: (2023)
Chat-Driven Text Generation and Interaction for Person Retrieval
by: Xie, Zequn, et al.
Published: (2025)
by: Xie, Zequn, et al.
Published: (2025)
Towards Localizing Structural Elements: Merging Geometrical Detection with Semantic Verification in RGB-D Data
by: Tourani, Ali, et al.
Published: (2024)
by: Tourani, Ali, et al.
Published: (2024)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
by: Feng, Yichen, et al.
Published: (2026)
by: Feng, Yichen, et al.
Published: (2026)
Training-Free Diffusion Priors for Text-to-Image Generation via Optimization-based Visual Inversion
by: Dell'Erba, Samuele, et al.
Published: (2025)
by: Dell'Erba, Samuele, et al.
Published: (2025)
Ego-Motion Aware Target Prediction Module for Robust Multi-Object Tracking
by: Mahdian, Navid, et al.
Published: (2024)
by: Mahdian, Navid, et al.
Published: (2024)
Positive Style Accumulation: A Style Screening and Continuous Utilization Framework for Federated DG-ReID
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
by: Radwan, Ahmed, et al.
Published: (2024)
by: Radwan, Ahmed, et al.
Published: (2024)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
by: Skripkin, Matvey, et al.
Published: (2025)
by: Skripkin, Matvey, et al.
Published: (2025)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
by: Shahin, Nada, et al.
Published: (2025)
by: Shahin, Nada, et al.
Published: (2025)
Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
by: Liu, Shanyuan, et al.
Published: (2023)
by: Liu, Shanyuan, et al.
Published: (2023)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
by: Dua, Karan, et al.
Published: (2025)
by: Dua, Karan, et al.
Published: (2025)
Facial Attribute Based Text Guided Face Anonymization
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
by: Seo, Huichan, et al.
Published: (2025)
by: Seo, Huichan, et al.
Published: (2025)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
by: Shahin, Nada, et al.
Published: (2025)
by: Shahin, Nada, et al.
Published: (2025)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
by: Li, Danyang, et al.
Published: (2025)
by: Li, Danyang, et al.
Published: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
by: Lim, Shoon Kit, et al.
Published: (2025)
by: Lim, Shoon Kit, et al.
Published: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
NumeriKontrol: Adding Numeric Control to Diffusion Transformers for Instruction-based Image Editing
by: Xu, Zhenyu, et al.
Published: (2025)
by: Xu, Zhenyu, et al.
Published: (2025)
Accelerating Post-Tornado Disaster Assessment Using Advanced Deep Learning Models
by: Umeike, Robinson, et al.
Published: (2024)
by: Umeike, Robinson, et al.
Published: (2024)
Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images
by: Käs, Stephanie, et al.
Published: (2025)
by: Käs, Stephanie, et al.
Published: (2025)
Human-Centric Anomaly Detection in Surveillance Videos Using YOLO-World and Spatio-Temporal Deep Learning
by: Naeen, Mohammad Ali Etemadi, et al.
Published: (2025)
by: Naeen, Mohammad Ali Etemadi, et al.
Published: (2025)
Palmistry-Informed Feature Extraction and Analysis using Machine Learning
by: Patil, Shweta
Published: (2025)
by: Patil, Shweta
Published: (2025)
Transformers for Image-Goal Navigation
by: Pelluri, Nikhilanj
Published: (2024)
by: Pelluri, Nikhilanj
Published: (2024)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
by: Bartkowiak, Patryk, et al.
Published: (2026)
by: Bartkowiak, Patryk, et al.
Published: (2026)
Advanced Long-term Earth System Forecasting
by: Wu, Hao, et al.
Published: (2025)
by: Wu, Hao, et al.
Published: (2025)
Dynamic Residual Encoding with Slide-Level Contrastive Learning for End-to-End Whole Slide Image Representation
by: Jin, Jing, et al.
Published: (2025)
by: Jin, Jing, et al.
Published: (2025)
A Light Perspective for 3D Object Detection
by: Pederiva, Marcelo Eduardo, et al.
Published: (2025)
by: Pederiva, Marcelo Eduardo, et al.
Published: (2025)
Combining Absolute and Semi-Generalized Relative Poses for Visual Localization
by: Panek, Vojtech, et al.
Published: (2024)
by: Panek, Vojtech, et al.
Published: (2024)
A Guide to Structureless Visual Localization
by: Panek, Vojtech, et al.
Published: (2025)
by: Panek, Vojtech, et al.
Published: (2025)
Reference Dataset and Benchmark for Reconstructing Laser Parameters from On-axis Video in Powder Bed Fusion of Bulk Stainless Steel
by: Blanc, Cyril, et al.
Published: (2024)
by: Blanc, Cyril, et al.
Published: (2024)
Privacy-Preserving Structureless Visual Localization via Image Obfuscation
by: Panek, Vojtech, et al.
Published: (2026)
by: Panek, Vojtech, et al.
Published: (2026)
Cross-Domain Adversarial Augmentation: Stabilizing GANs for Medical and Handwriting Data Scarcity
by: Soad, Md. Sohanuzzaman, et al.
Published: (2026)
by: Soad, Md. Sohanuzzaman, et al.
Published: (2026)
Learning Diffeomorphism for Image Registration with Time-Continuous Networks using Semigroup Regularization
by: Matinkia, Mohammadjavad, et al.
Published: (2024)
by: Matinkia, Mohammadjavad, et al.
Published: (2024)
EvoIQA - Explaining Image Distortions with Evolved White-Box Logic
by: Gupta, Ruchika, et al.
Published: (2026)
by: Gupta, Ruchika, et al.
Published: (2026)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
HATL: Hierarchical Adaptive-Transfer Learning Framework for Sign Language Machine Translation
by: Shahin, Nada, et al.
Published: (2026)
by: Shahin, Nada, et al.
Published: (2026)
SpectraNet: FFT-assisted Deep Learning Classifier for Deepfake Face Detection
by: Jayarathne, Nithira, et al.
Published: (2025)
by: Jayarathne, Nithira, et al.
Published: (2025)
Similar Items
-
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
by: Tourani, Ali, et al.
Published: (2023) -
Chat-Driven Text Generation and Interaction for Person Retrieval
by: Xie, Zequn, et al.
Published: (2025) -
Towards Localizing Structural Elements: Merging Geometrical Detection with Semantic Verification in RGB-D Data
by: Tourani, Ali, et al.
Published: (2024) -
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024) -
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
by: Feng, Yichen, et al.
Published: (2026)