VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Cao, Jingtao, Zhang, Zheng, Wang, Hongru, Wong, Kam-Fai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025)
di: Dai, Song, et al.
Pubblicazione: (2025)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
di: Le, Van-Truong
Pubblicazione: (2026)
di: Le, Van-Truong
Pubblicazione: (2026)
GroundCap: A Visually Grounded Image Captioning Dataset
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
di: Ji, Yikun, et al.
Pubblicazione: (2025)
di: Ji, Yikun, et al.
Pubblicazione: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
di: Portelance, Eva, et al.
Pubblicazione: (2023)
di: Portelance, Eva, et al.
Pubblicazione: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
di: Yang, Shan
Pubblicazione: (2026)
di: Yang, Shan
Pubblicazione: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
di: Bucher, Martin JJ., et al.
Pubblicazione: (2025)
di: Bucher, Martin JJ., et al.
Pubblicazione: (2025)
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
di: Niu, Yuwei, et al.
Pubblicazione: (2025)
di: Niu, Yuwei, et al.
Pubblicazione: (2025)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
di: Chen, Yuangong, et al.
Pubblicazione: (2026)
di: Chen, Yuangong, et al.
Pubblicazione: (2026)
What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation
di: Yang, Dingyi, et al.
Pubblicazione: (2024)
di: Yang, Dingyi, et al.
Pubblicazione: (2024)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
di: Freitas, Diogo, et al.
Pubblicazione: (2025)
di: Freitas, Diogo, et al.
Pubblicazione: (2025)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
di: He, Wei
Pubblicazione: (2026)
di: He, Wei
Pubblicazione: (2026)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
Evaluating Perspectival Biases in Cross-Modal Retrieval
di: Saengsukhiran, Teerapol, et al.
Pubblicazione: (2025)
di: Saengsukhiran, Teerapol, et al.
Pubblicazione: (2025)
Cora: Correspondence-aware image editing using few step diffusion
di: Alimohammadi, Amirhossein, et al.
Pubblicazione: (2025)
di: Alimohammadi, Amirhossein, et al.
Pubblicazione: (2025)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2025)
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2025)
Defending against Backdoor Attacks via Module Switching
di: Li, Weijun, et al.
Pubblicazione: (2025)
di: Li, Weijun, et al.
Pubblicazione: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2025)
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2025)
3DGEER: 3D Gaussian Rendering Made Exact and Efficient for Generic Cameras
di: Huang, Zixun, et al.
Pubblicazione: (2025)
di: Huang, Zixun, et al.
Pubblicazione: (2025)
Appearance-Invariant Detection of Suggestive Motion via Laban Movement Descriptors on SMPL Skeletons
di: Ahn, Jaehoon, et al.
Pubblicazione: (2026)
di: Ahn, Jaehoon, et al.
Pubblicazione: (2026)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
di: Cui, Shaoyang, et al.
Pubblicazione: (2026)
di: Cui, Shaoyang, et al.
Pubblicazione: (2026)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
di: Anh, Duy Le Dinh, et al.
Pubblicazione: (2024)
di: Anh, Duy Le Dinh, et al.
Pubblicazione: (2024)
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
di: Dehghani, Mahshid, et al.
Pubblicazione: (2024)
di: Dehghani, Mahshid, et al.
Pubblicazione: (2024)
Large Language Model for Qualitative Research -- A Systematic Mapping Study
di: Barros, Cauã Ferreira, et al.
Pubblicazione: (2024)
di: Barros, Cauã Ferreira, et al.
Pubblicazione: (2024)
Semantic Leakage from Image Embeddings
di: Chen, Yiyi, et al.
Pubblicazione: (2026)
di: Chen, Yiyi, et al.
Pubblicazione: (2026)
On the Limitations of Vision-Language Models in Understanding Image Transforms
di: Anis, Ahmad Mustafa, et al.
Pubblicazione: (2025)
di: Anis, Ahmad Mustafa, et al.
Pubblicazione: (2025)
SSD-GS: Scattering and Shadow Decomposition for Relightable 3D Gaussian Splatting
di: Zheng, Iris, et al.
Pubblicazione: (2026)
di: Zheng, Iris, et al.
Pubblicazione: (2026)
MSGS: Multispectral 3D Gaussian Splatting
di: Zheng, Iris, et al.
Pubblicazione: (2026)
di: Zheng, Iris, et al.
Pubblicazione: (2026)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
di: Masrourisaadat, Nila, et al.
Pubblicazione: (2024)
di: Masrourisaadat, Nila, et al.
Pubblicazione: (2024)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
di: Fan, Lin, et al.
Pubblicazione: (2026)
di: Fan, Lin, et al.
Pubblicazione: (2026)
A Scalable Pipeline Combining Procedural 3D Graphics and Guided Diffusion for Photorealistic Synthetic Training Data Generation in White Button Mushroom Segmentation
di: Károly, Artúr I., et al.
Pubblicazione: (2025)
di: Károly, Artúr I., et al.
Pubblicazione: (2025)
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
di: Koukounas, Andreas, et al.
Pubblicazione: (2024)
di: Koukounas, Andreas, et al.
Pubblicazione: (2024)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
di: Portelance, Eva, et al.
Pubblicazione: (2024)
di: Portelance, Eva, et al.
Pubblicazione: (2024)
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2021)
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2021)
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2025)
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2022)
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2022)
Documenti analoghi
-
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
di: Bian, Zhipeng, et al.
Pubblicazione: (2026) -
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
di: Tong, Jingqi, et al.
Pubblicazione: (2025) -
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025) -
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
di: Le, Van-Truong
Pubblicazione: (2026) -
GroundCap: A Visually Grounded Image Captioning Dataset
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)