Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Masrourisaadat, Nila, Sedaghatkish, Nazanin, Sarshartehrani, Fatemeh, Fox, Edward A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
Pointing-Based Object Recognition
von: Hajdúch, Lukáš, et al.
Veröffentlicht: (2026)
von: Hajdúch, Lukáš, et al.
Veröffentlicht: (2026)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
von: Semenov, Andrei, et al.
Veröffentlicht: (2024)
von: Semenov, Andrei, et al.
Veröffentlicht: (2024)
Survey Transfer Learning: Recycling Data with Silicon Responses
von: Amini, Ali
Veröffentlicht: (2025)
von: Amini, Ali
Veröffentlicht: (2025)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
HATL: Hierarchical Adaptive-Transfer Learning Framework for Sign Language Machine Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2026)
von: Shahin, Nada, et al.
Veröffentlicht: (2026)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
von: Boumber, Dainis, et al.
Veröffentlicht: (2024)
von: Boumber, Dainis, et al.
Veröffentlicht: (2024)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
von: Marmoret, Axel, et al.
Veröffentlicht: (2025)
von: Marmoret, Axel, et al.
Veröffentlicht: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
von: Azov, Guy, et al.
Veröffentlicht: (2026)
von: Azov, Guy, et al.
Veröffentlicht: (2026)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
von: Agia, Christopher, et al.
Veröffentlicht: (2024)
von: Agia, Christopher, et al.
Veröffentlicht: (2024)
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
von: Du, Guanchen, et al.
Veröffentlicht: (2025)
von: Du, Guanchen, et al.
Veröffentlicht: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2025)
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2025)
Learning 3D object-centric representation through prediction
von: Day, John, et al.
Veröffentlicht: (2024)
von: Day, John, et al.
Veröffentlicht: (2024)
S-HR-VQVAE: Sequential Hierarchical Residual Learning Vector Quantized Variational Autoencoder for Video Prediction
von: Adiban, Mohammad, et al.
Veröffentlicht: (2023)
von: Adiban, Mohammad, et al.
Veröffentlicht: (2023)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
Text-to-Events: Synthetic Event Camera Streams from Conditional Text Input
von: Ott, Joachim, et al.
Veröffentlicht: (2024)
von: Ott, Joachim, et al.
Veröffentlicht: (2024)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
von: Seo, Huichan, et al.
Veröffentlicht: (2025)
von: Seo, Huichan, et al.
Veröffentlicht: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
Cora: Correspondence-aware image editing using few step diffusion
von: Alimohammadi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Alimohammadi, Amirhossein, et al.
Veröffentlicht: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
WaveMix: A Resource-efficient Neural Network for Image Analysis
von: Jeevan, Pranav, et al.
Veröffentlicht: (2022)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2022)
Deployment-Time Reliability of Learned Robot Policies
von: Agia, Christopher
Veröffentlicht: (2026)
von: Agia, Christopher
Veröffentlicht: (2026)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment
von: Ye, Hua, et al.
Veröffentlicht: (2025)
von: Ye, Hua, et al.
Veröffentlicht: (2025)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
von: Li, Danyang, et al.
Veröffentlicht: (2025)
von: Li, Danyang, et al.
Veröffentlicht: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
Leum-VL Technical Report
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
FEDTAIL: Federated Long-Tailed Domain Generalization with Sharpness-Guided Gradient Matching
von: Gupta, Sunny, et al.
Veröffentlicht: (2025)
von: Gupta, Sunny, et al.
Veröffentlicht: (2025)
NV3D: Leveraging Spatial Shape Through Normal Vector-based 3D Object Detection
von: Chaowakarn, Krittin, et al.
Veröffentlicht: (2025)
von: Chaowakarn, Krittin, et al.
Veröffentlicht: (2025)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025) -
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023) -
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026) -
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025) -
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
von: Portelance, Eva, et al.
Veröffentlicht: (2024)