Position: Towards Implicit Prompt For Text-To-Image Models
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yue, Lin, Yuqi, Liu, Hong, Shao, Wenqi, Chen, Runjian, Shang, Hailong, Wang, Yu, Qiao, Yu, Zhang, Kaipeng, Luo, Ping |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
by: Zhao, Lirui, et al.
Published: (2024)
by: Zhao, Lirui, et al.
Published: (2024)
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
by: Yang, Yue, et al.
Published: (2024)
by: Yang, Yue, et al.
Published: (2024)
JiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World Data
by: Chen, Runjian, et al.
Published: (2025)
by: Chen, Runjian, et al.
Published: (2025)
TREND: Unsupervised 3D Representation Learning via Temporal Forecasting for LiDAR Perception
by: Chen, Runjian, et al.
Published: (2024)
by: Chen, Runjian, et al.
Published: (2024)
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
by: Lin, Yuqi, et al.
Published: (2025)
by: Lin, Yuqi, et al.
Published: (2025)
Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
by: Ying, Kaining, et al.
Published: (2024)
by: Ying, Kaining, et al.
Published: (2024)
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
by: Zhang, Hao, et al.
Published: (2023)
by: Zhang, Hao, et al.
Published: (2023)
CLAP: Unsupervised 3D Representation Learning for Fusion 3D Perception via Curvature Sampling and Prototype Learning
by: Chen, Runjian, et al.
Published: (2024)
by: Chen, Runjian, et al.
Published: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
by: Xie, Yuxuan, et al.
Published: (2024)
by: Xie, Yuxuan, et al.
Published: (2024)
Data Adaptive Traceback for Vision-Language Foundation Models in Image Classification
by: Peng, Wenshuo, et al.
Published: (2024)
by: Peng, Wenshuo, et al.
Published: (2024)
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
by: Zhou, Pengfei, et al.
Published: (2024)
by: Zhou, Pengfei, et al.
Published: (2024)
Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Investigating Disability Representations in Text-to-Image Models
by: Tian, Yang, et al.
Published: (2026)
by: Tian, Yang, et al.
Published: (2026)
MLLMs-Augmented Visual-Language Representation Learning
by: Liu, Yanqing, et al.
Published: (2023)
by: Liu, Yanqing, et al.
Published: (2023)
Towards Reliable Verification of Unauthorized Data Usage in Personalized Text-to-Image Diffusion Models
by: Li, Boheng, et al.
Published: (2024)
by: Li, Boheng, et al.
Published: (2024)
A Framework for Critical Evaluation of Text-to-Image Models: Integrating Art Historical Analysis, Artistic Exploration, and Critical Prompt Engineering
by: Foka, Amalia
Published: (2024)
by: Foka, Amalia
Published: (2024)
CO^3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving
by: Chen, Runjian, et al.
Published: (2022)
by: Chen, Runjian, et al.
Published: (2022)
Position: Universal Aesthetic Alignment Narrows Artistic Expression
by: Guo, Wenqi Marshall, et al.
Published: (2025)
by: Guo, Wenqi Marshall, et al.
Published: (2025)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
by: Hall, Melissa, et al.
Published: (2024)
by: Hall, Melissa, et al.
Published: (2024)
TIPO: Text to Image with Text Presampling for Prompt Optimization
by: Yeh, Shih-Ying, et al.
Published: (2024)
by: Yeh, Shih-Ying, et al.
Published: (2024)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
by: Shao, Wenqi, et al.
Published: (2023)
by: Shao, Wenqi, et al.
Published: (2023)
FIGURA: A Modular Prompt Engineering Method for Artistic Figure Photography in Safety-Filtered Text-to-Image Models
by: Cazzaniga, Luca
Published: (2026)
by: Cazzaniga, Luca
Published: (2026)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models
by: Jin, Di, et al.
Published: (2024)
by: Jin, Di, et al.
Published: (2024)
Masked and Permuted Implicit Context Learning for Scene Text Recognition
by: Yang, Xiaomeng, et al.
Published: (2023)
by: Yang, Xiaomeng, et al.
Published: (2023)
KG-FairDiff: Knowledge Graph-Guided Prompt Refinement for Demographically Fair Text-to-Image Generation
by: Davoodi, Farbod, et al.
Published: (2026)
by: Davoodi, Farbod, et al.
Published: (2026)
SP-Guard: Selective Prompt-adaptive Guidance for Safe Text-to-Image Generation
by: Yu, Sumin, et al.
Published: (2025)
by: Yu, Sumin, et al.
Published: (2025)
TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models
by: Chinchure, Aditya, et al.
Published: (2023)
by: Chinchure, Aditya, et al.
Published: (2023)
SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
by: Yan, Xiangchao, et al.
Published: (2023)
by: Yan, Xiangchao, et al.
Published: (2023)
Interpersonal Relationship Analysis with Dyadic EEG Signals via Learning Spatial-Temporal Patterns
by: Ji, Wenqi, et al.
Published: (2024)
by: Ji, Wenqi, et al.
Published: (2024)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Generated Bias: Auditing Internal Bias Dynamics of Text-To-Image Generative Models
by: Mandal, Abhishek, et al.
Published: (2024)
by: Mandal, Abhishek, et al.
Published: (2024)
Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities
by: Alsudais, Abdulkareem
Published: (2025)
by: Alsudais, Abdulkareem
Published: (2025)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
Similar Items
-
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
by: Meng, Fanqing, et al.
Published: (2024) -
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
by: Zhao, Lirui, et al.
Published: (2024) -
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
by: Zhang, Hao, et al.
Published: (2024) -
Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
by: Yang, Yue, et al.
Published: (2024) -
JiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World Data
by: Chen, Runjian, et al.
Published: (2025)