AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Hanjun, Huang, Zhimu, Chung, Sylvia, Wang, Yiran, Jin, Yingbin, Li, Jialin, Li, Jiang, Li, Xinfeng, Salam, Hanan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
by: Luo, Hanjun, et al.
Published: (2025)
by: Luo, Hanjun, et al.
Published: (2025)
BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models
by: Luo, Hanjun, et al.
Published: (2026)
by: Luo, Hanjun, et al.
Published: (2026)
PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction
by: Li, Jialin, et al.
Published: (2026)
by: Li, Jialin, et al.
Published: (2026)
AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
by: Luo, Hanjun, et al.
Published: (2025)
by: Luo, Hanjun, et al.
Published: (2025)
BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models
by: Luo, Hanjun, et al.
Published: (2024)
by: Luo, Hanjun, et al.
Published: (2024)
Improving Personalisation in Valence and Arousal Prediction using Data Augmentation
by: Nwadike, Munachiso, et al.
Published: (2024)
by: Nwadike, Munachiso, et al.
Published: (2024)
Beyond One-Size-Fits-All: A Survey of Personalized Affective Computing in Human-Agent Interaction
by: Li, Jialin, et al.
Published: (2023)
by: Li, Jialin, et al.
Published: (2023)
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
by: Sarangi, Sneheel, et al.
Published: (2025)
by: Sarangi, Sneheel, et al.
Published: (2025)
DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity Recognition
by: Luo, Hanjun, et al.
Published: (2024)
by: Luo, Hanjun, et al.
Published: (2024)
BatchEval: Towards Human-like Text Evaluation
by: Yuan, Peiwen, et al.
Published: (2023)
by: Yuan, Peiwen, et al.
Published: (2023)
T2IBias: Uncovering Societal Bias Encoded in the Latent Space of Text-to-Image Generative Models
by: Sufian, Abu, et al.
Published: (2025)
by: Sufian, Abu, et al.
Published: (2025)
FAIntbench: A Holistic and Precise Benchmark for Bias Evaluation in Text-to-Image Models
by: Luo, Hanjun, et al.
Published: (2024)
by: Luo, Hanjun, et al.
Published: (2024)
POMP: Probability-driven Meta-graph Prompter for LLMs in Low-resource Unsupervised Neural Machine Translation
by: Pan, Shilong, et al.
Published: (2024)
by: Pan, Shilong, et al.
Published: (2024)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
by: Paulus, Anselm, et al.
Published: (2024)
by: Paulus, Anselm, et al.
Published: (2024)
VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image Synthesis
by: Wu, Shiyu, et al.
Published: (2025)
by: Wu, Shiyu, et al.
Published: (2025)
CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making
by: Rahman, Hasibur, et al.
Published: (2025)
by: Rahman, Hasibur, et al.
Published: (2025)
Supporting Productivity Skill Development in College Students through Social Robot Coaching: A Proof-of-Concept
by: Lalwani, Himanshi, et al.
Published: (2025)
by: Lalwani, Himanshi, et al.
Published: (2025)
The Supportiveness-Safety Tradeoff in LLM Well-Being Agents
by: Lalwani, Himanshi, et al.
Published: (2026)
by: Lalwani, Himanshi, et al.
Published: (2026)
Ethically-Aware Participatory Design of a Productivity Social Robot for College Students
by: Lalwani, Himanshi, et al.
Published: (2025)
by: Lalwani, Himanshi, et al.
Published: (2025)
ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions
by: Li, Jialin, et al.
Published: (2026)
by: Li, Jialin, et al.
Published: (2026)
EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation
by: Han, Shuhao, et al.
Published: (2024)
by: Han, Shuhao, et al.
Published: (2024)
Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries
by: Li, Xinfeng, et al.
Published: (2025)
by: Li, Xinfeng, et al.
Published: (2025)
On Style: An Atelier
Published: (2019)
Published: (2019)
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
Informing Robot Wellbeing Coach Design through Longitudinal Analysis of Human-AI Dialogue
by: Shah, Keya, et al.
Published: (2026)
by: Shah, Keya, et al.
Published: (2026)
IP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual Prompting
by: Zhang, Yuxin, et al.
Published: (2025)
by: Zhang, Yuxin, et al.
Published: (2025)
CoPrompter: User-Centric Evaluation of LLM Instruction Alignment for Improved Prompt Engineering
by: Joshi, Ishika, et al.
Published: (2024)
by: Joshi, Ishika, et al.
Published: (2024)
DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models
by: Wang, Juntong, et al.
Published: (2026)
by: Wang, Juntong, et al.
Published: (2026)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
by: Cheng, Zhili, et al.
Published: (2025)
by: Cheng, Zhili, et al.
Published: (2025)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
by: Wang, Ganghua, et al.
Published: (2025)
by: Wang, Ganghua, et al.
Published: (2025)
Ateliers d'Anthropologie
Published: (2013)
Published: (2013)
Atelier de Kayar
by: Dioh, B., et al.
Published: (2002)
by: Dioh, B., et al.
Published: (2002)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023)
by: Shu, Lei, et al.
Published: (2023)
GPS: General Per-Sample Prompter
by: Batorski, Pawel, et al.
Published: (2025)
by: Batorski, Pawel, et al.
Published: (2025)
Mixture of In-Context Prompters for Tabular PFNs
by: Xu, Derek, et al.
Published: (2024)
by: Xu, Derek, et al.
Published: (2024)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
by: Zhou, Lingfeng, et al.
Published: (2025)
by: Zhou, Lingfeng, et al.
Published: (2025)
Aggregated Text Transformer for Scene Text Detection
by: Zhou, Zhao, et al.
Published: (2022)
by: Zhou, Zhao, et al.
Published: (2022)
SparQLe: Speech Queries to Text Translation Through LLMs
by: Djanibekov, Amirbek, et al.
Published: (2025)
by: Djanibekov, Amirbek, et al.
Published: (2025)
LLM as Prompter: Low-resource Inductive Reasoning on Arbitrary Knowledge Graphs
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
GROW: A Conversational AI Coach for Goals, Reflection, Optimism, and Well-Being
by: Shah, Keya, et al.
Published: (2026)
by: Shah, Keya, et al.
Published: (2026)
Similar Items
-
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
by: Luo, Hanjun, et al.
Published: (2025) -
BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models
by: Luo, Hanjun, et al.
Published: (2026) -
PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction
by: Li, Jialin, et al.
Published: (2026) -
AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
by: Luo, Hanjun, et al.
Published: (2025) -
BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models
by: Luo, Hanjun, et al.
Published: (2024)