Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Raphael, Zhang, Crystina, Li, Wenyan, Lai, Carmen, Stenetorp, Pontus, Lu, Yao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining
by: Shao, Jiandong, et al.
Published: (2026)
by: Shao, Jiandong, et al.
Published: (2026)
Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt Optimisation
by: Lu, Yao, et al.
Published: (2023)
by: Lu, Yao, et al.
Published: (2023)
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
by: Tang, Raphael, et al.
Published: (2024)
by: Tang, Raphael, et al.
Published: (2024)
Multilingual Pretraining Using a Large Corpus Machine-Translated from a Single Source Language
by: Wang, Jiayi, et al.
Published: (2024)
by: Wang, Jiayi, et al.
Published: (2024)
Quantifying Generative Media Bias with a Corpus of Real-world and Generated News Articles
by: Trhlik, Filip, et al.
Published: (2024)
by: Trhlik, Filip, et al.
Published: (2024)
Multilingual Language Model Pretraining using Machine-translated Data
by: Wang, Jiayi, et al.
Published: (2025)
by: Wang, Jiayi, et al.
Published: (2025)
Jet Expansions of Residual Computation
by: Chen, Yihong, et al.
Published: (2024)
by: Chen, Yihong, et al.
Published: (2024)
LTD-Bench: Evaluating Large Language Models by Letting Them Draw
by: Lin, Liuhao, et al.
Published: (2025)
by: Lin, Liuhao, et al.
Published: (2025)
DrawingBench: Evaluating Spatial Reasoning and UI Interaction Capabilities of Large Language Models through Mouse-Based Drawing Tasks
by: Kim, Hyunjun, et al.
Published: (2025)
by: Kim, Hyunjun, et al.
Published: (2025)
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
by: Madaan, Lovish, et al.
Published: (2024)
by: Madaan, Lovish, et al.
Published: (2024)
Constructing a Norm for Children's Scientific Drawing: Distribution Features Based on Semantic Similarity of Large Language Models
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Rethinking Human Preference Evaluation of LLM Rationales
by: Li, Ziang, et al.
Published: (2025)
by: Li, Ziang, et al.
Published: (2025)
Using Natural Language Explanations to Improve Robustness of In-context Learning
by: He, Xuanli, et al.
Published: (2023)
by: He, Xuanli, et al.
Published: (2023)
Gender-specific Machine Translation with Large Language Models
by: Sánchez, Eduardo, et al.
Published: (2023)
by: Sánchez, Eduardo, et al.
Published: (2023)
Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation
by: Cui, Zhiqing, et al.
Published: (2025)
by: Cui, Zhiqing, et al.
Published: (2025)
Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts?
by: Zhang, Crystina, et al.
Published: (2024)
by: Zhang, Crystina, et al.
Published: (2024)
Infrequent Child-Directed Speech Is Bursty and May Draw Infant Vocalizations
by: Cychosz, Margaret, et al.
Published: (2026)
by: Cychosz, Margaret, et al.
Published: (2026)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
by: Chiang, Wei-Lin, et al.
Published: (2024)
by: Chiang, Wei-Lin, et al.
Published: (2024)
PsyDraw: A Multi-Agent Multimodal System for Mental Health Screening in Left-Behind Children
by: Zhang, Yiqun, et al.
Published: (2024)
by: Zhang, Yiqun, et al.
Published: (2024)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction
by: Hu, Juncheng, et al.
Published: (2026)
by: Hu, Juncheng, et al.
Published: (2026)
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
by: Baral, Sami, et al.
Published: (2025)
by: Baral, Sami, et al.
Published: (2025)
Linguini: A benchmark for language-agnostic linguistic reasoning
by: Sánchez, Eduardo, et al.
Published: (2024)
by: Sánchez, Eduardo, et al.
Published: (2024)
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
by: Lucy, Li, et al.
Published: (2026)
by: Lucy, Li, et al.
Published: (2026)
Translate, then Detect: Leveraging Machine Translation for Cross-Lingual Toxicity Classification
by: Bell, Samuel J., et al.
Published: (2025)
by: Bell, Samuel J., et al.
Published: (2025)
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
by: Kundurthy, Srivatsa, et al.
Published: (2026)
by: Kundurthy, Srivatsa, et al.
Published: (2026)
Warmup Generations: A Task-Agnostic Approach for Guiding Sequence-to-Sequence Learning with Unsupervised Initial State Generation
by: Li, Senyu, et al.
Published: (2025)
by: Li, Senyu, et al.
Published: (2025)
Semantic Draw Engineering for Text-to-Image Creation
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
APM: Evaluating Style Personalization in LLMs with Arbitrary Preference Mappings
by: Spohn, Philipp, et al.
Published: (2026)
by: Spohn, Philipp, et al.
Published: (2026)
3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience
by: Xiao, Hongcan, et al.
Published: (2026)
by: Xiao, Hongcan, et al.
Published: (2026)
MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts
by: Li, Weiyue, et al.
Published: (2026)
by: Li, Weiyue, et al.
Published: (2026)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
by: Hakimov, Sherzod, et al.
Published: (2026)
by: Hakimov, Sherzod, et al.
Published: (2026)
GameArena: Evaluating LLM Reasoning through Live Computer Games
by: Hu, Lanxiang, et al.
Published: (2024)
by: Hu, Lanxiang, et al.
Published: (2024)
Style over Story: Measuring LLM Narrative Preferences via Structured Selection
by: Jung, Donghoon, et al.
Published: (2025)
by: Jung, Donghoon, et al.
Published: (2025)
Improving Language Plasticity via Pretraining with Active Forgetting
by: Chen, Yihong, et al.
Published: (2023)
by: Chen, Yihong, et al.
Published: (2023)
GraphArena: Evaluating and Exploring Large Language Models on Graph Computation
by: Tang, Jianheng, et al.
Published: (2024)
by: Tang, Jianheng, et al.
Published: (2024)
Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
by: Zhao, Ruochen, et al.
Published: (2024)
by: Zhao, Ruochen, et al.
Published: (2024)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences
by: Wu, Eric, et al.
Published: (2026)
by: Wu, Eric, et al.
Published: (2026)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
Similar Items
-
The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining
by: Shao, Jiandong, et al.
Published: (2026) -
Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt Optimisation
by: Lu, Yao, et al.
Published: (2023) -
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
by: Tang, Raphael, et al.
Published: (2024) -
Multilingual Pretraining Using a Large Corpus Machine-Translated from a Single Source Language
by: Wang, Jiayi, et al.
Published: (2024) -
Quantifying Generative Media Bias with a Corpus of Real-world and Generated News Articles
by: Trhlik, Filip, et al.
Published: (2024)