K-Sort Arena: Efficient and Reliable Benchmarking for Generative Models via K-wise Human Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhikai, Liu, Xuewen, Fu, Dongrong Joe, Li, Jianquan, Gu, Qingyi, Keutzer, Kurt, Dong, Zhen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge
by: Li, Zhikai, et al.
Published: (2026)
by: Li, Zhikai, et al.
Published: (2026)
SoK: Privacy Personalised -- Mapping Personal Attributes \& Preferences of Privacy Mechanisms for Shoulder Surfing
by: Farzand, Habiba, et al.
Published: (2024)
by: Farzand, Habiba, et al.
Published: (2024)
PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction
by: Li, Jialin, et al.
Published: (2026)
by: Li, Jialin, et al.
Published: (2026)
FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
by: Yang, Qinglong, et al.
Published: (2025)
by: Yang, Qinglong, et al.
Published: (2025)
Prefer2SD: A Human-in-the-Loop Approach to Balancing Similarity and Diversity in In-Game Friend Recommendations
by: Wang, Xiyuan, et al.
Published: (2025)
by: Wang, Xiyuan, et al.
Published: (2025)
Arena as Offline Reward: Efficient Fine-Grained Preference Optimization for Diffusion Models
by: Li, Zhikai, et al.
Published: (2026)
by: Li, Zhikai, et al.
Published: (2026)
EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?
by: Lian, Zheng, et al.
Published: (2025)
by: Lian, Zheng, et al.
Published: (2025)
Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons
by: Zhu, Banghua, et al.
Published: (2023)
by: Zhu, Banghua, et al.
Published: (2023)
"What's Happening"- A Human-centered Multimodal Interpreter Explaining the Actions of Autonomous Vehicles
by: Luo, Xuewen, et al.
Published: (2025)
by: Luo, Xuewen, et al.
Published: (2025)
On Representing Humans' Soft-Ethics Preferences As Dispositions
by: Donati, Donatella, et al.
Published: (2024)
by: Donati, Donatella, et al.
Published: (2024)
A Multi-Agent Framework for Democratizing XR Content Creation in K-12 Classrooms
by: Chang, Yuan, et al.
Published: (2026)
by: Chang, Yuan, et al.
Published: (2026)
ECNUClaw: A Learner-Profiled Intelligent Study Companion Framework for K-12 Personalized Education
by: Zhou, Yizhou, et al.
Published: (2026)
by: Zhou, Yizhou, et al.
Published: (2026)
DesignBridge: Bridging Designer Expertise and User Preferences through AI-Enhanced Co-Design for Fashion
by: Shao, Yuheng, et al.
Published: (2026)
by: Shao, Yuheng, et al.
Published: (2026)
SoK: Synthesizing Smart Home Privacy Protection Mechanisms Across Academic Proposals and Commercial Documentations
by: Zhang, Shuning, et al.
Published: (2025)
by: Zhang, Shuning, et al.
Published: (2025)
ActiveAI: Enabling K-12 AI Literacy Education & Analytics at Scale
by: Xiao, Ruiwei, et al.
Published: (2024)
by: Xiao, Ruiwei, et al.
Published: (2024)
Beyond Text: Probing K-12 Educators' Perspectives and Ideas for Learning Opportunities Leveraging Multimodal Large Language Models
by: Tseng, Tiffany, et al.
Published: (2025)
by: Tseng, Tiffany, et al.
Published: (2025)
What Social Media Use Do People Regret? An Analysis of 34K Smartphone Screenshots with Multimodal LLM
by: Guo, Longjie, et al.
Published: (2024)
by: Guo, Longjie, et al.
Published: (2024)
WeeCare: Towards Handheld Bladder Fullness Sensing with a Conformable Pad
by: Qin, Zhikai, et al.
Published: (2026)
by: Qin, Zhikai, et al.
Published: (2026)
K-QA: A Real-World Medical Q&A Benchmark
by: Manes, Itay, et al.
Published: (2024)
by: Manes, Itay, et al.
Published: (2024)
Understanding Password Preferences, Memorability, and Security through a Human-Centered Lens
by: Paker, Duru, et al.
Published: (2026)
by: Paker, Duru, et al.
Published: (2026)
Preference-Guided Prompt Optimization for Text-to-Image Generation
by: Li, Zhipeng, et al.
Published: (2026)
by: Li, Zhipeng, et al.
Published: (2026)
A Comprehensive Review of Human Error in Risk-Informed Decision Making: Integrating Human Reliability Assessment, Artificial Intelligence, and Human Performance Models
by: Xiao, Xingyu, et al.
Published: (2025)
by: Xiao, Xingyu, et al.
Published: (2025)
Learning Multimodal Confidence for Intention Recognition in Human-Robot Interaction
by: Zhao, Xiyuan, et al.
Published: (2024)
by: Zhao, Xiyuan, et al.
Published: (2024)
LLMs to Support K-12 Teachers in Culturally Relevant Pedagogy: An AI Literacy Example
by: Wang, Jiayi, et al.
Published: (2025)
by: Wang, Jiayi, et al.
Published: (2025)
Assessing Data Literacy in K-12 Education: Challenges and Opportunities
by: Goldman, Annabel, et al.
Published: (2026)
by: Goldman, Annabel, et al.
Published: (2026)
Towards Aligning Personalized Conversational Recommendation Agents with Users' Privacy Preferences
by: Zhang, Shuning, et al.
Published: (2025)
by: Zhang, Shuning, et al.
Published: (2025)
Efficient Visual Appearance Optimization by Learning from Prior Preferences
by: Li, Zhipeng, et al.
Published: (2025)
by: Li, Zhipeng, et al.
Published: (2025)
Card Sorting Simulator: Augmenting Design of Logical Information Architectures with Large Language Models
by: Kuric, Eduard, et al.
Published: (2025)
by: Kuric, Eduard, et al.
Published: (2025)
CWEFS: Brain volume conduction effects inspired channel-wise EEG feature selection for multi-dimensional emotion recognition
by: Xu, Xueyuan, et al.
Published: (2025)
by: Xu, Xueyuan, et al.
Published: (2025)
The Human Factor in Data Cleaning: Exploring Preferences and Biases
by: AbdElazim, Hazim, et al.
Published: (2026)
by: AbdElazim, Hazim, et al.
Published: (2026)
AI-Generated Rubric Interfaces: K-12 Teachers' Perceptions and Practices
by: Riahi, Bahare, et al.
Published: (2026)
by: Riahi, Bahare, et al.
Published: (2026)
Impact of Cognitive Load on Human Trust in Hybrid Human-Robot Collaboration
by: Guo, Hao, et al.
Published: (2024)
by: Guo, Hao, et al.
Published: (2024)
Understanding Dynamic Human-Robot Proxemics in the Case of Four-Legged Canine-Inspired Robots
by: Xu, Xiangmin, et al.
Published: (2023)
by: Xu, Xiangmin, et al.
Published: (2023)
Demonstrating HumanTHOR: A Simulation Platform and Benchmark for Human-Robot Collaboration in a Shared Workspace
by: Wang, Chenxu, et al.
Published: (2024)
by: Wang, Chenxu, et al.
Published: (2024)
Pick and Sort for Graphical Authentication
by: Rahartomo, Argianto, et al.
Published: (2026)
by: Rahartomo, Argianto, et al.
Published: (2026)
Emerging Patterns of GenAI Use in K-12 Science and Mathematics Education
by: Esbenshade, Lief, et al.
Published: (2025)
by: Esbenshade, Lief, et al.
Published: (2025)
A Benchmark to Assess Common Ground in Human-AI Collaboration
by: Poelitz, Christian, et al.
Published: (2026)
by: Poelitz, Christian, et al.
Published: (2026)
FARPLS: A Feature-Augmented Robot Trajectory Preference Labeling System to Assist Human Labelers' Preference Elicitation
by: Lyu, Hanfang, et al.
Published: (2024)
by: Lyu, Hanfang, et al.
Published: (2024)
AI Literacy Education for Older Adults: Motivations, Challenges and Preferences
by: KangJie, Eugene Tang, et al.
Published: (2025)
by: KangJie, Eugene Tang, et al.
Published: (2025)
The Effects of Generative AI on Computing Students' Help-Seeking Preferences
by: Hou, Irene, et al.
Published: (2024)
by: Hou, Irene, et al.
Published: (2024)
Similar Items
-
K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge
by: Li, Zhikai, et al.
Published: (2026) -
SoK: Privacy Personalised -- Mapping Personal Attributes \& Preferences of Privacy Mechanisms for Shoulder Surfing
by: Farzand, Habiba, et al.
Published: (2024) -
PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction
by: Li, Jialin, et al.
Published: (2026) -
FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
by: Yang, Qinglong, et al.
Published: (2025) -
Prefer2SD: A Human-in-the-Loop Approach to Balancing Similarity and Diversity in In-Game Friend Recommendations
by: Wang, Xiyuan, et al.
Published: (2025)