WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lin, Bill Yuchen, Deng, Yuntian, Chandu, Khyathi, Brahman, Faeze, Ravichander, Abhilasha, Pyatkin, Valentina, Dziri, Nouha, Bras, Ronan Le, Choi, Yejin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
RESTOR: Knowledge Recovery in Machine Unlearning
par: Rezaei, Keivan, et autres
Publié: (2024)
par: Rezaei, Keivan, et autres
Publié: (2024)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
par: Yin, Da, et autres
Publié: (2023)
par: Yin, Da, et autres
Publié: (2023)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
par: Zhao, Wenting, et autres
Publié: (2024)
par: Zhao, Wenting, et autres
Publié: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
par: Brahman, Faeze, et autres
Publié: (2024)
par: Brahman, Faeze, et autres
Publié: (2024)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
par: Rao, Kavel, et autres
Publié: (2023)
par: Rao, Kavel, et autres
Publié: (2023)
RewardBench: Evaluating Reward Models for Language Modeling
par: Lambert, Nathan, et autres
Publié: (2024)
par: Lambert, Nathan, et autres
Publié: (2024)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
par: Jiang, Liwei, et autres
Publié: (2024)
par: Jiang, Liwei, et autres
Publié: (2024)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
par: Han, Seungju, et autres
Publié: (2024)
par: Han, Seungju, et autres
Publié: (2024)
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
par: Baheti, Ashutosh, et autres
Publié: (2024)
par: Baheti, Ashutosh, et autres
Publié: (2024)
MacGyver: Are Large Language Models Creative Problem Solvers?
par: Tian, Yufei, et autres
Publié: (2023)
par: Tian, Yufei, et autres
Publié: (2023)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
par: Jung, Jaehun, et autres
Publié: (2024)
par: Jung, Jaehun, et autres
Publié: (2024)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
par: Graf, Victoria, et autres
Publié: (2026)
par: Graf, Victoria, et autres
Publié: (2026)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
par: Lu, Yujie, et autres
Publié: (2024)
par: Lu, Yujie, et autres
Publié: (2024)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
par: Ravichander, Abhilasha, et autres
Publié: (2025)
par: Ravichander, Abhilasha, et autres
Publié: (2025)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
par: Röttger, Paul, et autres
Publié: (2025)
par: Röttger, Paul, et autres
Publié: (2025)
WildChat: 1M ChatGPT Interaction Logs in the Wild
par: Zhao, Wenting, et autres
Publié: (2024)
par: Zhao, Wenting, et autres
Publié: (2024)
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
par: Deng, Yuntian, et autres
Publié: (2024)
par: Deng, Yuntian, et autres
Publié: (2024)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
par: Lu, Ximing, et autres
Publié: (2024)
par: Lu, Ximing, et autres
Publié: (2024)
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
par: Yamada, Yutaro, et autres
Publié: (2024)
par: Yamada, Yutaro, et autres
Publié: (2024)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
par: Sun, Yiyou, et autres
Publié: (2025)
par: Sun, Yiyou, et autres
Publié: (2025)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
par: Srinivasan, Tejas, et autres
Publié: (2024)
par: Srinivasan, Tejas, et autres
Publié: (2024)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
par: Lin, Bill Yuchen, et autres
Publié: (2025)
par: Lin, Bill Yuchen, et autres
Publié: (2025)
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions
par: Zhou, Xuhui, et autres
Publié: (2024)
par: Zhou, Xuhui, et autres
Publié: (2024)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
par: Zhang, Linhao, et autres
Publié: (2025)
par: Zhang, Linhao, et autres
Publié: (2025)
Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
par: Baheti, Ashutosh, et autres
Publié: (2023)
par: Baheti, Ashutosh, et autres
Publié: (2023)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
par: Balepur, Nishant, et autres
Publié: (2024)
par: Balepur, Nishant, et autres
Publié: (2024)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
par: Qiu, Linlu, et autres
Publié: (2023)
par: Qiu, Linlu, et autres
Publié: (2023)
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
par: Brahman, Faeze, et autres
Publié: (2023)
par: Brahman, Faeze, et autres
Publié: (2023)
WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
par: Wang, Pengyu, et autres
Publié: (2026)
par: Wang, Pengyu, et autres
Publié: (2026)
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
par: Xu, Zhangchen, et autres
Publié: (2024)
par: Xu, Zhangchen, et autres
Publié: (2024)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
par: Li, Jing-Jing, et autres
Publié: (2024)
par: Li, Jing-Jing, et autres
Publié: (2024)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
par: Bennion, Jonathan, et autres
Publié: (2025)
par: Bennion, Jonathan, et autres
Publié: (2025)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
par: Chandu, Khyathi Raghavi, et autres
Publié: (2024)
par: Chandu, Khyathi Raghavi, et autres
Publié: (2024)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
par: Mendelsohn, Julia, et autres
Publié: (2023)
par: Mendelsohn, Julia, et autres
Publié: (2023)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
par: Li, Huihan, et autres
Publié: (2024)
par: Li, Huihan, et autres
Publié: (2024)
What Has Been Lost with Synthetic Evaluation?
par: Gill, Alexander, et autres
Publié: (2025)
par: Gill, Alexander, et autres
Publié: (2025)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
par: Sorensen, Taylor, et autres
Publié: (2023)
par: Sorensen, Taylor, et autres
Publié: (2023)
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
par: Ravichander, Abhilasha, et autres
Publié: (2025)
par: Ravichander, Abhilasha, et autres
Publié: (2025)
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
par: Wang, Jiayu, et autres
Publié: (2025)
par: Wang, Jiayu, et autres
Publié: (2025)
Information-Theoretic Distillation for Reference-less Summarization
par: Jung, Jaehun, et autres
Publié: (2024)
par: Jung, Jaehun, et autres
Publié: (2024)
Documents similaires
-
RESTOR: Knowledge Recovery in Machine Unlearning
par: Rezaei, Keivan, et autres
Publié: (2024) -
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
par: Yin, Da, et autres
Publié: (2023) -
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
par: Zhao, Wenting, et autres
Publié: (2024) -
The Art of Saying No: Contextual Noncompliance in Language Models
par: Brahman, Faeze, et autres
Publié: (2024) -
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
par: Rao, Kavel, et autres
Publié: (2023)