LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | White, Colin, Dooley, Samuel, Roberts, Manley, Pal, Arka, Feuer, Ben, Jain, Siddhartha, Shwartz-Ziv, Ravid, Jain, Neel, Saifullah, Khalid, Dey, Sreemanti, Shubh-Agrawal, Sandha, Sandeep Singh, Naidu, Siddartha, Hegde, Chinmay, LeCun, Yann, Goldstein, Tom, Neiswanger, Willie, Goldblum, Micah |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Training in Imagination
by: Timor, Nadav, et al.
Published: (2026)
by: Timor, Nadav, et al.
Published: (2026)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
Just How Flexible are Neural Networks in Practice?
by: Shwartz-Ziv, Ravid, et al.
Published: (2024)
by: Shwartz-Ziv, Ravid, et al.
Published: (2024)
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
by: Goldfeder, Judah, et al.
Published: (2026)
by: Goldfeder, Judah, et al.
Published: (2026)
The Entropy Enigma: Success and Failure of Entropy Minimization
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
by: Skean, Oscar, et al.
Published: (2024)
by: Skean, Oscar, et al.
Published: (2024)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
by: Pal, Arka, et al.
Published: (2024)
by: Pal, Arka, et al.
Published: (2024)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
by: Shani, Chen, et al.
Published: (2025)
by: Shani, Chen, et al.
Published: (2025)
Variance-Covariance Regularization Improves Representation Learning
by: Zhu, Jiachen, et al.
Published: (2023)
by: Zhu, Jiachen, et al.
Published: (2023)
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training
by: Feuer, Benjamin, et al.
Published: (2025)
by: Feuer, Benjamin, et al.
Published: (2025)
An Information-Theoretic Perspective on Variance-Invariance-Covariance Regularization
by: Shwartz-Ziv, Ravid, et al.
Published: (2023)
by: Shwartz-Ziv, Ravid, et al.
Published: (2023)
Rate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty Estimation
by: Zeevi, Tal, et al.
Published: (2024)
by: Zeevi, Tal, et al.
Published: (2024)
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
by: Patel, Niket, et al.
Published: (2024)
by: Patel, Niket, et al.
Published: (2024)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
by: Ioannides, Georgios, et al.
Published: (2025)
by: Ioannides, Georgios, et al.
Published: (2025)
TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks
by: Feuer, Benjamin, et al.
Published: (2024)
by: Feuer, Benjamin, et al.
Published: (2024)
Layer by Layer: Uncovering Hidden Representations in Language Models
by: Skean, Oscar, et al.
Published: (2025)
by: Skean, Oscar, et al.
Published: (2025)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
by: Arefin, Md Rifat, et al.
Published: (2024)
by: Arefin, Md Rifat, et al.
Published: (2024)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
by: Ioannides, Georgios, et al.
Published: (2026)
by: Ioannides, Georgios, et al.
Published: (2026)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
by: Devic, Siddartha, et al.
Published: (2025)
by: Devic, Siddartha, et al.
Published: (2025)
When Do Neural Nets Outperform Boosted Trees on Tabular Data?
by: McElfresh, Duncan, et al.
Published: (2023)
by: McElfresh, Duncan, et al.
Published: (2023)
MARVIS: Modality Adaptive Reasoning over VISualizations
by: Feuer, Benjamin, et al.
Published: (2025)
by: Feuer, Benjamin, et al.
Published: (2025)
ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models
by: Feuer, Benjamin, et al.
Published: (2023)
by: Feuer, Benjamin, et al.
Published: (2023)
Large Language Models Must Be Taught to Know What They Don't Know
by: Kapoor, Sanyam, et al.
Published: (2024)
by: Kapoor, Sanyam, et al.
Published: (2024)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
by: Zhang, Eva, et al.
Published: (2024)
by: Zhang, Eva, et al.
Published: (2024)
Dynamic Delayed Tree Expansion For Improved Multi-Path Speculative Decoding
by: Thomas, Rahul, et al.
Published: (2026)
by: Thomas, Rahul, et al.
Published: (2026)
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
by: Feuer, Benjamin, et al.
Published: (2024)
by: Feuer, Benjamin, et al.
Published: (2024)
Closing the Train-Test Gap in World Models for Gradient-Based Planning
by: Parthasarathy, Arjun, et al.
Published: (2025)
by: Parthasarathy, Arjun, et al.
Published: (2025)
Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models
by: Paech, Samuel, et al.
Published: (2025)
by: Paech, Samuel, et al.
Published: (2025)
FaceCloak: Learning to Protect Face Templates
by: Banerjee, Sudipta, et al.
Published: (2025)
by: Banerjee, Sudipta, et al.
Published: (2025)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
by: Sanyal, Sunny, et al.
Published: (2024)
by: Sanyal, Sunny, et al.
Published: (2024)
Privacy-Preserving Mechanisms Enable Cheap Verifiable Inference of LLMs
by: Pal, Arka, et al.
Published: (2026)
by: Pal, Arka, et al.
Published: (2026)
Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
by: Pal, Arka, et al.
Published: (2025)
by: Pal, Arka, et al.
Published: (2025)
Hidden in the Noise: Two-Stage Robust Watermarking for Images
by: Arabi, Kasra, et al.
Published: (2024)
by: Arabi, Kasra, et al.
Published: (2024)
Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces
by: Shaar, Eitan, et al.
Published: (2026)
by: Shaar, Eitan, et al.
Published: (2026)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
by: Jain, Neel, et al.
Published: (2024)
by: Jain, Neel, et al.
Published: (2024)
Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
by: Yost, Alexandra, et al.
Published: (2025)
by: Yost, Alexandra, et al.
Published: (2025)
SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification
by: Feuer, Benjamin, et al.
Published: (2024)
by: Feuer, Benjamin, et al.
Published: (2024)
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs
by: Chen, Angelica, et al.
Published: (2023)
by: Chen, Angelica, et al.
Published: (2023)
Re-analysis of the Human Transcription Factor Atlas Recovers TF-Specific Signatures from Pooled Single-Cell Screens with Missing Controls
by: Jain, Arka, et al.
Published: (2026)
by: Jain, Arka, et al.
Published: (2026)
Similar Items
-
On Training in Imagination
by: Timor, Nadav, et al.
Published: (2026) -
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024) -
Just How Flexible are Neural Networks in Practice?
by: Shwartz-Ziv, Ravid, et al.
Published: (2024) -
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
by: Goldfeder, Judah, et al.
Published: (2026) -
The Entropy Enigma: Success and Failure of Entropy Minimization
by: Press, Ori, et al.
Published: (2024)