ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Ghosh, Adhiraj, Dziadzio, Sebastian, Prabhu, Ameya, Udandarao, Vishaal, Albanie, Samuel, Bethge, Matthias |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Merge Your Multimodal Models Over Time?
by: Dziadzio, Sebastian, et al.
Published: (2024)
by: Dziadzio, Sebastian, et al.
Published: (2024)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025)
by: Hochlehnert, Andreas, et al.
Published: (2025)
A Practitioner's Guide to Continual Multimodal Pretraining
by: Roth, Karsten, et al.
Published: (2024)
by: Roth, Karsten, et al.
Published: (2024)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
CiteME: Can Language Models Accurately Cite Scientific Claims?
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
by: Udandarao, Vishaal, et al.
Published: (2025)
by: Udandarao, Vishaal, et al.
Published: (2025)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
by: Skorobogat, Ronald, et al.
Published: (2026)
by: Skorobogat, Ronald, et al.
Published: (2026)
Solving Spatial Supersensing Without Spatial Supersensing
by: Udandarao, Vishaal, et al.
Published: (2025)
by: Udandarao, Vishaal, et al.
Published: (2025)
Scaling Open-Ended Reasoning to Predict the Future
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Concept-Aware Batch Sampling Improves Language-Image Pretraining
by: Ghosh, Adhiraj, et al.
Published: (2025)
by: Ghosh, Adhiraj, et al.
Published: (2025)
Mapping Post-Training Forgetting in Language Models at Scale
by: Harmon, Jackson, et al.
Published: (2025)
by: Harmon, Jackson, et al.
Published: (2025)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
by: Sinha, Shiven, et al.
Published: (2024)
by: Sinha, Shiven, et al.
Published: (2024)
LLM generation novelty through the lens of semantic similarity
by: Davydov, Philipp, et al.
Published: (2025)
by: Davydov, Philipp, et al.
Published: (2025)
Prompt-Level Reward Specifications for Open-Ended Post-Training
by: Weng, Zijun, et al.
Published: (2026)
by: Weng, Zijun, et al.
Published: (2026)
How Long Is a Piece of String? A Brief Empirical Analysis of Tokenizers
by: Roberts, Jonathan, et al.
Published: (2026)
by: Roberts, Jonathan, et al.
Published: (2026)
Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks?
by: Roberts, Jonathan, et al.
Published: (2024)
by: Roberts, Jonathan, et al.
Published: (2024)
Are We Done with Object-Centric Learning?
by: Rubinstein, Alexander, et al.
Published: (2025)
by: Rubinstein, Alexander, et al.
Published: (2025)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization
by: Banerjee, Adhiraj, et al.
Published: (2026)
by: Banerjee, Adhiraj, et al.
Published: (2026)
OpenEP: Open-Ended Future Event Prediction
by: Guan, Yong, et al.
Published: (2024)
by: Guan, Yong, et al.
Published: (2024)
OpenSIR: Open-Ended Self-Improving Reasoner
by: Kwan, Wai-Chung, et al.
Published: (2025)
by: Kwan, Wai-Chung, et al.
Published: (2025)
ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
by: Lange, Robert Tjarko, et al.
Published: (2025)
by: Lange, Robert Tjarko, et al.
Published: (2025)
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
by: Li, Hengzhi, et al.
Published: (2025)
by: Li, Hengzhi, et al.
Published: (2025)
Great Models Think Alike and this Undermines AI Oversight
by: Goel, Shashwat, et al.
Published: (2025)
by: Goel, Shashwat, et al.
Published: (2025)
OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation
by: Zhang, Hanning, et al.
Published: (2025)
by: Zhang, Hanning, et al.
Published: (2025)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
by: Xie, Tianbao, et al.
Published: (2024)
by: Xie, Tianbao, et al.
Published: (2024)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
by: Demchak, Nathaniel, et al.
Published: (2024)
by: Demchak, Nathaniel, et al.
Published: (2024)
Decoding Decoded: Understanding Hyperparameter Effects in Open-Ended Text Generation
by: Arias, Esteban Garces, et al.
Published: (2024)
by: Arias, Esteban Garces, et al.
Published: (2024)
Open-Ended Wargames with Large Language Models
by: Hogan, Daniel P., et al.
Published: (2024)
by: Hogan, Daniel P., et al.
Published: (2024)
REAL Sampling: Boosting Factuality and Diversity of Open-Ended Generation via Asymptotic Entropy
by: Chang, Haw-Shiuan, et al.
Published: (2024)
by: Chang, Haw-Shiuan, et al.
Published: (2024)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
A Benchmark for Open-Domain Numerical Fact-Checking Enhanced by Claim Decomposition
by: Venktesh, V, et al.
Published: (2025)
by: Venktesh, V, et al.
Published: (2025)
GUARD: Glocal Uncertainty-Aware Robust Decoding for Effective and Efficient Open-Ended Text Generation
by: Ding, Yuanhao, et al.
Published: (2025)
by: Ding, Yuanhao, et al.
Published: (2025)
Semantic Agreement Enables Efficient Open-Ended LLM Cascades
by: Soiffer, Duncan, et al.
Published: (2025)
by: Soiffer, Duncan, et al.
Published: (2025)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
by: Jiang, Liwei, et al.
Published: (2025)
by: Jiang, Liwei, et al.
Published: (2025)
MASIVE: Open-Ended Affective State Identification in English and Spanish
by: Deas, Nicholas, et al.
Published: (2024)
by: Deas, Nicholas, et al.
Published: (2024)
Improving Open-Ended Text Generation via Adaptive Decoding
by: Zhu, Wenhong, et al.
Published: (2024)
by: Zhu, Wenhong, et al.
Published: (2024)
Bias Association Discovery Framework for Open-Ended LLM Generations
by: Pan, Jinhao, et al.
Published: (2025)
by: Pan, Jinhao, et al.
Published: (2025)
Similar Items
-
How to Merge Your Multimodal Models Over Time?
by: Dziadzio, Sebastian, et al.
Published: (2024) -
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
by: Udandarao, Vishaal, et al.
Published: (2024) -
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025) -
A Practitioner's Guide to Continual Multimodal Pretraining
by: Roth, Karsten, et al.
Published: (2024) -
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024)