Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
Fuente:
arXiv
Guardado en:
| Autores principales: | Padlewski, Piotr, Bain, Max, Henderson, Matthew, Zhu, Zhongkai, Relan, Nishant, Pham, Hai, Ong, Donovan, Aleksiev, Kaloyan, Ormazabal, Aitor, Phua, Samuel, Yeo, Ethan, Lamprecht, Eugenie, Liu, Qi, Wang, Yuqi, Chen, Eric, Fu, Deyu, Li, Lei, Zheng, Che, d'Autume, Cyprien de Masson, Yogatama, Dani, Artetxe, Mikel, Tay, Yi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models
por: Reka Team, et al.
Publicado: (2024)
por: Reka Team, et al.
Publicado: (2024)
Latxa: An Open Language Model and Evaluation Suite for Basque
por: Etxaniz, Julen, et al.
Publicado: (2024)
por: Etxaniz, Julen, et al.
Publicado: (2024)
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
por: Chiang, Ting-Rui, et al.
Publicado: (2025)
por: Chiang, Ting-Rui, et al.
Publicado: (2025)
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
por: Chiang, Ting-Rui, et al.
Publicado: (2024)
por: Chiang, Ting-Rui, et al.
Publicado: (2024)
FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale
por: Lee, Isabelle, et al.
Publicado: (2025)
por: Lee, Isabelle, et al.
Publicado: (2025)
Improving the Efficiency of Visually Augmented Language Models
por: Ontalvilla, Paula, et al.
Publicado: (2024)
por: Ontalvilla, Paula, et al.
Publicado: (2024)
Multimodal LLMs Do Not Compose Skills Optimally Across Modalities
por: Ontalvilla, Paula, et al.
Publicado: (2025)
por: Ontalvilla, Paula, et al.
Publicado: (2025)
BertaQA: How Much Do Language Models Know About Local Culture?
por: Etxaniz, Julen, et al.
Publicado: (2024)
por: Etxaniz, Julen, et al.
Publicado: (2024)
Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
por: Elhady, Ahmed, et al.
Publicado: (2025)
por: Elhady, Ahmed, et al.
Publicado: (2025)
WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More Challenging
por: Elhady, Ahmed, et al.
Publicado: (2025)
por: Elhady, Ahmed, et al.
Publicado: (2025)
Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models
por: Elhady, Ahmed, et al.
Publicado: (2026)
por: Elhady, Ahmed, et al.
Publicado: (2026)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
por: Liu, Ollie, et al.
Publicado: (2024)
por: Liu, Ollie, et al.
Publicado: (2024)
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
por: Lee, Isabelle, et al.
Publicado: (2024)
por: Lee, Isabelle, et al.
Publicado: (2024)
LocateBench: Evaluating the Locating Ability of Vision Language Models
por: Chiang, Ting-Rui, et al.
Publicado: (2024)
por: Chiang, Ting-Rui, et al.
Publicado: (2024)
The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities
por: Wu, Zhaofeng, et al.
Publicado: (2024)
por: Wu, Zhaofeng, et al.
Publicado: (2024)
Interpretable Diffusion via Information Decomposition
por: Kong, Xianghao, et al.
Publicado: (2023)
por: Kong, Xianghao, et al.
Publicado: (2023)
Bronchial Dieulafoy: Diagnostic Utility of Endobronchial Ultrasound
por: Chee Kiang Phua, et al.
Publicado: (2025)
por: Chee Kiang Phua, et al.
Publicado: (2025)
ENDOSCOPIC ENDONASAL TRANSSPHENOIDAL APPROACH IN THE TREATMENT OF OPTIC CHIASM GLIOMAS – A REVIEW OF THE LITERATURE
por: Bechev K., et al.
Publicado: (2026)
por: Bechev K., et al.
Publicado: (2026)
Observation of amplitude-driven nonreciprocity for energy guiding
por: Padlewski, Mathieu, et al.
Publicado: (2024)
por: Padlewski, Mathieu, et al.
Publicado: (2024)
Can We Test Consciousness Theories on AI? Ablations, Markers, and Robustness
por: Phua, Yin Jun
Publicado: (2025)
por: Phua, Yin Jun
Publicado: (2025)
Enantioselective Anion‐Binding‐Catalyzed Nucleophilic Addition to 4‐Quinolones
por: Martin Aleksiev, et al.
Publicado: (2025)
por: Martin Aleksiev, et al.
Publicado: (2025)
PCN: A Deep Learning Approach to Jet Tagging Utilizing Novel Graph Construction Methods and Chebyshev Graph Convolutions
por: Semlani, Yash, et al.
Publicado: (2023)
por: Semlani, Yash, et al.
Publicado: (2023)
MathRobust-LV: Evaluation of Large Language Models' Robustness to Linguistic Variations in Mathematical Reasoning
por: Kirtane, Neeraja, et al.
Publicado: (2025)
por: Kirtane, Neeraja, et al.
Publicado: (2025)
An application of random plane slicing to counting $\mathbb{F}_q$-points on hypersurfaces
por: Slavov, Kaloyan
Publicado: (2017)
por: Slavov, Kaloyan
Publicado: (2017)
The moduli space of hypersurfaces whose singular locus has high dimension
por: Slavov, Kaloyan
Publicado: (2012)
por: Slavov, Kaloyan
Publicado: (2012)
Factorization type probabilities of polynomials with prescribed coefficients over a finite field
por: Slavov, Kaloyan
Publicado: (2019)
por: Slavov, Kaloyan
Publicado: (2019)
The Hilbert polynomial of a symbolic square
por: Slavov, Kaloyan
Publicado: (2012)
por: Slavov, Kaloyan
Publicado: (2012)
An algebraic geometry version of the Kakeya problem
por: Slavov, Kaloyan
Publicado: (2014)
por: Slavov, Kaloyan
Publicado: (2014)
Main-sequence exoplanet systems: tidal evolution
por: Penev, Kaloyan
Publicado: (2024)
por: Penev, Kaloyan
Publicado: (2024)
Improved Lang--Weil bounds for a geometrically irreducible hypersurface over a finite field
por: Slavov, Kaloyan
Publicado: (2021)
por: Slavov, Kaloyan
Publicado: (2021)
Square values of several polynomials over a finite field
por: Slavov, Kaloyan
Publicado: (2024)
por: Slavov, Kaloyan
Publicado: (2024)
Gender-specific Machine Translation with Large Language Models
por: Sánchez, Eduardo, et al.
Publicado: (2023)
por: Sánchez, Eduardo, et al.
Publicado: (2023)
Decision trees for estimating osteological sex from the skull using an expanded suite of morphological traits
por: Morgan J. Ferrell, et al.
Publicado: (2025)
por: Morgan J. Ferrell, et al.
Publicado: (2025)
SAVIANO, Roberto (2008). Gomorra. Infiltrado no Império Económico da Máfia Napolitana, Caderno, 2008, Lisboa, 3ª. Ed.: 351 pp
por: René Luis Tapia Ormazábal
Publicado: (2010)
por: René Luis Tapia Ormazábal
Publicado: (2010)
On Retrieval Augmentation and the Limitations of Language Model Training
por: Chiang, Ting-Rui, et al.
Publicado: (2023)
por: Chiang, Ting-Rui, et al.
Publicado: (2023)
Rigorous Interpretation Is a Form of Evaluation
por: Lee, Isabelle, et al.
Publicado: (2026)
por: Lee, Isabelle, et al.
Publicado: (2026)
Position: Vibe Coding Needs Vibe Reasoning: Improving Vibe Coding with Formal Verification
por: Mitchell, Jacqueline, et al.
Publicado: (2025)
por: Mitchell, Jacqueline, et al.
Publicado: (2025)
Evaluating Large Language Models for Fair and Reliable Organ Allocation
por: Kim, Brian Hyeongseok, et al.
Publicado: (2025)
por: Kim, Brian Hyeongseok, et al.
Publicado: (2025)
VibeContract: The Missing Quality Assurance Piece in Vibe Coding
por: Wang, Song
Publicado: (2026)
por: Wang, Song
Publicado: (2026)
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
por: Bansal, Srijan, et al.
Publicado: (2026)
por: Bansal, Srijan, et al.
Publicado: (2026)
Ejemplares similares
-
Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models
por: Reka Team, et al.
Publicado: (2024) -
Latxa: An Open Language Model and Evaluation Suite for Basque
por: Etxaniz, Julen, et al.
Publicado: (2024) -
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
por: Chiang, Ting-Rui, et al.
Publicado: (2025) -
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
por: Chiang, Ting-Rui, et al.
Publicado: (2024) -
FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale
por: Lee, Isabelle, et al.
Publicado: (2025)