DevBench: A multimodal developmental benchmark for language learning
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Alvin Wei Ming, Yu, Sunny, Long, Bria, Ma, Wanjing Anya, Murray, Tonya, Silverman, Rebecca D., Yeatman, Jason D., Frank, Michael C. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
The Virtuous Cycle between Education and Neuroscience
by: Jason D. Yeatman, et al.
Published: (2025)
by: Jason D. Yeatman, et al.
Published: (2025)
Automatic Generation of Inference Making Questions for Reading Comprehension Assessments
by: Ma, Wanjing Anya, et al.
Published: (2025)
by: Ma, Wanjing Anya, et al.
Published: (2025)
Characterizing the visual representation of objects from the child's view
by: Yang, Jane, et al.
Published: (2026)
by: Yang, Jane, et al.
Published: (2026)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
Context informs pragmatic interpretation in vision-language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
by: Jassim, Serwan, et al.
Published: (2023)
by: Jassim, Serwan, et al.
Published: (2023)
MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
by: Yang, Junjie, et al.
Published: (2025)
by: Yang, Junjie, et al.
Published: (2025)
GraphBench: Next-generation graph learning benchmarking
by: Stoll, Timo, et al.
Published: (2025)
by: Stoll, Timo, et al.
Published: (2025)
La caballerosidad como mediador entre el autoritarismo y los roles de género
by: Paula Bria
Published: (2020)
by: Paula Bria
Published: (2020)
iSafetyBench: A video-language benchmark for safety in industrial environment
by: Abdullah, Raiyaan, et al.
Published: (2025)
by: Abdullah, Raiyaan, et al.
Published: (2025)
Developing Vocabulary and Oral Language in Young Children. The Essential Library of PreK-2 Literacy
by: Silverman, Rebecca D., et al.
Published: (2014)
by: Silverman, Rebecca D., et al.
Published: (2014)
Navigating Data Corruption in Machine Learning: Balancing Quality, Quantity, and Imputation Strategies
by: Liu, Qi, et al.
Published: (2024)
by: Liu, Qi, et al.
Published: (2024)
The Epochal Sawtooth Phenomenon: Unveiling Training Loss Oscillations in Adam and Other Optimizers
by: Liu, Qi, et al.
Published: (2024)
by: Liu, Qi, et al.
Published: (2024)
Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models
by: He-Yueya, Joy, et al.
Published: (2024)
by: He-Yueya, Joy, et al.
Published: (2024)
On the robustness of multimodal language model towards distractions
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
by: Lu, Pengrui, et al.
Published: (2026)
by: Lu, Pengrui, et al.
Published: (2026)
Reseña de "Copy-left. Manual de uso" de VV.AA.
by: Marc Bría Ramírez
Published: (2007)
by: Marc Bría Ramírez
Published: (2007)
SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
by: Rashid, Muhammad Shihab, et al.
Published: (2025)
by: Rashid, Muhammad Shihab, et al.
Published: (2025)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
The role of translation equivalents in bilingual word learning
by: Alvin W. M. Tan, et al.
Published: (2024)
by: Alvin W. M. Tan, et al.
Published: (2024)
ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows
by: Padigela, Harshith, et al.
Published: (2025)
by: Padigela, Harshith, et al.
Published: (2025)
MobileDev-Bench: A Benchmark for Issue Resolution in Mobile Application Development
by: Fakorede, Moshood A., et al.
Published: (2026)
by: Fakorede, Moshood A., et al.
Published: (2026)
A small lens on timescales and multimodality in classroom language learning emotions
by: Richard J. Sampson
Published: (2024)
by: Richard J. Sampson
Published: (2024)
A Scoping Review of Neighborhood Effects and Health Among Children and Adolescents: Measurement and Design Characteristics
by: Bria Gresham, et al.
Published: (2025)
by: Bria Gresham, et al.
Published: (2025)
REVISTA DE LIBROS
by: Fernando Javier Guida Bria
Published: (2019)
by: Fernando Javier Guida Bria
Published: (2019)
Is your multimodal large language model a good science tutor?
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
MotifBench: A standardized protein design benchmark for motif-scaffolding problems
by: Zheng, Zhuoqi, et al.
Published: (2025)
by: Zheng, Zhuoqi, et al.
Published: (2025)
Retrieval-augmented in-context learning for multimodal large language models in disease classification
by: Zhan, Zaifu, et al.
Published: (2025)
by: Zhan, Zaifu, et al.
Published: (2025)
Acceptability judgements in Colloquial Singaporean English
by: Alvin Wei Ming Tan, et al.
Published: (2026)
by: Alvin Wei Ming Tan, et al.
Published: (2026)
GATSim: Urban Mobility Simulation with Generative Agents
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
Effectiveness Evaluation of Signal Coordination Based on Spatially Sparse Trajectory Data
by: Yuxuan Sun, et al.
Published: (2025)
by: Yuxuan Sun, et al.
Published: (2025)
Comprehensive benchmarking of large language models for RNA secondary structure prediction
by: Zablocki, L. I., et al.
Published: (2024)
by: Zablocki, L. I., et al.
Published: (2024)
A quantum inspired predictor of Parkinsons disease built on a diverse, multimodal dataset
by: Vatsavai, Diya, et al.
Published: (2024)
by: Vatsavai, Diya, et al.
Published: (2024)
MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models
by: Yoshitake, Michiko, et al.
Published: (2026)
by: Yoshitake, Michiko, et al.
Published: (2026)
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
A comprehensive multimodal dataset and benchmark for ulcerative colitis scoring in endoscopy
by: Ghatwary, Noha, et al.
Published: (2026)
by: Ghatwary, Noha, et al.
Published: (2026)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
by: Saxena, Siddhant, et al.
Published: (2026)
by: Saxena, Siddhant, et al.
Published: (2026)
How predictable is language model benchmark performance?
by: Owen, David
Published: (2024)
by: Owen, David
Published: (2024)
Similar Items
-
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
by: Kumarappan, Adarsh, et al.
Published: (2026) -
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025) -
The Virtuous Cycle between Education and Neuroscience
by: Jason D. Yeatman, et al.
Published: (2025) -
Automatic Generation of Inference Making Questions for Reading Comprehension Assessments
by: Ma, Wanjing Anya, et al.
Published: (2025) -
Characterizing the visual representation of objects from the child's view
by: Yang, Jane, et al.
Published: (2026)