KidsArtBench: Multi-Dimensional Children's Art Evaluation with Attribute-Aware MLLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Ye, Mingrui, Zheng, Chanjin, Yu, Zengyi, Xiang, Chenyu, Zhao, Zhixue, Yuan, Zheng, Yannakoudakis, Helen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ArtMentor: AI-Assisted Evaluation of Artworks to Explore Multimodal Large Language Models Capabilities
por: Zheng, Chanjin, et al.
Publicado: (2025)
por: Zheng, Chanjin, et al.
Publicado: (2025)
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
por: Ye, Hengwei, et al.
Publicado: (2026)
por: Ye, Hengwei, et al.
Publicado: (2026)
Incorporating Attribution Importance for Improving Faithfulness Metrics
por: Zhao, Zhixue, et al.
Publicado: (2023)
por: Zhao, Zhixue, et al.
Publicado: (2023)
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
por: Zhao, Zhixue, et al.
Publicado: (2024)
por: Zhao, Zhixue, et al.
Publicado: (2024)
AttributionBench: How Hard is Automatic Attribution Evaluation?
por: Li, Yifei, et al.
Publicado: (2024)
por: Li, Yifei, et al.
Publicado: (2024)
Agentic Problem Frames: A Systematic Approach to Engineering Reliable Domain Agents
por: Park, Chanjin
Publicado: (2026)
por: Park, Chanjin
Publicado: (2026)
ArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling
por: Zhu, William Yicheng, et al.
Publicado: (2024)
por: Zhu, William Yicheng, et al.
Publicado: (2024)
ArtBrain: An Explainable end-to-end Toolkit for Classification and Attribution of AI-Generated Art and Style
por: Silva, Ravidu Suien Rammuni, et al.
Publicado: (2024)
por: Silva, Ravidu Suien Rammuni, et al.
Publicado: (2024)
MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
por: Guo, Shengyu, et al.
Publicado: (2026)
por: Guo, Shengyu, et al.
Publicado: (2026)
Strat-LLM: Stratified Strategy Alignment for LLM-based Stock Trading with Real-time Multi-Source Signals
por: Huang, Wenliang, et al.
Publicado: (2026)
por: Huang, Wenliang, et al.
Publicado: (2026)
CALM: A Causal Analysis Language Model for Tabular Data in Complex Systems with Local Scores, Conditional Independence Tests, and Relation Attributes
por: Fan, Zhenjiang, et al.
Publicado: (2025)
por: Fan, Zhenjiang, et al.
Publicado: (2025)
MileBench: Benchmarking MLLMs in Long Context
por: Song, Dingjie, et al.
Publicado: (2024)
por: Song, Dingjie, et al.
Publicado: (2024)
TeamLLM: A Human-Like Team-Oriented Collaboration Framework for Multi-Step Contextualized Tasks
por: Wang, Xiangyu, et al.
Publicado: (2026)
por: Wang, Xiangyu, et al.
Publicado: (2026)
PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms
por: Wang, Wei, et al.
Publicado: (2026)
por: Wang, Wei, et al.
Publicado: (2026)
A Functional Perspective on Knowledge Distillation in Neural Networks
por: Mason-Williams, Israel, et al.
Publicado: (2025)
por: Mason-Williams, Israel, et al.
Publicado: (2025)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
por: Jiang, Fengqing, et al.
Publicado: (2024)
por: Jiang, Fengqing, et al.
Publicado: (2024)
The Art of Tool Interface Design
por: Wu, Yunnan, et al.
Publicado: (2025)
por: Wu, Yunnan, et al.
Publicado: (2025)
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
por: Li, Caorui, et al.
Publicado: (2025)
por: Li, Caorui, et al.
Publicado: (2025)
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
por: Liang, Shanchao, et al.
Publicado: (2025)
por: Liang, Shanchao, et al.
Publicado: (2025)
A Function-Centric Perspective on Flat and Sharp Minima
por: Mason-Williams, Israel, et al.
Publicado: (2025)
por: Mason-Williams, Israel, et al.
Publicado: (2025)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
por: Lin, Junming, et al.
Publicado: (2024)
por: Lin, Junming, et al.
Publicado: (2024)
CSR-Bench: A Benchmark for Evaluating the Cross-modal Safety and Reliability of MLLMs
por: Liu, Yuxuan, et al.
Publicado: (2026)
por: Liu, Yuxuan, et al.
Publicado: (2026)
A Forced-Choice Neural Cognitive Diagnostic Model of Personality Testing
por: Li, Xiaoyu, et al.
Publicado: (2025)
por: Li, Xiaoyu, et al.
Publicado: (2025)
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension
por: Zhang, Zhi, et al.
Publicado: (2023)
por: Zhang, Zhi, et al.
Publicado: (2023)
TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning
por: Li, Yize, et al.
Publicado: (2026)
por: Li, Yize, et al.
Publicado: (2026)
The Pleasure Principle: Where is it in Kids' Art Books
por: Wilton, Shirley M.
Publicado: (1977)
por: Wilton, Shirley M.
Publicado: (1977)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
por: Qiu, Yansheng, et al.
Publicado: (2025)
por: Qiu, Yansheng, et al.
Publicado: (2025)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
por: Xu, Pengju, et al.
Publicado: (2025)
por: Xu, Pengju, et al.
Publicado: (2025)
Wired Perspectives: Multi-View Wire Art Embraces Generative AI
por: Qu, Zhiyu, et al.
Publicado: (2023)
por: Qu, Zhiyu, et al.
Publicado: (2023)
Position: State-of-the-Art Claims Require State-of-the-Art Evidence
por: Oh, YongKyung
Publicado: (2026)
por: Oh, YongKyung
Publicado: (2026)
A Scoping Review of Energy-Efficient Driving Behaviors and Applied State-of-the-Art AI Methods
por: Ma, Zhipeng, et al.
Publicado: (2024)
por: Ma, Zhipeng, et al.
Publicado: (2024)
Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment
por: Hong, Jiaying, et al.
Publicado: (2025)
por: Hong, Jiaying, et al.
Publicado: (2025)
VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage
por: Alfarano, A., et al.
Publicado: (2025)
por: Alfarano, A., et al.
Publicado: (2025)
Holistic Evaluation of State-of-the-Art LLMs for Code Generation
por: Zhang, Le, et al.
Publicado: (2025)
por: Zhang, Le, et al.
Publicado: (2025)
ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams
por: Xu, Qiang, et al.
Publicado: (2026)
por: Xu, Qiang, et al.
Publicado: (2026)
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
por: Kirch, Nathalie, et al.
Publicado: (2024)
por: Kirch, Nathalie, et al.
Publicado: (2024)
M-ArtAgent: Evidence-Based Multimodal Agent for Implicit Art Influence Discovery
por: Liu, Hanyi, et al.
Publicado: (2026)
por: Liu, Hanyi, et al.
Publicado: (2026)
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
por: Jing, Pengfei, et al.
Publicado: (2024)
por: Jing, Pengfei, et al.
Publicado: (2024)
Differentiating Student Feedbacks for Knowledge Tracing
por: Cui, Jiajun, et al.
Publicado: (2022)
por: Cui, Jiajun, et al.
Publicado: (2022)
Quantifying Compositionality of Classic and State-of-the-Art Embeddings
por: Guo, Zhijin, et al.
Publicado: (2025)
por: Guo, Zhijin, et al.
Publicado: (2025)
Ejemplares similares
-
ArtMentor: AI-Assisted Evaluation of Artworks to Explore Multimodal Large Language Models Capabilities
por: Zheng, Chanjin, et al.
Publicado: (2025) -
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
por: Ye, Hengwei, et al.
Publicado: (2026) -
Incorporating Attribution Importance for Improving Faithfulness Metrics
por: Zhao, Zhixue, et al.
Publicado: (2023) -
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
por: Zhao, Zhixue, et al.
Publicado: (2024) -
AttributionBench: How Hard is Automatic Attribution Evaluation?
por: Li, Yifei, et al.
Publicado: (2024)