MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models
Fuente:
arXiv
Saved in:
| Main Authors: | Yoshitake, Michiko, Suzuki, Yuta, Igarashi, Ryo, Ushiku, Yoshitaka, Nagato, Keisuke |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MaterialBENCH: Evaluating College-Level Materials Science Problem-Solving Abilities of Large Language Models
by: Yoshitake, Michiko, et al.
Published: (2024)
by: Yoshitake, Michiko, et al.
Published: (2024)
Bridging Text and Crystal Structures: Literature-driven Contrastive Learning for Materials Science
by: Suzuki, Yuta, et al.
Published: (2025)
by: Suzuki, Yuta, et al.
Published: (2025)
Crystalformer: Infinitely Connected Attention for Periodic Structure Encoding
by: Taniai, Tatsunori, et al.
Published: (2024)
by: Taniai, Tatsunori, et al.
Published: (2024)
Rethinking Symbolic Regression Datasets and Benchmarks for Scientific Discovery
by: Matsubara, Yoshitomo, et al.
Published: (2022)
by: Matsubara, Yoshitomo, et al.
Published: (2022)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
CrystalFramer: Rethinking the Role of Frames for SE(3)-Invariant Crystal Structure Modeling
by: Ito, Yusei, et al.
Published: (2025)
by: Ito, Yusei, et al.
Published: (2025)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
by: Li, Kunning, et al.
Published: (2025)
by: Li, Kunning, et al.
Published: (2025)
A Transformer Model for Symbolic Regression towards Scientific Discovery
by: Lalande, Florian, et al.
Published: (2023)
by: Lalande, Florian, et al.
Published: (2023)
Framework for evaluating code generation ability of large language models
by: Sangyeop Yeo, et al.
Published: (2024)
by: Sangyeop Yeo, et al.
Published: (2024)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
by: Jassim, Serwan, et al.
Published: (2023)
by: Jassim, Serwan, et al.
Published: (2023)
Is your multimodal large language model a good science tutor?
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
A dataset and benchmark for hospital course summarization with adapted large language models
by: Aali, Asad, et al.
Published: (2024)
by: Aali, Asad, et al.
Published: (2024)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
by: Wen, Zichen, et al.
Published: (2024)
by: Wen, Zichen, et al.
Published: (2024)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
by: Matos, João, et al.
Published: (2024)
by: Matos, João, et al.
Published: (2024)
LongTail-Swap: benchmarking language models' abilities on rare words
by: Algayres, Robin, et al.
Published: (2025)
by: Algayres, Robin, et al.
Published: (2025)
DevBench: A multimodal developmental benchmark for language learning
by: Tan, Alvin Wei Ming, et al.
Published: (2024)
by: Tan, Alvin Wei Ming, et al.
Published: (2024)
Type-Based Verification of Connectivity Constraints in Lattice Surgery
by: Wakizaka, Ryo, et al.
Published: (2024)
by: Wakizaka, Ryo, et al.
Published: (2024)
A comprehensive multimodal dataset and benchmark for ulcerative colitis scoring in endoscopy
by: Ghatwary, Noha, et al.
Published: (2026)
by: Ghatwary, Noha, et al.
Published: (2026)
Embodiment in multimodal large language models
by: Kadambi, Akila, et al.
Published: (2025)
by: Kadambi, Akila, et al.
Published: (2025)
Qwen-BIM: developing large language model for BIM-based design with domain-specific benchmark and dataset
by: Lin, Jia-Rui, et al.
Published: (2026)
by: Lin, Jia-Rui, et al.
Published: (2026)
FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models
by: Shamsfard, Mehrnoush, et al.
Published: (2025)
by: Shamsfard, Mehrnoush, et al.
Published: (2025)
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Visual cognition in multimodal large language models
by: Buschoff, Luca M. Schulze, et al.
Published: (2023)
by: Buschoff, Luca M. Schulze, et al.
Published: (2023)
Matching prior pairs connecting Maximum A Posteriori estimation and posterior expectation
by: Okudo, Michiko, et al.
Published: (2023)
by: Okudo, Michiko, et al.
Published: (2023)
Notificação espontânea de erros de medicação em hospital universitário pediátrico
by: Michiko Suzuki Yamamoto
Published: (2011)
by: Michiko Suzuki Yamamoto
Published: (2011)
A figure-of-merit-based framework to evaluate photovoltaic materials
by: Crovetto, Andrea
Published: (2024)
by: Crovetto, Andrea
Published: (2024)
SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters
by: Tanaka, Shohei, et al.
Published: (2025)
by: Tanaka, Shohei, et al.
Published: (2025)
SciPostLayout: A Dataset for Layout Analysis and Layout Generation of Scientific Posters
by: Tanaka, Shohei, et al.
Published: (2024)
by: Tanaka, Shohei, et al.
Published: (2024)
Fístula axilo-cava para hemodiálise: relato de caso
by: Yosio Nagato
Published: (2009)
by: Yosio Nagato
Published: (2009)
A large dataset curation and benchmark for drug target interaction
by: Golts, Alex, et al.
Published: (2024)
by: Golts, Alex, et al.
Published: (2024)
AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
by: Schmidgall, Samuel, et al.
Published: (2024)
by: Schmidgall, Samuel, et al.
Published: (2024)
A large-scale image-text dataset benchmark for farmland segmentation
by: Tao, Chao, et al.
Published: (2025)
by: Tao, Chao, et al.
Published: (2025)
Reinforcement-Learning-Designed Field-Free Sub-Nanosecond Spin-Orbit-Torque Switching
by: Igarashi, Yuta, et al.
Published: (2025)
by: Igarashi, Yuta, et al.
Published: (2025)
Creativity Benchmark: A benchmark for marketing creativity for large language models
by: Bhat, Ninad, et al.
Published: (2025)
by: Bhat, Ninad, et al.
Published: (2025)
Comprehensive benchmarking of large language models for RNA secondary structure prediction
by: Zablocki, L. I., et al.
Published: (2024)
by: Zablocki, L. I., et al.
Published: (2024)
Protecting multimodal large language models against misleading visualizations
by: Tonglet, Jonathan, et al.
Published: (2025)
by: Tonglet, Jonathan, et al.
Published: (2025)
Probing the limitations of multimodal language models for chemistry and materials research
by: Alampara, Nawaf, et al.
Published: (2024)
by: Alampara, Nawaf, et al.
Published: (2024)
CerraData-4MM: A multimodal benchmark dataset on Cerrado for land use and land cover classification
by: Miranda, Mateus de Souza, et al.
Published: (2025)
by: Miranda, Mateus de Souza, et al.
Published: (2025)
Evaluating point-light biological motion in multimodal large language models
by: Kadambi, Akila, et al.
Published: (2025)
by: Kadambi, Akila, et al.
Published: (2025)
Design, Control, and Motion Strategy of TRADY: Tilted-Rotor-Equipped Aerial Robot With Autonomous In-Flight Assembly and Disassembly Ability
by: Sugihara, Junichiro, et al.
Published: (2023)
by: Sugihara, Junichiro, et al.
Published: (2023)
Similar Items
-
MaterialBENCH: Evaluating College-Level Materials Science Problem-Solving Abilities of Large Language Models
by: Yoshitake, Michiko, et al.
Published: (2024) -
Bridging Text and Crystal Structures: Literature-driven Contrastive Learning for Materials Science
by: Suzuki, Yuta, et al.
Published: (2025) -
Crystalformer: Infinitely Connected Attention for Periodic Structure Encoding
by: Taniai, Tatsunori, et al.
Published: (2024) -
Rethinking Symbolic Regression Datasets and Benchmarks for Scientific Discovery
by: Matsubara, Yoshitomo, et al.
Published: (2022) -
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)