Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Yixuan, He, Peng, Liu, Honglu, Fan, Jinxuan, Ji, Yuyang, Li, Tingting, Chen, Tianlong, Xu, Kaidi, Liu, Feng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
di: Lai, Zhengzhao, et al.
Pubblicazione: (2025)
di: Lai, Zhengzhao, et al.
Pubblicazione: (2025)
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
di: Zhou, Honglu, et al.
Pubblicazione: (2025)
di: Zhou, Honglu, et al.
Pubblicazione: (2025)
TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs
di: Wang, Yunxiao, et al.
Pubblicazione: (2025)
di: Wang, Yunxiao, et al.
Pubblicazione: (2025)
SciEval: A Benchmark for Automatic Evaluation of K-12 Science Instructional Materials
di: Li, Zhaohui, et al.
Pubblicazione: (2026)
di: Li, Zhaohui, et al.
Pubblicazione: (2026)
IDSelect: A RL-Based Cost-Aware Selection Agent for Video-based Multi-Modal Person Recognition
di: Ji, Yuyang, et al.
Pubblicazione: (2026)
di: Ji, Yuyang, et al.
Pubblicazione: (2026)
DrawSim-PD: Simulating Student Science Drawings to Support NGSS-Aligned Teacher Diagnostic Reasoning
di: Chakma, Arijit, et al.
Pubblicazione: (2026)
di: Chakma, Arijit, et al.
Pubblicazione: (2026)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
di: Ma, David, et al.
Pubblicazione: (2025)
di: Ma, David, et al.
Pubblicazione: (2025)
Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior
di: Platt, Nolan, et al.
Pubblicazione: (2026)
di: Platt, Nolan, et al.
Pubblicazione: (2026)
Non-Colliding Biometric Identities for Digital Entities: Geometry, Capacity, and Million-Scale Virtual Identity Provisioning
di: Ji, Yuyang, et al.
Pubblicazione: (2026)
di: Ji, Yuyang, et al.
Pubblicazione: (2026)
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
di: Liang, Yijun, et al.
Pubblicazione: (2025)
di: Liang, Yijun, et al.
Pubblicazione: (2025)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
EMCompress: Video-LLMs with Endomorphic Multimodal Compression
di: Fan, Zheyu, et al.
Pubblicazione: (2025)
di: Fan, Zheyu, et al.
Pubblicazione: (2025)
Visual Reasoning Benchmark: Evaluating Multimodal LLMs on Classroom-Authentic Visual Problems from Primary Education
di: Huti, Mohamed, et al.
Pubblicazione: (2026)
di: Huti, Mohamed, et al.
Pubblicazione: (2026)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
di: Wang, Youze, et al.
Pubblicazione: (2025)
di: Wang, Youze, et al.
Pubblicazione: (2025)
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
di: Guo, Zhongbin, et al.
Pubblicazione: (2026)
di: Guo, Zhongbin, et al.
Pubblicazione: (2026)
From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching
di: Ji, Yuyang, et al.
Pubblicazione: (2026)
di: Ji, Yuyang, et al.
Pubblicazione: (2026)
LLMs can Find Mathematical Reasoning Mistakes by Pedagogical Chain-of-Thought
di: Jiang, Zhuoxuan, et al.
Pubblicazione: (2024)
di: Jiang, Zhuoxuan, et al.
Pubblicazione: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials
di: He, Peng, et al.
Pubblicazione: (2026)
di: He, Peng, et al.
Pubblicazione: (2026)
GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
di: Lei, Yiming, et al.
Pubblicazione: (2025)
di: Lei, Yiming, et al.
Pubblicazione: (2025)
Can GRPO Help LLMs Transcend Their Pretraining Origin?
di: Ni, Kangqi, et al.
Pubblicazione: (2025)
di: Ni, Kangqi, et al.
Pubblicazione: (2025)
Rewarding How Models Think Pedagogically: Integrating Pedagogical Reasoning and Thinking Rewards for LLMs in Education
di: Lee, Unggi, et al.
Pubblicazione: (2026)
di: Lee, Unggi, et al.
Pubblicazione: (2026)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
di: Cheng, Zixu, et al.
Pubblicazione: (2025)
di: Cheng, Zixu, et al.
Pubblicazione: (2025)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
di: Liu, Hongtao, et al.
Pubblicazione: (2025)
di: Liu, Hongtao, et al.
Pubblicazione: (2025)
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
di: Javaji, Shashidhar Reddy, et al.
Pubblicazione: (2025)
di: Javaji, Shashidhar Reddy, et al.
Pubblicazione: (2025)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
di: Liu, Chengzhi, et al.
Pubblicazione: (2025)
di: Liu, Chengzhi, et al.
Pubblicazione: (2025)
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
di: Hou, Wenjin, et al.
Pubblicazione: (2026)
di: Hou, Wenjin, et al.
Pubblicazione: (2026)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
di: Imam, Mohamed Fazli, et al.
Pubblicazione: (2025)
di: Imam, Mohamed Fazli, et al.
Pubblicazione: (2025)
Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
di: Guo, Xingang, et al.
Pubblicazione: (2025)
di: Guo, Xingang, et al.
Pubblicazione: (2025)
The Instructional Technology Support Center at MTSU: Integrating Technology into K-12 and University Classrooms.
di: Schmidt, Constance R.
Pubblicazione: (1996)
di: Schmidt, Constance R.
Pubblicazione: (1996)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
di: Jiang, Ruixiang, et al.
Pubblicazione: (2025)
di: Jiang, Ruixiang, et al.
Pubblicazione: (2025)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
di: Upadhyay, Ujjwal, et al.
Pubblicazione: (2025)
di: Upadhyay, Ujjwal, et al.
Pubblicazione: (2025)
Evaluating the Effectiveness of LLMs in Introductory Computer Science Education: A Semester-Long Field Study
di: Lyu, Wenhan, et al.
Pubblicazione: (2024)
di: Lyu, Wenhan, et al.
Pubblicazione: (2024)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
di: Zhu, Kejian, et al.
Pubblicazione: (2025)
di: Zhu, Kejian, et al.
Pubblicazione: (2025)
Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations
di: Liu, Zijie, et al.
Pubblicazione: (2025)
di: Liu, Zijie, et al.
Pubblicazione: (2025)
Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction
di: Shen, Yiqing, et al.
Pubblicazione: (2025)
di: Shen, Yiqing, et al.
Pubblicazione: (2025)
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
di: Gan, Ziliang, et al.
Pubblicazione: (2024)
di: Gan, Ziliang, et al.
Pubblicazione: (2024)
Multimodal LLMs See Sentiment
di: da Silva, Neemias B., et al.
Pubblicazione: (2025)
di: da Silva, Neemias B., et al.
Pubblicazione: (2025)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
di: Zhou, Xingcheng, et al.
Pubblicazione: (2026)
di: Zhou, Xingcheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
di: Wang, Wenxuan, et al.
Pubblicazione: (2025) -
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
di: Lai, Zhengzhao, et al.
Pubblicazione: (2025) -
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
di: Zhou, Honglu, et al.
Pubblicazione: (2025) -
TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs
di: Wang, Yunxiao, et al.
Pubblicazione: (2025) -
SciEval: A Benchmark for Automatic Evaluation of K-12 Science Instructional Materials
di: Li, Zhaohui, et al.
Pubblicazione: (2026)