Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Yibo, Wang, Shen, Huo, Jiahao, Ye, Jingheng, Chu, Zhendong, Hu, Xuming, Yu, Philip S., Gomes, Carla, Selman, Bart, Wen, Qingsong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
von: Yan, Yibo, et al.
Veröffentlicht: (2024)
von: Yan, Yibo, et al.
Veröffentlicht: (2024)
MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model
von: Huo, Jiahao, et al.
Veröffentlicht: (2024)
von: Huo, Jiahao, et al.
Veröffentlicht: (2024)
LLM Agents for Education: Advances and Applications
von: Chu, Zhendong, et al.
Veröffentlicht: (2025)
von: Chu, Zhendong, et al.
Veröffentlicht: (2025)
Position: LLMs Can be Good Tutors in English Education
von: Ye, Jingheng, et al.
Veröffentlicht: (2025)
von: Ye, Jingheng, et al.
Veröffentlicht: (2025)
MINER: Mining the Underlying Pattern of Modality-Specific Neurons in Multimodal Large Language Models
von: Huang, Kaichen, et al.
Veröffentlicht: (2024)
von: Huang, Kaichen, et al.
Veröffentlicht: (2024)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
von: Huo, Jiahao, et al.
Veröffentlicht: (2025)
von: Huo, Jiahao, et al.
Veröffentlicht: (2025)
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
von: Yan, Yibo, et al.
Veröffentlicht: (2024)
von: Yan, Yibo, et al.
Veröffentlicht: (2024)
Multimodal AI Teacher: Integrating Edge Computing and Reasoning Models for Enhanced Student Error Analysis
von: Tianlong Xu, et al.
Veröffentlicht: (2025)
von: Tianlong Xu, et al.
Veröffentlicht: (2025)
AI-Driven Virtual Teacher for Enhanced Educational Efficiency: Leveraging Large Pretrain Models for Autonomous Error Analysis and Correction
von: Xu, Tianlong, et al.
Veröffentlicht: (2024)
von: Xu, Tianlong, et al.
Veröffentlicht: (2024)
Agentic Neurosymbolic Collaboration for Mathematical Discovery: A Case Study in Combinatorial Design
von: Xia, Hai, et al.
Veröffentlicht: (2026)
von: Xia, Hai, et al.
Veröffentlicht: (2026)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025)
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025)
UniEDU: A Unified Language and Vision Assistant for Education Applications
von: Chu, Zhendong, et al.
Veröffentlicht: (2025)
von: Chu, Zhendong, et al.
Veröffentlicht: (2025)
Corrections Meet Explanations: A Unified Framework for Explainable Grammatical Error Correction
von: Ye, Jingheng, et al.
Veröffentlicht: (2025)
von: Ye, Jingheng, et al.
Veröffentlicht: (2025)
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
von: Huo, Jiahao, et al.
Veröffentlicht: (2026)
von: Huo, Jiahao, et al.
Veröffentlicht: (2026)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis
von: Huang, Haoming, et al.
Veröffentlicht: (2025)
von: Huang, Haoming, et al.
Veröffentlicht: (2025)
ARM2: Adaptive Reasoning Model with Vision Understanding and Executable Code
von: Xie, Jian, et al.
Veröffentlicht: (2025)
von: Xie, Jian, et al.
Veröffentlicht: (2025)
SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation
von: Guo, Longteng, et al.
Veröffentlicht: (2026)
von: Guo, Longteng, et al.
Veröffentlicht: (2026)
EffiReason-Bench: A Unified Benchmark for Evaluating and Advancing Efficient Reasoning in Large Language Models
von: Huang, Junquan, et al.
Veröffentlicht: (2025)
von: Huang, Junquan, et al.
Veröffentlicht: (2025)
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey
von: Dang, Yunkai, et al.
Veröffentlicht: (2024)
von: Dang, Yunkai, et al.
Veröffentlicht: (2024)
CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding
von: Huo, Jiahao, et al.
Veröffentlicht: (2026)
von: Huo, Jiahao, et al.
Veröffentlicht: (2026)
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
von: Zhou, Guanyu, et al.
Veröffentlicht: (2024)
von: Zhou, Guanyu, et al.
Veröffentlicht: (2024)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
von: Huo, Jiahao, et al.
Veröffentlicht: (2025)
von: Huo, Jiahao, et al.
Veröffentlicht: (2025)
LocalEscaper: A Weakly-supervised Framework with Regional Reconstruction for Scalable Neural TSP Solvers
von: Wen, Junrui, et al.
Veröffentlicht: (2025)
von: Wen, Junrui, et al.
Veröffentlicht: (2025)
SciMDR: Advancing Scientific Multimodal Document Reasoning
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation
von: Cui, Zhiqing, et al.
Veröffentlicht: (2025)
von: Cui, Zhiqing, et al.
Veröffentlicht: (2025)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
von: Qin, Jialong, et al.
Veröffentlicht: (2025)
von: Qin, Jialong, et al.
Veröffentlicht: (2025)
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning
von: Chen, Junkai, et al.
Veröffentlicht: (2025)
von: Chen, Junkai, et al.
Veröffentlicht: (2025)
Mind Scramble: Unveiling Large Language Model Psychology Via Typoglycemia
von: Yu, Miao, et al.
Veröffentlicht: (2024)
von: Yu, Miao, et al.
Veröffentlicht: (2024)
Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
von: Dongfang, Zihao, et al.
Veröffentlicht: (2025)
von: Dongfang, Zihao, et al.
Veröffentlicht: (2025)
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
Reciprocal longitudinal relations between child routines and parenting
von: Saliha B. Selman, et al.
Veröffentlicht: (2025)
von: Saliha B. Selman, et al.
Veröffentlicht: (2025)
A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation
von: Wang, Liping, et al.
Veröffentlicht: (2026)
von: Wang, Liping, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
von: Yan, Yibo, et al.
Veröffentlicht: (2025) -
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
von: Su, Jiamin, et al.
Veröffentlicht: (2025) -
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
von: Yan, Yibo, et al.
Veröffentlicht: (2024) -
MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model
von: Huo, Jiahao, et al.
Veröffentlicht: (2024) -
LLM Agents for Education: Advances and Applications
von: Chu, Zhendong, et al.
Veröffentlicht: (2025)