TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ku, Max, Chong, Thomas, Leung, Jonathan, Shah, Krish, Yu, Alvin, Chen, Wenhu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
von: Ku, Max, et al.
Veröffentlicht: (2023)
von: Ku, Max, et al.
Veröffentlicht: (2023)
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
von: Maji, Arijit, et al.
Veröffentlicht: (2025)
von: Maji, Arijit, et al.
Veröffentlicht: (2025)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
Visual Set Program Synthesizer
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
Sentiment-enhanced Graph-based Sarcasm Explanation in Dialogue
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction
von: Ali, Mai, et al.
Veröffentlicht: (2025)
von: Ali, Mai, et al.
Veröffentlicht: (2025)
Towards Event-oriented Long Video Understanding
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
von: Ku, Max, et al.
Veröffentlicht: (2024)
von: Ku, Max, et al.
Veröffentlicht: (2024)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
Towards Robust Multimodal Sentiment Analysis with Incomplete Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences
von: Qian, Fan, et al.
Veröffentlicht: (2024)
von: Qian, Fan, et al.
Veröffentlicht: (2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Multimodal Model for Fake News Detection
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing
von: Wang, Bryan, et al.
Veröffentlicht: (2024)
von: Wang, Bryan, et al.
Veröffentlicht: (2024)
Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph Completion
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
ChemDFM-X: Towards Large Multimodal Model for Chemistry
von: Zhao, Zihan, et al.
Veröffentlicht: (2024)
von: Zhao, Zihan, et al.
Veröffentlicht: (2024)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
von: Shah, Siddhant Bikram, et al.
Veröffentlicht: (2024)
von: Shah, Siddhant Bikram, et al.
Veröffentlicht: (2024)
Shapley Value-based Contrastive Alignment for Multimodal Information Extraction
von: Luo, Wen, et al.
Veröffentlicht: (2024)
von: Luo, Wen, et al.
Veröffentlicht: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
Multimodal Sentiment Analysis Based on Causal Reasoning
von: Chen, Fuhai, et al.
Veröffentlicht: (2024)
von: Chen, Fuhai, et al.
Veröffentlicht: (2024)
VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations
von: Galarnyk, Michael, et al.
Veröffentlicht: (2025)
von: Galarnyk, Michael, et al.
Veröffentlicht: (2025)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
von: Wu, Zichen, et al.
Veröffentlicht: (2024)
von: Wu, Zichen, et al.
Veröffentlicht: (2024)
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
von: Ma, Ziyang, et al.
Veröffentlicht: (2026)
von: Ma, Ziyang, et al.
Veröffentlicht: (2026)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
von: Liu, Peipei, et al.
Veröffentlicht: (2023)
von: Liu, Peipei, et al.
Veröffentlicht: (2023)
Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck
von: Cui, Shiyao, et al.
Veröffentlicht: (2023)
von: Cui, Shiyao, et al.
Veröffentlicht: (2023)
Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video Ads
von: Zhang, Kunpeng, et al.
Veröffentlicht: (2026)
von: Zhang, Kunpeng, et al.
Veröffentlicht: (2026)
Towards Understanding Camera Motions in Any Video
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2025)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2025)
Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
von: Chen, Xiaolin, et al.
Veröffentlicht: (2025)
von: Chen, Xiaolin, et al.
Veröffentlicht: (2025)
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
von: Wang, Bing, et al.
Veröffentlicht: (2025)
von: Wang, Bing, et al.
Veröffentlicht: (2025)
TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis
von: Quan, Weize, et al.
Veröffentlicht: (2024)
von: Quan, Weize, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
von: Ku, Max, et al.
Veröffentlicht: (2023) -
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
von: Maji, Arijit, et al.
Veröffentlicht: (2025) -
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
von: Liang, Hao, et al.
Veröffentlicht: (2024) -
Visual Set Program Synthesizer
von: Cheng, Zehua, et al.
Veröffentlicht: (2026) -
Sentiment-enhanced Graph-based Sarcasm Explanation in Dialogue
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)