A Review of Mechanistic Models of Event Comprehension
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Nguyen, Tan T. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Role of Language Models in Modern Healthcare: A Comprehensive Review
von: Khalid, Amna, et al.
Veröffentlicht: (2024)
von: Khalid, Amna, et al.
Veröffentlicht: (2024)
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
A Comprehensive Review of Sign Language Recognition: Different Types, Modalities, and Datasets
von: Madhiarasan, M., et al.
Veröffentlicht: (2022)
von: Madhiarasan, M., et al.
Veröffentlicht: (2022)
Vision-Language Models for Edge Networks: A Comprehensive Survey
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025)
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025)
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
von: He, Zheqi, et al.
Veröffentlicht: (2025)
von: He, Zheqi, et al.
Veröffentlicht: (2025)
UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
von: Lee, Segyu, et al.
Veröffentlicht: (2026)
von: Lee, Segyu, et al.
Veröffentlicht: (2026)
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
Large Multi-modal Model Cartographic Map Comprehension for Textual Locality Georeferencing
von: Wijegunarathna, Kalana, et al.
Veröffentlicht: (2025)
von: Wijegunarathna, Kalana, et al.
Veröffentlicht: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
The Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language Models
von: Nguyen, Phuc Minh, et al.
Veröffentlicht: (2025)
von: Nguyen, Phuc Minh, et al.
Veröffentlicht: (2025)
YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
von: Nandy, Abhilash, et al.
Veröffentlicht: (2024)
von: Nandy, Abhilash, et al.
Veröffentlicht: (2024)
In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
von: Dönmez, Esra, et al.
Veröffentlicht: (2026)
von: Dönmez, Esra, et al.
Veröffentlicht: (2026)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation
von: Shang, Fangxin, et al.
Veröffentlicht: (2025)
von: Shang, Fangxin, et al.
Veröffentlicht: (2025)
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
von: Li, Shilong, et al.
Veröffentlicht: (2025)
von: Li, Shilong, et al.
Veröffentlicht: (2025)
Fostering Video Reasoning via Next-Event Prediction
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
von: Qi, Yukun, et al.
Veröffentlicht: (2025)
von: Qi, Yukun, et al.
Veröffentlicht: (2025)
Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration
von: Zhang, Yuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyi, et al.
Veröffentlicht: (2025)
Improved GUI Grounding via Iterative Narrowing
von: Nguyen, Anthony
Veröffentlicht: (2024)
von: Nguyen, Anthony
Veröffentlicht: (2024)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
von: Huang, Han, et al.
Veröffentlicht: (2024)
von: Huang, Han, et al.
Veröffentlicht: (2024)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting
von: Islam, Md Mofijul, et al.
Veröffentlicht: (2026)
von: Islam, Md Mofijul, et al.
Veröffentlicht: (2026)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
von: Qin, Lixiong, et al.
Veröffentlicht: (2025)
A Systematic Review of Open Datasets Used in Text-to-Image (T2I) Gen AI Model Safety
von: Rouf, Rakeen, et al.
Veröffentlicht: (2025)
von: Rouf, Rakeen, et al.
Veröffentlicht: (2025)
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
von: Nguyen, Pha, et al.
Veröffentlicht: (2025)
von: Nguyen, Pha, et al.
Veröffentlicht: (2025)
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
von: De, Anik, et al.
Veröffentlicht: (2025)
von: De, Anik, et al.
Veröffentlicht: (2025)
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
von: Yao, Jihan, et al.
Veröffentlicht: (2025)
von: Yao, Jihan, et al.
Veröffentlicht: (2025)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding
von: Willemsen, Bram, et al.
Veröffentlicht: (2024)
von: Willemsen, Bram, et al.
Veröffentlicht: (2024)
ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
von: Hu, Yizhi, et al.
Veröffentlicht: (2025)
von: Hu, Yizhi, et al.
Veröffentlicht: (2025)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024)
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Role of Language Models in Modern Healthcare: A Comprehensive Review
von: Khalid, Amna, et al.
Veröffentlicht: (2024) -
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024) -
A Comprehensive Review of Sign Language Recognition: Different Types, Modalities, and Datasets
von: Madhiarasan, M., et al.
Veröffentlicht: (2022) -
Vision-Language Models for Edge Networks: A Comprehensive Survey
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025) -
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
von: He, Zheqi, et al.
Veröffentlicht: (2025)