MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yim, Wen-wai, Abacha, Asma Ben, Yu, Zixuan, Doerning, Robert, Xia, Fei, Yetisgen, Meliha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
The MSR-Video to Text Dataset with Clean Annotations
von: Chen, Haoran, et al.
Veröffentlicht: (2021)
von: Chen, Haoran, et al.
Veröffentlicht: (2021)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
von: Sáez, Arnau Igualde, et al.
Veröffentlicht: (2025)
von: Sáez, Arnau Igualde, et al.
Veröffentlicht: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
von: Farzulla, Murad
Veröffentlicht: (2026)
von: Farzulla, Murad
Veröffentlicht: (2026)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
von: Qi, Dekang, et al.
Veröffentlicht: (2026)
von: Qi, Dekang, et al.
Veröffentlicht: (2026)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Does CLIP perceive art the same way we do?
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
von: Chen, Kewei, et al.
Veröffentlicht: (2025)
von: Chen, Kewei, et al.
Veröffentlicht: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
von: Chen, Kewei, et al.
Veröffentlicht: (2026)
von: Chen, Kewei, et al.
Veröffentlicht: (2026)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2025)
von: Patel, Urjitkumar, et al.
Veröffentlicht: (2025)
Tricks and Plug-ins for Gradient Boosting in Image Classification
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
von: Gkountouras, John, et al.
Veröffentlicht: (2025)
von: Gkountouras, John, et al.
Veröffentlicht: (2025)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
von: Papyan, Narek, et al.
Veröffentlicht: (2024)
von: Papyan, Narek, et al.
Veröffentlicht: (2024)
ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
von: Liu, Xueyi, et al.
Veröffentlicht: (2025)
von: Liu, Xueyi, et al.
Veröffentlicht: (2025)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
von: Zeng, Zhitao, et al.
Veröffentlicht: (2026)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2026)
RACAS: Controlling Diverse Robots With a Single Agentic System
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
von: Koh, Hyunseo, et al.
Veröffentlicht: (2026)
von: Koh, Hyunseo, et al.
Veröffentlicht: (2026)
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
von: Ouyang, Rongxin, et al.
Veröffentlicht: (2024)
von: Ouyang, Rongxin, et al.
Veröffentlicht: (2024)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
von: Chahine, Makram, et al.
Veröffentlicht: (2024)
von: Chahine, Makram, et al.
Veröffentlicht: (2024)
Multimodal Structure-Aware Quantum Data Processing
von: Hawashin, Hala, et al.
Veröffentlicht: (2024)
von: Hawashin, Hala, et al.
Veröffentlicht: (2024)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
von: Yang, Baoyao, et al.
Veröffentlicht: (2025)
von: Yang, Baoyao, et al.
Veröffentlicht: (2025)
OpenFusion++: An Open-vocabulary Real-time Scene Understanding System
von: Jin, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Jin, Xiaofeng, et al.
Veröffentlicht: (2025)
Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving
von: Da, Longchao, et al.
Veröffentlicht: (2025)
von: Da, Longchao, et al.
Veröffentlicht: (2025)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
PDFMathTranslate: Scientific Document Translation Preserving Layouts
von: Ouyang, Rongxin, et al.
Veröffentlicht: (2025)
von: Ouyang, Rongxin, et al.
Veröffentlicht: (2025)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
von: Marian, Vasile, et al.
Veröffentlicht: (2026)
von: Marian, Vasile, et al.
Veröffentlicht: (2026)
Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
von: Kalušev, Vladimir, et al.
Veröffentlicht: (2026)
von: Kalušev, Vladimir, et al.
Veröffentlicht: (2026)
RefineFormer3D: Efficient 3D Medical Image Segmentation via Adaptive Multi-Scale Transformer with Cross Attention Fusion
von: Tyagi, Kavyansh, et al.
Veröffentlicht: (2026)
von: Tyagi, Kavyansh, et al.
Veröffentlicht: (2026)
AUTHENTICATION: Identifying Rare Failure Modes in Autonomous Vehicle Perception Systems using Adversarially Guided Diffusion Models
von: Zarei, Mohammad, et al.
Veröffentlicht: (2025)
von: Zarei, Mohammad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025) -
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
von: Huo, Dongjie, et al.
Veröffentlicht: (2026) -
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
von: Viveiros, André G., et al.
Veröffentlicht: (2025) -
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025) -
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)