DeVisE: Behavioral Testing of Medical Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tagliabue, Camila Zurdo, Boll, Heloisa Oss, Erdem, Aykut, Erdem, Erkut, Calixto, Iacer |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FewMMBench: A Benchmark for Multimodal Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2026)
by: Dogan, Mustafa, et al.
Published: (2026)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024)
by: Dogan, Mustafa, et al.
Published: (2024)
DistillNote: Toward a Functional Evaluation Framework of LLM-Generated Clinical Note Summaries
by: Boll, Heloisa Oss, et al.
Published: (2025)
by: Boll, Heloisa Oss, et al.
Published: (2025)
Sequential Compositional Generalization in Multimodal Models
by: Yagcioglu, Semih, et al.
Published: (2024)
by: Yagcioglu, Semih, et al.
Published: (2024)
Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
by: Acikgoz, Emre Can, et al.
Published: (2024)
by: Acikgoz, Emre Can, et al.
Published: (2024)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025)
by: Karanfil, Enes, et al.
Published: (2025)
Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding
by: Vural, Hatice Merve, et al.
Published: (2026)
by: Vural, Hatice Merve, et al.
Published: (2026)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
by: Sanli, Enes, et al.
Published: (2025)
by: Sanli, Enes, et al.
Published: (2025)
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)
by: Ercan, Burak, et al.
Published: (2024)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA
by: Testoni, Alberto, et al.
Published: (2026)
by: Testoni, Alberto, et al.
Published: (2026)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
by: Bond, Andrew, et al.
Published: (2026)
by: Bond, Andrew, et al.
Published: (2026)
Mind the Gap: Benchmarking LLM Uncertainty and Calibration with Specialty-Aware Clinical QA and Reasoning-Based Behavioural Features
by: Testoni, Alberto, et al.
Published: (2025)
by: Testoni, Alberto, et al.
Published: (2025)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
by: Ekin, Yigit, et al.
Published: (2024)
by: Ekin, Yigit, et al.
Published: (2024)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
by: Bond, Andrew, et al.
Published: (2025)
by: Bond, Andrew, et al.
Published: (2025)
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)
by: Kizil, Muhammed Burak, et al.
Published: (2025)
AnyMatch -- Efficient Zero-Shot Entity Matching with a Small Language Model
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking
by: Mishra, Nishant, et al.
Published: (2025)
by: Mishra, Nishant, et al.
Published: (2025)
What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs
by: Yan, Xinlan, et al.
Published: (2025)
by: Yan, Xinlan, et al.
Published: (2025)
Domain Specific Specialization in Low-Resource Settings: The Efficacy of Offline Response-Based Knowledge Distillation in Large Language Models
by: Aslan, Erdem, et al.
Published: (2026)
by: Aslan, Erdem, et al.
Published: (2026)
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
by: Çapuk, Hakan, et al.
Published: (2025)
by: Çapuk, Hakan, et al.
Published: (2025)
SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
by: Biner, Burak Can, et al.
Published: (2024)
by: Biner, Burak Can, et al.
Published: (2024)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026)
by: Kizil, Muhammed Burak, et al.
Published: (2026)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)
by: Cokelek, Mert, et al.
Published: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
by: Er, Yakup Abrek, et al.
Published: (2025)
by: Er, Yakup Abrek, et al.
Published: (2025)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024)
by: Anees, Abdul Basit, et al.
Published: (2024)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
by: Ali, Moayed Haji, et al.
Published: (2023)
by: Ali, Moayed Haji, et al.
Published: (2023)
Object and Relation Centric Representations for Push Effect Prediction
by: Tekden, Ahmet E., et al.
Published: (2021)
by: Tekden, Ahmet E., et al.
Published: (2021)
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
by: Parcalabescu, Letitia, et al.
Published: (2021)
by: Parcalabescu, Letitia, et al.
Published: (2021)
Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation
by: Miranda, Michele, et al.
Published: (2026)
by: Miranda, Michele, et al.
Published: (2026)
Which books do I like?
by: Rosenbusch, Hannes, et al.
Published: (2025)
by: Rosenbusch, Hannes, et al.
Published: (2025)
Core-based Hierarchies for Efficient GraphRAG
by: Hossain, Jakir, et al.
Published: (2026)
by: Hossain, Jakir, et al.
Published: (2026)
Efficient Learning Content Retrieval with Knowledge Injection
by: Sariturk, Batuhan, et al.
Published: (2024)
by: Sariturk, Batuhan, et al.
Published: (2024)
Assessing GPTZero's Accuracy in Identifying AI vs. Human-Written Essays
by: Dik, Selin, et al.
Published: (2025)
by: Dik, Selin, et al.
Published: (2025)
VisEval: A Benchmark for Data Visualization in the Era of Large Language Models
by: Chen, Nan, et al.
Published: (2024)
by: Chen, Nan, et al.
Published: (2024)
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
by: Jaggi, Harbani, et al.
Published: (2024)
by: Jaggi, Harbani, et al.
Published: (2024)
ChatVis: Automating Scientific Visualization with a Large Language Model
by: Mallick, Tanwi, et al.
Published: (2024)
by: Mallick, Tanwi, et al.
Published: (2024)
Detection of Adverse Drug Events in Dutch clinical free text documents using Transformer Models: benchmark study
by: Murphy, Rachel M., et al.
Published: (2025)
by: Murphy, Rachel M., et al.
Published: (2025)
GLOCON Database: Design Decisions and User Manual (v1.0)
by: Hürriyetoğlu, Ali, et al.
Published: (2024)
by: Hürriyetoğlu, Ali, et al.
Published: (2024)
Similar Items
-
FewMMBench: A Benchmark for Multimodal Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2026) -
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024) -
DistillNote: Toward a Functional Evaluation Framework of LLM-Generated Clinical Note Summaries
by: Boll, Heloisa Oss, et al.
Published: (2025) -
Sequential Compositional Generalization in Multimodal Models
by: Yagcioglu, Semih, et al.
Published: (2024) -
Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
by: Acikgoz, Emre Can, et al.
Published: (2024)