CCS: Clinical Consensus Selection for Radiology Report Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xi, Li, Yingshu, Meng, Zaiqiao, Lever, Jake, Ho, Edmond S. L. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
by: Zhang, Xi, et al.
Published: (2025)
by: Zhang, Xi, et al.
Published: (2025)
Libra: Leveraging Temporal Images for Biomedical Radiology Analysis
by: Zhang, Xi, et al.
Published: (2024)
by: Zhang, Xi, et al.
Published: (2024)
Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation
by: Zhang, Xi, et al.
Published: (2024)
by: Zhang, Xi, et al.
Published: (2024)
DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models
by: Wu, Biao, et al.
Published: (2026)
by: Wu, Biao, et al.
Published: (2026)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation
by: Yang, Qi, et al.
Published: (2026)
by: Yang, Qi, et al.
Published: (2026)
HalalBench: A Multilingual OCR Benchmark for Food Packaging Ingredient Extraction
by: Arief, Hasan
Published: (2026)
by: Arief, Hasan
Published: (2026)
Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals
by: Baghel, Shruti Singh, et al.
Published: (2025)
by: Baghel, Shruti Singh, et al.
Published: (2025)
The American Sign Language Knowledge Graph: Infusing ASL Models with Linguistic Knowledge
by: Kezar, Lee, et al.
Published: (2024)
by: Kezar, Lee, et al.
Published: (2024)
GroundCap: A Visually Grounded Image Captioning Dataset
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
by: Freitas, Diogo, et al.
Published: (2025)
by: Freitas, Diogo, et al.
Published: (2025)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
by: He, Wei
Published: (2026)
by: He, Wei
Published: (2026)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
by: Ji, Yikun, et al.
Published: (2025)
by: Ji, Yikun, et al.
Published: (2025)
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
by: Dehghani, Mahshid, et al.
Published: (2024)
by: Dehghani, Mahshid, et al.
Published: (2024)
SemEval-2025 Task 1: AdMIRe -- Advancing Multimodal Idiomaticity Representation
by: Pickard, Thomas, et al.
Published: (2025)
by: Pickard, Thomas, et al.
Published: (2025)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
by: Cui, Shaoyang, et al.
Published: (2026)
by: Cui, Shaoyang, et al.
Published: (2026)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
by: Cao, Jingtao, et al.
Published: (2024)
by: Cao, Jingtao, et al.
Published: (2024)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
by: Arad, Dana, et al.
Published: (2023)
by: Arad, Dana, et al.
Published: (2023)
Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation Perspective
by: Weng, Zhaotian, et al.
Published: (2024)
by: Weng, Zhaotian, et al.
Published: (2024)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
R-Genie: Reasoning-Guided Generative Image Editing
by: Zhang, Dong, et al.
Published: (2025)
by: Zhang, Dong, et al.
Published: (2025)
SOLAR: Communication-Efficient Model Adaptation via Subspace-Oriented Latent Adapter Reparametrization
by: Mohammadabadi, Seyed Mahmoud Sajjadi, et al.
Published: (2026)
by: Mohammadabadi, Seyed Mahmoud Sajjadi, et al.
Published: (2026)
Semantic Leakage from Image Embeddings
by: Chen, Yiyi, et al.
Published: (2026)
by: Chen, Yiyi, et al.
Published: (2026)
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
by: Parcalabescu, Letitia, et al.
Published: (2021)
by: Parcalabescu, Letitia, et al.
Published: (2021)
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2025)
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2025)
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
by: Parcalabescu, Letitia, et al.
Published: (2022)
by: Parcalabescu, Letitia, et al.
Published: (2022)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection
by: Kim, Songsoo, et al.
Published: (2025)
by: Kim, Songsoo, et al.
Published: (2025)
Depthwise Separable Convolutions with Deep Residual Convolutions
by: Hasan, Md Arid, et al.
Published: (2024)
by: Hasan, Md Arid, et al.
Published: (2024)
Content Significance Distribution of Sub-Text Blocks in Articles and Its Application to Article-Organization Assessment
by: Zhou, You, et al.
Published: (2023)
by: Zhou, You, et al.
Published: (2023)
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
by: Li, Jianing, et al.
Published: (2024)
by: Li, Jianing, et al.
Published: (2024)
Evaluating Large Language Models for Zero-Shot Disease Labeling in CT Radiology Reports Across Organ Systems
by: Garcia-Alcoser, Michael E., et al.
Published: (2025)
by: Garcia-Alcoser, Michael E., et al.
Published: (2025)
A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports
by: Schäfer, Henning, et al.
Published: (2025)
by: Schäfer, Henning, et al.
Published: (2025)
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
by: Deichler, Anna, et al.
Published: (2026)
by: Deichler, Anna, et al.
Published: (2026)
Using Deep Learning to Generate Semantically Correct Hindi Captions
by: Khan, Wasim Akram, et al.
Published: (2026)
by: Khan, Wasim Akram, et al.
Published: (2026)
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
by: Niu, Yuwei, et al.
Published: (2025)
by: Niu, Yuwei, et al.
Published: (2025)
SeRpEnt: Selective Resampling for Expressive State Space Models
by: Rando, Stefano, et al.
Published: (2025)
by: Rando, Stefano, et al.
Published: (2025)
Unsupervised Band Selection Using Fused HSI and LiDAR Attention Integrating With Autoencoder
by: Yang, Judy X, et al.
Published: (2024)
by: Yang, Judy X, et al.
Published: (2024)
Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
by: Ji, Yikun, et al.
Published: (2025)
by: Ji, Yikun, et al.
Published: (2025)
Similar Items
-
CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
by: Zhang, Xi, et al.
Published: (2025) -
Libra: Leveraging Temporal Images for Biomedical Radiology Analysis
by: Zhang, Xi, et al.
Published: (2024) -
Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation
by: Zhang, Xi, et al.
Published: (2024) -
DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models
by: Wu, Biao, et al.
Published: (2026) -
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
by: Oliveira, Daniel A. P., et al.
Published: (2025)