SERPENT-VLM : Self-Refining Radiology Report Generation Using Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kapadnis, Manav Nitin, Patnaik, Sohan, Nandy, Abhilash, Ray, Sourjyadip, Goyal, Pawan, Sheet, Debdoot |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$FastDoc$: Domain-Specific Fast Continual Pre-training Technique using Document-Level Metadata and Taxonomy
by: Nandy, Abhilash, et al.
Published: (2023)
by: Nandy, Abhilash, et al.
Published: (2023)
REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback
by: Roy, Aniruddha, et al.
Published: (2025)
by: Roy, Aniruddha, et al.
Published: (2025)
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
by: Ray, Sourjyadip, et al.
Published: (2025)
by: Ray, Sourjyadip, et al.
Published: (2025)
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
by: Ray, Sourjyadip, et al.
Published: (2024)
by: Ray, Sourjyadip, et al.
Published: (2024)
ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
Order-Based Pre-training Strategies for Procedural Text Understanding
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
Knowledge Distillation of Convolutional Neural Networks through Feature Map Transformation using Decision Trees
by: Srinivas, Maddimsetti, et al.
Published: (2024)
by: Srinivas, Maddimsetti, et al.
Published: (2024)
Known Intents, New Combinations: Clause-Factorized Decoding for Compositional Multi-Intent Detection
by: Nandy, Abhilash
Published: (2026)
by: Nandy, Abhilash
Published: (2026)
A Pointer Network-based Approach for Joint Extraction and Detection of Multi-Label Multi-Class Intents
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
Leveraging Self-Attention for Input-Dependent Soft Prompting in LLMs
by: Muppidi, Ananth, et al.
Published: (2025)
by: Muppidi, Ananth, et al.
Published: (2025)
$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
by: Das, Trishanu, et al.
Published: (2025)
by: Das, Trishanu, et al.
Published: (2025)
CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
by: Kapadnis, Manav Nitin, et al.
Published: (2025)
by: Kapadnis, Manav Nitin, et al.
Published: (2025)
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages
by: Ray, Pretam, et al.
Published: (2024)
by: Ray, Pretam, et al.
Published: (2024)
Leveraging Large Language Models for Predictive Analysis of Human Misery
by: Seal, Bishanka, et al.
Published: (2025)
by: Seal, Bishanka, et al.
Published: (2025)
It Helps to Take a Second Opinion: Teaching Smaller LLMs to Deliberate Mutually via Selective Rationale Optimisation
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages
by: Salmè, Marco, et al.
Published: (2025)
by: Salmè, Marco, et al.
Published: (2025)
Architecture, Not Scale: Circuit Localization in Large Language Models
by: Venkatesh, Sohan
Published: (2026)
by: Venkatesh, Sohan
Published: (2026)
Negative Before Positive: Asymmetric Valence Processing in Large Language Models
by: Venkatesh, Sohan
Published: (2026)
by: Venkatesh, Sohan
Published: (2026)
PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics
by: Naik, Atharva, et al.
Published: (2025)
by: Naik, Atharva, et al.
Published: (2025)
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
by: Pellegrini, Chantal, et al.
Published: (2023)
by: Pellegrini, Chantal, et al.
Published: (2023)
VLM-KG: Multimodal Radiology Knowledge Graph Generation
by: Abdullah, Abdullah, et al.
Published: (2025)
by: Abdullah, Abdullah, et al.
Published: (2025)
Generative Large Language Models Trained for Detecting Errors in Radiology Reports
by: Sun, Cong, et al.
Published: (2025)
by: Sun, Cong, et al.
Published: (2025)
Enhancing Low-Resource NMT with a Multilingual Encoder and Knowledge Distillation: A Case Study
by: Roy, Aniruddha, et al.
Published: (2024)
by: Roy, Aniruddha, et al.
Published: (2024)
CYCLE: Learning to Self-Refine the Code Generation
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semantics
by: R V, Kavin, et al.
Published: (2025)
by: R V, Kavin, et al.
Published: (2025)
CABINET: Content Relevance based Noise Reduction for Table Question Answering
by: Patnaik, Sohan, et al.
Published: (2024)
by: Patnaik, Sohan, et al.
Published: (2024)
Automated Structured Radiology Report Generation
by: Delbrouck, Jean-Benoit, et al.
Published: (2025)
by: Delbrouck, Jean-Benoit, et al.
Published: (2025)
RoRA-VLM: Robust Retrieval-Augmented Vision Language Models
by: Qi, Jingyuan, et al.
Published: (2024)
by: Qi, Jingyuan, et al.
Published: (2024)
Development and Validation of a Large Language Model for Generating Fully-Structured Radiology Reports
by: Niu, Chuang, et al.
Published: (2024)
by: Niu, Chuang, et al.
Published: (2024)
Error-Aware Curriculum Learning for Biomedical Relation Classification
by: Chakraborty, Sinchani, et al.
Published: (2025)
by: Chakraborty, Sinchani, et al.
Published: (2025)
Intent Detection and Entity Extraction from BioMedical Literature
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
Large Language Models are Algorithmically Blind
by: Venkatesh, Sohan, et al.
Published: (2026)
by: Venkatesh, Sohan, et al.
Published: (2026)
Calibrated Confidence Expression for Radiology Report Generation
by: Bani-Harouni, David, et al.
Published: (2026)
by: Bani-Harouni, David, et al.
Published: (2026)
High-Fidelity Pseudo-label Generation by Large Language Models for Training Robust Radiology Report Classifiers
by: Wong, Brian, et al.
Published: (2025)
by: Wong, Brian, et al.
Published: (2025)
Radiology-GPT: A Large Language Model for Radiology
by: Liu, Zhengliang, et al.
Published: (2023)
by: Liu, Zhengliang, et al.
Published: (2023)
Chandomitra: Towards Generating Structured Sanskrit Poetry from Natural Language Inputs
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
RaTEScore: A Metric for Radiology Report Generation
by: Zhao, Weike, et al.
Published: (2024)
by: Zhao, Weike, et al.
Published: (2024)
Tracking Cancer Through Text: Longitudinal Extraction From Radiology Reports Using Open-Source Large Language Models
by: Builtjes, Luc, et al.
Published: (2026)
by: Builtjes, Luc, et al.
Published: (2026)
Similar Items
-
$FastDoc$: Domain-Specific Fast Continual Pre-training Technique using Document-Level Metadata and Taxonomy
by: Nandy, Abhilash, et al.
Published: (2023) -
REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback
by: Roy, Aniruddha, et al.
Published: (2025) -
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
by: Ray, Sourjyadip, et al.
Published: (2025) -
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
by: Ray, Sourjyadip, et al.
Published: (2024) -
ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2026)