Prompt-Guided Generation of Structured Chest X-Ray Report Using a Pre-trained LLM
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Hongzhao, Wang, Hongyu, Sun, Xia, He, Hua, Feng, Jun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster
por: Tang, Fenghe, et al.
Publicado: (2025)
por: Tang, Fenghe, et al.
Publicado: (2025)
Reinforcing Pre-trained Models Using Counterfactual Images
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
por: Wang, Xiao, et al.
Publicado: (2023)
por: Wang, Xiao, et al.
Publicado: (2023)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
por: Zhao, Xinyu, et al.
Publicado: (2026)
por: Zhao, Xinyu, et al.
Publicado: (2026)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
por: Han, Junlin, et al.
Publicado: (2025)
por: Han, Junlin, et al.
Publicado: (2025)
DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
por: He, Huiguo, et al.
Publicado: (2024)
por: He, Huiguo, et al.
Publicado: (2024)
Generalized Face Forgery Detection via Adaptive Learning for Pre-trained Vision Transformer
por: Luo, Anwei, et al.
Publicado: (2023)
por: Luo, Anwei, et al.
Publicado: (2023)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
por: Dong, Hao, et al.
Publicado: (2026)
por: Dong, Hao, et al.
Publicado: (2026)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
por: Sun, Zeyi, et al.
Publicado: (2024)
por: Sun, Zeyi, et al.
Publicado: (2024)
Modularized Zero-shot VQA with Pre-trained Models
por: Cao, Rui, et al.
Publicado: (2023)
por: Cao, Rui, et al.
Publicado: (2023)
LLM-EvRep: Learning an LLM-Compatible Event Representation Using a Self-Supervised Framework
por: Yu, Zongyou, et al.
Publicado: (2025)
por: Yu, Zongyou, et al.
Publicado: (2025)
A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
por: Hei, Nailei, et al.
Publicado: (2024)
por: Hei, Nailei, et al.
Publicado: (2024)
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
por: Zhou, Pengyuan, et al.
Publicado: (2024)
por: Zhou, Pengyuan, et al.
Publicado: (2024)
Scaling up Multimodal Pre-training for Sign Language Understanding
por: Zhou, Wengang, et al.
Publicado: (2024)
por: Zhou, Wengang, et al.
Publicado: (2024)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
por: Hsiao, Teng-Fang, et al.
Publicado: (2024)
por: Hsiao, Teng-Fang, et al.
Publicado: (2024)
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
por: Khac, Phúc H. Le, et al.
Publicado: (2024)
por: Khac, Phúc H. Le, et al.
Publicado: (2024)
AKiRa: Augmentation Kit on Rays for optical video generation
por: Wang, Xi, et al.
Publicado: (2024)
por: Wang, Xi, et al.
Publicado: (2024)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
por: Gao, Jiayi, et al.
Publicado: (2025)
por: Gao, Jiayi, et al.
Publicado: (2025)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
por: Cai, Minghong, et al.
Publicado: (2024)
por: Cai, Minghong, et al.
Publicado: (2024)
DyRoNet: Dynamic Routing and Low-Rank Adapters for Autonomous Driving Streaming Perception
por: Huang, Xiang, et al.
Publicado: (2024)
por: Huang, Xiang, et al.
Publicado: (2024)
Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
por: Zhang, Tong, et al.
Publicado: (2025)
por: Zhang, Tong, et al.
Publicado: (2025)
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection
por: Wang, Yihao, et al.
Publicado: (2024)
por: Wang, Yihao, et al.
Publicado: (2024)
Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
por: Phung, Thu Hang, et al.
Publicado: (2026)
por: Phung, Thu Hang, et al.
Publicado: (2026)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
por: Baraldi, Lorenzo, et al.
Publicado: (2023)
por: Baraldi, Lorenzo, et al.
Publicado: (2023)
Apollo: Unified Multi-Task Audio-Video Joint Generation
por: Wang, Jun, et al.
Publicado: (2026)
por: Wang, Jun, et al.
Publicado: (2026)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
por: Lin, Haokun, et al.
Publicado: (2024)
por: Lin, Haokun, et al.
Publicado: (2024)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
por: Ukai, Mahiro, et al.
Publicado: (2024)
por: Ukai, Mahiro, et al.
Publicado: (2024)
SNP-S3: Shared Network Pre-training and Significant Semantic Strengthening for Various Video-Text Tasks
por: Dong, Xingning, et al.
Publicado: (2024)
por: Dong, Xingning, et al.
Publicado: (2024)
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
por: Cui, Can, et al.
Publicado: (2024)
por: Cui, Can, et al.
Publicado: (2024)
GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
por: Sun, Hao, et al.
Publicado: (2025)
por: Sun, Hao, et al.
Publicado: (2025)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
por: Yuan, Hangjie, et al.
Publicado: (2025)
por: Yuan, Hangjie, et al.
Publicado: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
por: Wang, Yi, et al.
Publicado: (2025)
por: Wang, Yi, et al.
Publicado: (2025)
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
por: Zi, Xing, et al.
Publicado: (2025)
por: Zi, Xing, et al.
Publicado: (2025)
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
por: Cao, Pu, et al.
Publicado: (2023)
por: Cao, Pu, et al.
Publicado: (2023)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
por: Lv, Zheqi, et al.
Publicado: (2025)
por: Lv, Zheqi, et al.
Publicado: (2025)
Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
por: Xuan, Yunyi, et al.
Publicado: (2024)
por: Xuan, Yunyi, et al.
Publicado: (2024)
Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
por: Wu, Ruiqi, et al.
Publicado: (2024)
por: Wu, Ruiqi, et al.
Publicado: (2024)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
por: Ling, Jun, et al.
Publicado: (2024)
por: Ling, Jun, et al.
Publicado: (2024)
Enhancing Generalization in Medical Visual Question Answering Tasks via Gradient-Guided Model Perturbation
por: Liu, Gang, et al.
Publicado: (2024)
por: Liu, Gang, et al.
Publicado: (2024)
Generative Preprocessing for Image Compression with Pre-trained Diffusion Models
por: Guo, Mengxi, et al.
Publicado: (2025)
por: Guo, Mengxi, et al.
Publicado: (2025)
Ejemplares similares
-
Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster
por: Tang, Fenghe, et al.
Publicado: (2025) -
Reinforcing Pre-trained Models Using Counterfactual Images
por: Li, Xiang, et al.
Publicado: (2024) -
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
por: Wang, Xiao, et al.
Publicado: (2023) -
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
por: Zhao, Xinyu, et al.
Publicado: (2026) -
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
por: Han, Junlin, et al.
Publicado: (2025)