Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
Fuente:
arXiv
Guardado en:
| Autores principales: | Ko, Hanbin, Cho, Gihun, Baek, Inhyeok, Kim, Donguk, Koo, Joonbeom, Kim, Changi, Lee, Dongheon, Park, Chang Min |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MaP-AVR: A Meta-Action Planner for Agents Leveraging Vision Language Models and Retrieval-Augmented Generation
por: Guo, Zhenglong, et al.
Publicado: (2025)
por: Guo, Zhenglong, et al.
Publicado: (2025)
Sequence Matters: Harnessing Video Models in 3D Super-Resolution
por: Ko, Hyun-kyu, et al.
Publicado: (2024)
por: Ko, Hyun-kyu, et al.
Publicado: (2024)
PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation
por: Alanazi, Ahmed, et al.
Publicado: (2025)
por: Alanazi, Ahmed, et al.
Publicado: (2025)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
por: Syed, Shahram Najam, et al.
Publicado: (2025)
por: Syed, Shahram Najam, et al.
Publicado: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
por: Li, Jinhao, et al.
Publicado: (2024)
por: Li, Jinhao, et al.
Publicado: (2024)
OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment
por: Zhao, Weiyi, et al.
Publicado: (2025)
por: Zhao, Weiyi, et al.
Publicado: (2025)
Learning Continuous Receive Apodization Weights via Implicit Neural Representation for Ultrafast ICE Ultrasound Imaging
por: Delaunay, Rémi, et al.
Publicado: (2025)
por: Delaunay, Rémi, et al.
Publicado: (2025)
Modulated INR with Prior Embeddings for Ultrasound Imaging Reconstruction
por: Delaunay, Rémi, et al.
Publicado: (2025)
por: Delaunay, Rémi, et al.
Publicado: (2025)
Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models
por: Chaichuk, Mikhail, et al.
Publicado: (2025)
por: Chaichuk, Mikhail, et al.
Publicado: (2025)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
por: Koh, Hyunseo, et al.
Publicado: (2026)
por: Koh, Hyunseo, et al.
Publicado: (2026)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
por: Chahine, Makram, et al.
Publicado: (2024)
por: Chahine, Makram, et al.
Publicado: (2024)
VDPP: Video Depth Post-Processing for Speed and Scalability
por: Yoon, Daewon, et al.
Publicado: (2026)
por: Yoon, Daewon, et al.
Publicado: (2026)
Banana Ripeness Level Classification using a Simple CNN Model Trained with Real and Synthetic Datasets
por: Chuquimarca, Luis, et al.
Publicado: (2025)
por: Chuquimarca, Luis, et al.
Publicado: (2025)
BreastDCEDL: A Comprehensive Breast Cancer DCE-MRI Dataset and Transformer Implementation for Treatment Response Prediction
por: Fridman, Naomi, et al.
Publicado: (2025)
por: Fridman, Naomi, et al.
Publicado: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
por: Adžemović, Momir
Publicado: (2025)
por: Adžemović, Momir
Publicado: (2025)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
por: Chen, Kewei, et al.
Publicado: (2025)
por: Chen, Kewei, et al.
Publicado: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
por: Chen, Kewei, et al.
Publicado: (2026)
por: Chen, Kewei, et al.
Publicado: (2026)
Tricks and Plug-ins for Gradient Boosting in Image Classification
por: Fang, Biyi, et al.
Publicado: (2025)
por: Fang, Biyi, et al.
Publicado: (2025)
Edged USLAM: Edge-Aware Event-Based SLAM with Learning-Based Depth Priors
por: Sarıözkan, Şebnem, et al.
Publicado: (2026)
por: Sarıözkan, Şebnem, et al.
Publicado: (2026)
Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
por: Wang, Wentao, et al.
Publicado: (2025)
por: Wang, Wentao, et al.
Publicado: (2025)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
por: Zeng, Zhitao, et al.
Publicado: (2026)
por: Zeng, Zhitao, et al.
Publicado: (2026)
Breast Cell Segmentation Under Extreme Data Constraints: Quantum Enhancement Meets Adaptive Loss Stabilization
por: Dasoju, Varun Kumar, et al.
Publicado: (2025)
por: Dasoju, Varun Kumar, et al.
Publicado: (2025)
MATEX: Multi-scale Attention and Text-guided Explainability of Medical Vision-Language Models
por: Imran, Muhammad, et al.
Publicado: (2026)
por: Imran, Muhammad, et al.
Publicado: (2026)
When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
por: Allakhverdov, Eduard, et al.
Publicado: (2025)
por: Allakhverdov, Eduard, et al.
Publicado: (2025)
Image Reconstruction as a Tool for Feature Analysis
por: Allakhverdov, Eduard, et al.
Publicado: (2025)
por: Allakhverdov, Eduard, et al.
Publicado: (2025)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
por: Komurcu, Kursat, et al.
Publicado: (2026)
por: Komurcu, Kursat, et al.
Publicado: (2026)
Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning
por: Koubaa, Anis, et al.
Publicado: (2025)
por: Koubaa, Anis, et al.
Publicado: (2025)
Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Systematic Review and Practical Design Guidelines
por: Wimalasiri, Chathura
Publicado: (2026)
por: Wimalasiri, Chathura
Publicado: (2026)
Learning Sign Language Representation using CNN LSTM, 3DCNN, CNN RNN LSTM and CCN TD
por: Louison, Nikita, et al.
Publicado: (2024)
por: Louison, Nikita, et al.
Publicado: (2024)
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
A Novel Approach to Breast Cancer Segmentation using U-Net Model with Attention Mechanisms and FedProx
por: Gad, Eyad, et al.
Publicado: (2025)
por: Gad, Eyad, et al.
Publicado: (2025)
DSER: Spectral Epipolar Representation for Efficient Light Field Depth Estimation
por: Mohammad, Noor Islam S., et al.
Publicado: (2025)
por: Mohammad, Noor Islam S., et al.
Publicado: (2025)
Hierarchical Spatial Algorithms for High-Resolution Image Quantization and Feature Extraction
por: Mohammad, Noor Islam S.
Publicado: (2025)
por: Mohammad, Noor Islam S.
Publicado: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
por: Kim, Soyeon, et al.
Publicado: (2026)
por: Kim, Soyeon, et al.
Publicado: (2026)
VGA: Vision GUI Assistant -- Minimizing Hallucinations through Image-Centric Fine-Tuning
por: Meng, Ziyang, et al.
Publicado: (2024)
por: Meng, Ziyang, et al.
Publicado: (2024)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
The MSR-Video to Text Dataset with Clean Annotations
por: Chen, Haoran, et al.
Publicado: (2021)
por: Chen, Haoran, et al.
Publicado: (2021)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
por: Marian, Vasile, et al.
Publicado: (2026)
por: Marian, Vasile, et al.
Publicado: (2026)
Low Dose CT for Stroke Diagnosis: A Dual Pipeline Deep Learning Framework for Portable Neuroimaging
por: Ghosal, Rhea, et al.
Publicado: (2026)
por: Ghosal, Rhea, et al.
Publicado: (2026)
Ejemplares similares
-
MaP-AVR: A Meta-Action Planner for Agents Leveraging Vision Language Models and Retrieval-Augmented Generation
por: Guo, Zhenglong, et al.
Publicado: (2025) -
Sequence Matters: Harnessing Video Models in 3D Super-Resolution
por: Ko, Hyun-kyu, et al.
Publicado: (2024) -
PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation
por: Alanazi, Ahmed, et al.
Publicado: (2025) -
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
por: Syed, Shahram Najam, et al.
Publicado: (2025) -
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)