Sumotosima: A Framework and Dataset for Classifying and Summarizing Otoscopic Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khan, Eram Anwarul, Khan, Anas Anwarul Haq |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
von: Khan, Anas Anwarul Haq, et al.
Veröffentlicht: (2025)
von: Khan, Anas Anwarul Haq, et al.
Veröffentlicht: (2025)
RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture
von: Khan, Anas Anwarul Haq, et al.
Veröffentlicht: (2026)
von: Khan, Anas Anwarul Haq, et al.
Veröffentlicht: (2026)
Improving Pediatric Pneumonia Diagnosis with Adult Chest X-ray Images Utilizing Contrastive Learning and Embedding Similarity
von: Zunaed, Mohammad, et al.
Veröffentlicht: (2024)
von: Zunaed, Mohammad, et al.
Veröffentlicht: (2024)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script
von: Malik, Haq Nawaz
Veröffentlicht: (2026)
von: Malik, Haq Nawaz
Veröffentlicht: (2026)
MambaFormer: Token-Level Guided Routing Mixture-of-Experts for Accurate and Efficient Clinical Assistance
von: Khan, Hamad, et al.
Veröffentlicht: (2026)
von: Khan, Hamad, et al.
Veröffentlicht: (2026)
See, Explain, and Intervene: A Few-Shot Multimodal Agent Framework for Hateful Meme Moderation
von: Rizwan, Naquee, et al.
Veröffentlicht: (2026)
von: Rizwan, Naquee, et al.
Veröffentlicht: (2026)
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
von: Liu, Dongqi, et al.
Veröffentlicht: (2025)
von: Liu, Dongqi, et al.
Veröffentlicht: (2025)
A Paradigm Shift in Mouza Map Vectorization: A Human-Machine Collaboration Approach
von: Dhrubo, Mahir Shahriar, et al.
Veröffentlicht: (2024)
von: Dhrubo, Mahir Shahriar, et al.
Veröffentlicht: (2024)
ReMI: A Dataset for Reasoning with Multiple Images
von: Kazemi, Mehran, et al.
Veröffentlicht: (2024)
von: Kazemi, Mehran, et al.
Veröffentlicht: (2024)
PALO: A Polyglot Large Multimodal Model for 5B People
von: Maaz, Muhammad, et al.
Veröffentlicht: (2024)
von: Maaz, Muhammad, et al.
Veröffentlicht: (2024)
Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization
von: Islam, Md Moinul, et al.
Veröffentlicht: (2025)
von: Islam, Md Moinul, et al.
Veröffentlicht: (2025)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
Wolf: Dense Video Captioning with a World Summarization Framework
von: Li, Boyi, et al.
Veröffentlicht: (2024)
von: Li, Boyi, et al.
Veröffentlicht: (2024)
Multimodal Human-AI Synergy for Medical Imaging Quality Control: A Hybrid Intelligence Framework with Adaptive Dataset Curation and Closed-Loop Evaluation
von: Qin, Zhi, et al.
Veröffentlicht: (2025)
von: Qin, Zhi, et al.
Veröffentlicht: (2025)
#PraCegoVer: A Large Dataset for Image Captioning in Portuguese
von: Santos, Gabriel Oliveira dos, et al.
Veröffentlicht: (2021)
von: Santos, Gabriel Oliveira dos, et al.
Veröffentlicht: (2021)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
von: Lin, Jingyang, et al.
Veröffentlicht: (2023)
von: Lin, Jingyang, et al.
Veröffentlicht: (2023)
Chitrakshara: A Large Multilingual Multimodal Dataset for Indian languages
von: Khan, Shaharukh, et al.
Veröffentlicht: (2026)
von: Khan, Shaharukh, et al.
Veröffentlicht: (2026)
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
AutoGen Driven Multi Agent Framework for Iterative Crime Data Analysis and Prediction
von: Fatima, Syeda Kisaa, et al.
Veröffentlicht: (2025)
von: Fatima, Syeda Kisaa, et al.
Veröffentlicht: (2025)
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
von: Chau, Phan Phuong Mai, et al.
Veröffentlicht: (2024)
von: Chau, Phan Phuong Mai, et al.
Veröffentlicht: (2024)
Realizing Video Summarization from the Path of Language-based Semantic Understanding
von: Mu, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Mu, Kuan-Chen, et al.
Veröffentlicht: (2024)
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
von: Malik, Haq Nawaz, et al.
Veröffentlicht: (2026)
von: Malik, Haq Nawaz, et al.
Veröffentlicht: (2026)
BIMCV-R: A Landmark Dataset for 3D CT Text-Image Retrieval
von: Chen, Yinda, et al.
Veröffentlicht: (2024)
von: Chen, Yinda, et al.
Veröffentlicht: (2024)
Performance Analysis of Image Classification on Bangladeshi Datasets
von: Khan, Mohammed Sami, et al.
Veröffentlicht: (2026)
von: Khan, Mohammed Sami, et al.
Veröffentlicht: (2026)
Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
von: Gigant, Théo, et al.
Veröffentlicht: (2025)
von: Gigant, Théo, et al.
Veröffentlicht: (2025)
Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)
Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction
von: Pennec, Galann, et al.
Veröffentlicht: (2025)
von: Pennec, Galann, et al.
Veröffentlicht: (2025)
Less Is More? Selective Visual Attention to High-Importance Regions for Multimodal Radiology Summarization
von: Naznin, Mst. Fahmida Sultana, et al.
Veröffentlicht: (2026)
von: Naznin, Mst. Fahmida Sultana, et al.
Veröffentlicht: (2026)
A Unified Agentic Framework for Evaluating Conditional Image Generation
von: Wang, Jifang, et al.
Veröffentlicht: (2025)
von: Wang, Jifang, et al.
Veröffentlicht: (2025)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
von: Pham, Thang M., et al.
Veröffentlicht: (2024)
von: Pham, Thang M., et al.
Veröffentlicht: (2024)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
CGRA-DeBERTa Concept Guided Residual Augmentation Transformer for Theologically Islamic Understanding
von: Hussain, Tahir, et al.
Veröffentlicht: (2026)
von: Hussain, Tahir, et al.
Veröffentlicht: (2026)
Frame Sampling Strategies Matter: A Benchmark for small vision language models
von: Brkic, Marija, et al.
Veröffentlicht: (2025)
von: Brkic, Marija, et al.
Veröffentlicht: (2025)
LADDER: Language-Driven Slice Discovery and Error Rectification in Vision Classifiers
von: Ghosh, Shantanu, et al.
Veröffentlicht: (2024)
von: Ghosh, Shantanu, et al.
Veröffentlicht: (2024)
Train a Unified Multimodal Data Quality Classifier with Synthetic Data
von: Wang, Weizhi, et al.
Veröffentlicht: (2025)
von: Wang, Weizhi, et al.
Veröffentlicht: (2025)
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
von: Khan, Anas Anwarul Haq, et al.
Veröffentlicht: (2025) -
RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture
von: Khan, Anas Anwarul Haq, et al.
Veröffentlicht: (2026) -
Improving Pediatric Pneumonia Diagnosis with Adult Chest X-ray Images Utilizing Contrastive Learning and Embedding Similarity
von: Zunaed, Mohammad, et al.
Veröffentlicht: (2024) -
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
von: Zang, Yuan, et al.
Veröffentlicht: (2025) -
600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script
von: Malik, Haq Nawaz
Veröffentlicht: (2026)