Guardado en:
| Autores principales: | Townsend, Benjamin, May, Madison, Mackowiak, Katherine, Wells, Christopher |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2403.20101 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
por: Park, Jonggwon, et al.
Publicado: (2025)
por: Park, Jonggwon, et al.
Publicado: (2025)
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
por: Shen, Huawen, et al.
Publicado: (2024)
por: Shen, Huawen, et al.
Publicado: (2024)
Model Interpretability and Rationale Extraction by Input Mask Optimization
por: Brinner, Marc, et al.
Publicado: (2025)
por: Brinner, Marc, et al.
Publicado: (2025)
Improving MLLM Historical Record Extraction with Test-Time Image
por: Archibald, Taylor, et al.
Publicado: (2025)
por: Archibald, Taylor, et al.
Publicado: (2025)
Robustness of Structured Data Extraction from Perspectively Distorted Documents
por: Nakada, Hyakka, et al.
Publicado: (2025)
por: Nakada, Hyakka, et al.
Publicado: (2025)
Text-Enhanced Data-free Approach for Federated Class-Incremental Learning
por: Tran, Minh-Tuan, et al.
Publicado: (2024)
por: Tran, Minh-Tuan, et al.
Publicado: (2024)
Overconfidence is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models
por: Groot, Tobias, et al.
Publicado: (2024)
por: Groot, Tobias, et al.
Publicado: (2024)
Improving Resnet-9 Generalization Trained on Small Datasets
por: Awad, Omar Mohamed, et al.
Publicado: (2023)
por: Awad, Omar Mohamed, et al.
Publicado: (2023)
Real-time Bangla Sign Language Translator
por: Pranto, Rotan Hawlader, et al.
Publicado: (2024)
por: Pranto, Rotan Hawlader, et al.
Publicado: (2024)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
por: Li, Ang, et al.
Publicado: (2025)
por: Li, Ang, et al.
Publicado: (2025)
Omnimodal Dataset Distillation via High-order Proxy Alignment
por: Gao, Yuxuan, et al.
Publicado: (2026)
por: Gao, Yuxuan, et al.
Publicado: (2026)
From Pixels to Prose: A Large Dataset of Dense Image Captions
por: Singla, Vasu, et al.
Publicado: (2024)
por: Singla, Vasu, et al.
Publicado: (2024)
One Category One Prompt: Dataset Distillation using Diffusion Models
por: Abbasi, Ali, et al.
Publicado: (2024)
por: Abbasi, Ali, et al.
Publicado: (2024)
Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
por: Madaan, Divyam, et al.
Publicado: (2025)
por: Madaan, Divyam, et al.
Publicado: (2025)
Neural Style Transfer for Synthesising a Dataset of Ancient Egyptian Hieroglyphs
por: Creed, Lewis Matheson
Publicado: (2025)
por: Creed, Lewis Matheson
Publicado: (2025)
SciGA: A Comprehensive Dataset for Designing Graphical Abstracts in Academic Papers
por: Kawada, Takuro, et al.
Publicado: (2025)
por: Kawada, Takuro, et al.
Publicado: (2025)
WLASL-LEX: a Dataset for Recognising Phonological Properties in American Sign Language
por: Tavella, Federico, et al.
Publicado: (2022)
por: Tavella, Federico, et al.
Publicado: (2022)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
por: Lù, Xing Han, et al.
Publicado: (2024)
por: Lù, Xing Han, et al.
Publicado: (2024)
Brazilian Portuguese Image Captioning with Transformers: A Study on Cross-Native-Translated Dataset
por: Bromonschenkel, Gabriel, et al.
Publicado: (2026)
por: Bromonschenkel, Gabriel, et al.
Publicado: (2026)
E-TSL: A Continuous Educational Turkish Sign Language Dataset with Baseline Methods
por: Öztürk, Şükrü, et al.
Publicado: (2024)
por: Öztürk, Şükrü, et al.
Publicado: (2024)
TechING: Towards Real World Technical Image Understanding via VLMs
por: Nadeem, Tafazzul, et al.
Publicado: (2026)
por: Nadeem, Tafazzul, et al.
Publicado: (2026)
DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
por: Picón, Ginés Carreto, et al.
Publicado: (2025)
por: Picón, Ginés Carreto, et al.
Publicado: (2025)
Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
por: Qian, Yusu, et al.
Publicado: (2025)
por: Qian, Yusu, et al.
Publicado: (2025)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
por: Heyward, Joseph, et al.
Publicado: (2024)
por: Heyward, Joseph, et al.
Publicado: (2024)
ImplicitAVE: An Open-Source Dataset and Multimodal LLMs Benchmark for Implicit Attribute Value Extraction
por: Zou, Henry Peng, et al.
Publicado: (2024)
por: Zou, Henry Peng, et al.
Publicado: (2024)
AiGen-FoodReview: A Multimodal Dataset of Machine-Generated Restaurant Reviews and Images on Social Media
por: Gambetti, Alessandro, et al.
Publicado: (2024)
por: Gambetti, Alessandro, et al.
Publicado: (2024)
BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce
por: Shamael, Mohammad Nazmush, et al.
Publicado: (2024)
por: Shamael, Mohammad Nazmush, et al.
Publicado: (2024)
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models
por: Yi, Hao, et al.
Publicado: (2024)
por: Yi, Hao, et al.
Publicado: (2024)
Improving Multimodal Large Language Models Using Continual Learning
por: Srivastava, Shikhar, et al.
Publicado: (2024)
por: Srivastava, Shikhar, et al.
Publicado: (2024)
GRASP: A Rehearsal Policy for Efficient Online Continual Learning
por: Harun, Md Yousuf, et al.
Publicado: (2023)
por: Harun, Md Yousuf, et al.
Publicado: (2023)
Multi-Modal Hallucination Control by Visual Information Grounding
por: Favero, Alessandro, et al.
Publicado: (2024)
por: Favero, Alessandro, et al.
Publicado: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
por: Yu, Eric Yang, et al.
Publicado: (2024)
por: Yu, Eric Yang, et al.
Publicado: (2024)
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
por: Javadi, Farnoosh, et al.
Publicado: (2023)
por: Javadi, Farnoosh, et al.
Publicado: (2023)
A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models
por: Xiu, Lixin, et al.
Publicado: (2026)
por: Xiu, Lixin, et al.
Publicado: (2026)
Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
por: Hao, Tianxiang, et al.
Publicado: (2023)
por: Hao, Tianxiang, et al.
Publicado: (2023)
BloomVQA: Assessing Hierarchical Multi-modal Comprehension
por: Gong, Yunye, et al.
Publicado: (2023)
por: Gong, Yunye, et al.
Publicado: (2023)
Improve Academic Query Resolution through BERT-based Question Extraction from Images
por: Kamal, Nidhi, et al.
Publicado: (2024)
por: Kamal, Nidhi, et al.
Publicado: (2024)
Exploring Attention Mechanisms in Integration of Multi-Modal Information for Sign Language Recognition and Translation
por: Hakim, Zaber Ibn Abdul, et al.
Publicado: (2023)
por: Hakim, Zaber Ibn Abdul, et al.
Publicado: (2023)
FisherMask: Enhancing Neural Network Labeling Efficiency in Image Classification Using Fisher Information
por: Gul, Shreen, et al.
Publicado: (2024)
por: Gul, Shreen, et al.
Publicado: (2024)
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
por: Ye, Suyu, et al.
Publicado: (2025)
por: Ye, Suyu, et al.
Publicado: (2025)
Ejemplares similares
-
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
por: Park, Jonggwon, et al.
Publicado: (2025) -
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
por: Shen, Huawen, et al.
Publicado: (2024) -
Model Interpretability and Rationale Extraction by Input Mask Optimization
por: Brinner, Marc, et al.
Publicado: (2025) -
Improving MLLM Historical Record Extraction with Test-Time Image
por: Archibald, Taylor, et al.
Publicado: (2025) -
Robustness of Structured Data Extraction from Perspectively Distorted Documents
por: Nakada, Hyakka, et al.
Publicado: (2025)