Image captioning in different languages
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | van Miltenburg, Emiel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Image captioning for Brazilian Portuguese using GRIT model
von: de Alencar, Rafael Silva, et al.
Veröffentlicht: (2024)
von: de Alencar, Rafael Silva, et al.
Veröffentlicht: (2024)
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025)
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025)
Dual use issues in the field of Natural Language Generation
von: van Miltenburg, Emiel
Veröffentlicht: (2025)
von: van Miltenburg, Emiel
Veröffentlicht: (2025)
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
von: Zhang, Chenhui, et al.
Veröffentlicht: (2024)
von: Zhang, Chenhui, et al.
Veröffentlicht: (2024)
Natural Language Generation
von: van Miltenburg, Emiel, et al.
Veröffentlicht: (2025)
von: van Miltenburg, Emiel, et al.
Veröffentlicht: (2025)
VCRScore: Image captioning metric based on V\&L Transformers, CLIP, and precision-recall
von: Ruiz, Guillermo, et al.
Veröffentlicht: (2025)
von: Ruiz, Guillermo, et al.
Veröffentlicht: (2025)
Evaluating Task-oriented Dialogue Systems: A Systematic Review of Measures, Constructs and their Operationalisations
von: Braggaar, Anouck, et al.
Veröffentlicht: (2023)
von: Braggaar, Anouck, et al.
Veröffentlicht: (2023)
Learning text-to-video retrieval from image captioning
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
Linear Alignment of Vision-language Models for Image Captioning
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
The in-context inductive biases of vision-language models differ across modalities
von: Allen, Kelsey, et al.
Veröffentlicht: (2025)
von: Allen, Kelsey, et al.
Veröffentlicht: (2025)
Towards Deployable OCR models for Indic languages
von: Mathew, Minesh, et al.
Veröffentlicht: (2022)
von: Mathew, Minesh, et al.
Veröffentlicht: (2022)
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
Leveraging image captions for selective whole slide image annotation
von: Qiu, Jingna, et al.
Veröffentlicht: (2024)
von: Qiu, Jingna, et al.
Veröffentlicht: (2024)
Fine-grained length controllable video captioning with ordinal embeddings
von: Nitta, Tomoya, et al.
Veröffentlicht: (2024)
von: Nitta, Tomoya, et al.
Veröffentlicht: (2024)
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models
von: Ye-Bin, Moon, et al.
Veröffentlicht: (2024)
von: Ye-Bin, Moon, et al.
Veröffentlicht: (2024)
SignBart -- New approach with the skeleton sequence for Isolated Sign language Recognition
von: Nguyen, Tinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Tinh, et al.
Veröffentlicht: (2025)
Frame Sampling Strategies Matter: A Benchmark for small vision language models
von: Brkic, Marija, et al.
Veröffentlicht: (2025)
von: Brkic, Marija, et al.
Veröffentlicht: (2025)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
von: Tan, Alvin Wei Ming, et al.
Veröffentlicht: (2025)
von: Tan, Alvin Wei Ming, et al.
Veröffentlicht: (2025)
ProfVLM: A lightweight video-language model for multi-view proficiency estimation
von: Bianchi, Edoardo, et al.
Veröffentlicht: (2025)
von: Bianchi, Edoardo, et al.
Veröffentlicht: (2025)
PubMed-Ophtha: An open resource for training ophthalmology vision-language models on scientific literature
von: Hallitschke, Verena Jasmin, et al.
Veröffentlicht: (2026)
von: Hallitschke, Verena Jasmin, et al.
Veröffentlicht: (2026)
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2022)
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2022)
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
von: Malik, Haq Nawaz, et al.
Veröffentlicht: (2026)
von: Malik, Haq Nawaz, et al.
Veröffentlicht: (2026)
Multi-Modal interpretable automatic video captioning
von: Hanna-Asaad, Antoine, et al.
Veröffentlicht: (2024)
von: Hanna-Asaad, Antoine, et al.
Veröffentlicht: (2024)
Interpreting the linear structure of vision-language model embedding spaces
von: Papadimitriou, Isabel, et al.
Veröffentlicht: (2025)
von: Papadimitriou, Isabel, et al.
Veröffentlicht: (2025)
Object-oriented backdoor attack against image captioning
von: Li, Meiling, et al.
Veröffentlicht: (2024)
von: Li, Meiling, et al.
Veröffentlicht: (2024)
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
von: Garg, Roopal, et al.
Veröffentlicht: (2024)
von: Garg, Roopal, et al.
Veröffentlicht: (2024)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
Superhuman performance in urology board questions by an explainable large language model enabled for context integration of the European Association of Urology guidelines: the UroBot study
von: Hetz, Martin J., et al.
Veröffentlicht: (2024)
von: Hetz, Martin J., et al.
Veröffentlicht: (2024)
Multimodal Evaluation of Russian-language Architectures
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Images Speak Volumes: User-Centric Assessment of Image Generation for Accessible Communication
von: Anschütz, Miriam, et al.
Veröffentlicht: (2024)
von: Anschütz, Miriam, et al.
Veröffentlicht: (2024)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
von: Cheng, Sheng, et al.
Veröffentlicht: (2024)
von: Cheng, Sheng, et al.
Veröffentlicht: (2024)
Quality-Aware Image-Text Alignment for Opinion-Unaware Image Quality Assessment
von: Agnolucci, Lorenzo, et al.
Veröffentlicht: (2024)
von: Agnolucci, Lorenzo, et al.
Veröffentlicht: (2024)
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
Subobject-level Image Tokenization
von: Chen, Delong, et al.
Veröffentlicht: (2024)
von: Chen, Delong, et al.
Veröffentlicht: (2024)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Towards Automatic Evaluation for Image Transcreation
von: Khanuja, Simran, et al.
Veröffentlicht: (2024)
von: Khanuja, Simran, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Image captioning for Brazilian Portuguese using GRIT model
von: de Alencar, Rafael Silva, et al.
Veröffentlicht: (2024) -
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025) -
Dual use issues in the field of Natural Language Generation
von: van Miltenburg, Emiel
Veröffentlicht: (2025) -
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
von: Zhang, Chenhui, et al.
Veröffentlicht: (2024) -
Natural Language Generation
von: van Miltenburg, Emiel, et al.
Veröffentlicht: (2025)