Image captioning for Brazilian Portuguese using GRIT model
Fuente:
arXiv
Salvato in:
| Autori principali: | de Alencar, Rafael Silva, Castañeda, William Alberto Cruz, Amadeus, Marcellus |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Amadeus-Verbo Technical Report: The powerful Qwen2.5 family models trained in Portuguese
di: Cruz-Castañeda, William Alberto, et al.
Pubblicazione: (2025)
di: Cruz-Castañeda, William Alberto, et al.
Pubblicazione: (2025)
From Pampas to Pixels: Fine-Tuning Diffusion Models for Gaúcho Heritage
di: Amadeus, Marcellus, et al.
Pubblicazione: (2024)
di: Amadeus, Marcellus, et al.
Pubblicazione: (2024)
An Inpainting-Infused Pipeline for Attire and Background Replacement
di: Perche-Mahlow, Felipe Rodrigues, et al.
Pubblicazione: (2024)
di: Perche-Mahlow, Felipe Rodrigues, et al.
Pubblicazione: (2024)
GRIT: Teaching MLLMs to Think with Images
di: Fan, Yue, et al.
Pubblicazione: (2025)
di: Fan, Yue, et al.
Pubblicazione: (2025)
Evaluation Metrics for Text Data Augmentation in NLP
di: Amadeus, Marcellus, et al.
Pubblicazione: (2024)
di: Amadeus, Marcellus, et al.
Pubblicazione: (2024)
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
di: Albadarneh, Israa A., et al.
Pubblicazione: (2025)
di: Albadarneh, Israa A., et al.
Pubblicazione: (2025)
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
di: Zhang, Chenhui, et al.
Pubblicazione: (2024)
di: Zhang, Chenhui, et al.
Pubblicazione: (2024)
Phonetically rich corpus construction for a low-resourced language
di: Amadeus, Marcellus, et al.
Pubblicazione: (2024)
di: Amadeus, Marcellus, et al.
Pubblicazione: (2024)
Image captioning in different languages
di: van Miltenburg, Emiel
Pubblicazione: (2024)
di: van Miltenburg, Emiel
Pubblicazione: (2024)
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
di: Anugraha, David, et al.
Pubblicazione: (2025)
di: Anugraha, David, et al.
Pubblicazione: (2025)
Multi-Modal interpretable automatic video captioning
di: Hanna-Asaad, Antoine, et al.
Pubblicazione: (2024)
di: Hanna-Asaad, Antoine, et al.
Pubblicazione: (2024)
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
di: Du, Zhicheng, et al.
Pubblicazione: (2024)
di: Du, Zhicheng, et al.
Pubblicazione: (2024)
Object-oriented backdoor attack against image captioning
di: Li, Meiling, et al.
Pubblicazione: (2024)
di: Li, Meiling, et al.
Pubblicazione: (2024)
Leveraging LLMs for On-the-Fly Instruction Guided Image Editing
di: Santos, Rodrigo, et al.
Pubblicazione: (2024)
di: Santos, Rodrigo, et al.
Pubblicazione: (2024)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
di: Sharma, Aditya, et al.
Pubblicazione: (2024)
di: Sharma, Aditya, et al.
Pubblicazione: (2024)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
di: Fu, Xingyu, et al.
Pubblicazione: (2024)
di: Fu, Xingyu, et al.
Pubblicazione: (2024)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
di: Feng, Weixi, et al.
Pubblicazione: (2024)
di: Feng, Weixi, et al.
Pubblicazione: (2024)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
di: Li, Xiujun, et al.
Pubblicazione: (2023)
di: Li, Xiujun, et al.
Pubblicazione: (2023)
Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts
di: Rykov, Elisei, et al.
Pubblicazione: (2025)
di: Rykov, Elisei, et al.
Pubblicazione: (2025)
GaussianVision: Vision-Language Alignment from Compressed Image Representations using 2D Gaussian Splatting
di: Omri, Yasmine, et al.
Pubblicazione: (2025)
di: Omri, Yasmine, et al.
Pubblicazione: (2025)
Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)
di: Saxon, Michael, et al.
Pubblicazione: (2024)
di: Saxon, Michael, et al.
Pubblicazione: (2024)
Instruct-Imagen: Image Generation with Multi-modal Instruction
di: Hu, Hexiang, et al.
Pubblicazione: (2024)
di: Hu, Hexiang, et al.
Pubblicazione: (2024)
All in an Aggregated Image for In-Image Learning
di: Wang, Lei, et al.
Pubblicazione: (2024)
di: Wang, Lei, et al.
Pubblicazione: (2024)
Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation
di: Chen, Wenting, et al.
Pubblicazione: (2023)
di: Chen, Wenting, et al.
Pubblicazione: (2023)
Thinking with Generated Images
di: Chern, Ethan, et al.
Pubblicazione: (2025)
di: Chern, Ethan, et al.
Pubblicazione: (2025)
Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval
di: Liu, Delong, et al.
Pubblicazione: (2023)
di: Liu, Delong, et al.
Pubblicazione: (2023)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
On the Interplay of Human-AI Alignment,Fairness, and Performance Trade-offs in Medical Imaging
di: Luo, Haozhe, et al.
Pubblicazione: (2025)
di: Luo, Haozhe, et al.
Pubblicazione: (2025)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
di: Chen, Wenting, et al.
Pubblicazione: (2024)
di: Chen, Wenting, et al.
Pubblicazione: (2024)
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
di: Zhang, Leixin, et al.
Pubblicazione: (2024)
di: Zhang, Leixin, et al.
Pubblicazione: (2024)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
di: Ye, Junyan, et al.
Pubblicazione: (2025)
di: Ye, Junyan, et al.
Pubblicazione: (2025)
Text-only Synthesis for Image Captioning
di: Zhou, Qing, et al.
Pubblicazione: (2024)
di: Zhou, Qing, et al.
Pubblicazione: (2024)
Is Your Image a Good Storyteller?
di: Song, Xiujie, et al.
Pubblicazione: (2024)
di: Song, Xiujie, et al.
Pubblicazione: (2024)
NL-Eye: Abductive NLI for Images
di: Ventura, Mor, et al.
Pubblicazione: (2024)
di: Ventura, Mor, et al.
Pubblicazione: (2024)
Chatting with Images for Introspective Visual Thinking
di: Wu, Junfei, et al.
Pubblicazione: (2026)
di: Wu, Junfei, et al.
Pubblicazione: (2026)
The Role of Data Curation in Image Captioning
di: Li, Wenyan, et al.
Pubblicazione: (2023)
di: Li, Wenyan, et al.
Pubblicazione: (2023)
Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion
di: Allgeuer, Philipp, et al.
Pubblicazione: (2024)
di: Allgeuer, Philipp, et al.
Pubblicazione: (2024)
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
di: Yamabe, Shojiro, et al.
Pubblicazione: (2025)
di: Yamabe, Shojiro, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Amadeus-Verbo Technical Report: The powerful Qwen2.5 family models trained in Portuguese
di: Cruz-Castañeda, William Alberto, et al.
Pubblicazione: (2025) -
From Pampas to Pixels: Fine-Tuning Diffusion Models for Gaúcho Heritage
di: Amadeus, Marcellus, et al.
Pubblicazione: (2024) -
An Inpainting-Infused Pipeline for Attire and Background Replacement
di: Perche-Mahlow, Felipe Rodrigues, et al.
Pubblicazione: (2024) -
GRIT: Teaching MLLMs to Think with Images
di: Fan, Yue, et al.
Pubblicazione: (2025) -
Evaluation Metrics for Text Data Augmentation in NLP
di: Amadeus, Marcellus, et al.
Pubblicazione: (2024)