From Pampas to Pixels: Fine-Tuning Diffusion Models for Gaúcho Heritage
Fuente:
arXiv
Guardado en:
| Autores principales: | Amadeus, Marcellus, Castañeda, William Alberto Cruz, Zanella, André Felipe, Mahlow, Felipe Rodrigues Perche |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Inpainting-Infused Pipeline for Attire and Background Replacement
por: Perche-Mahlow, Felipe Rodrigues, et al.
Publicado: (2024)
por: Perche-Mahlow, Felipe Rodrigues, et al.
Publicado: (2024)
Image captioning for Brazilian Portuguese using GRIT model
por: de Alencar, Rafael Silva, et al.
Publicado: (2024)
por: de Alencar, Rafael Silva, et al.
Publicado: (2024)
Evaluation Metrics for Text Data Augmentation in NLP
por: Amadeus, Marcellus, et al.
Publicado: (2024)
por: Amadeus, Marcellus, et al.
Publicado: (2024)
Amadeus-Verbo Technical Report: The powerful Qwen2.5 family models trained in Portuguese
por: Cruz-Castañeda, William Alberto, et al.
Publicado: (2025)
por: Cruz-Castañeda, William Alberto, et al.
Publicado: (2025)
Phonetically rich corpus construction for a low-resourced language
por: Amadeus, Marcellus, et al.
Publicado: (2024)
por: Amadeus, Marcellus, et al.
Publicado: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
por: Lu, Yujie, et al.
Publicado: (2024)
por: Lu, Yujie, et al.
Publicado: (2024)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
por: Yang, Cheng, et al.
Publicado: (2026)
por: Yang, Cheng, et al.
Publicado: (2026)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
por: Irawan, Patrick Amadeus, et al.
Publicado: (2024)
por: Irawan, Patrick Amadeus, et al.
Publicado: (2024)
From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
por: Cai, Zefan, et al.
Publicado: (2025)
por: Cai, Zefan, et al.
Publicado: (2025)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
por: Shang, Yuying, et al.
Publicado: (2024)
por: Shang, Yuying, et al.
Publicado: (2024)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
por: Yuan, Huizhuo, et al.
Publicado: (2024)
por: Yuan, Huizhuo, et al.
Publicado: (2024)
Expert Pyramid Tuning: Efficient Parameter Fine-Tuning for Expertise-Driven Task Allocation
por: Zhang, Jia-Chen, et al.
Publicado: (2026)
por: Zhang, Jia-Chen, et al.
Publicado: (2026)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
por: Irawan, Patrick Amadeus, et al.
Publicado: (2026)
por: Irawan, Patrick Amadeus, et al.
Publicado: (2026)
Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning
por: Cao, Qinglong, et al.
Publicado: (2026)
por: Cao, Qinglong, et al.
Publicado: (2026)
Beyond Accuracy Optimization: Computer Vision Losses for Large Language Model Fine-Tuning
por: Cambrin, Daniele Rege, et al.
Publicado: (2024)
por: Cambrin, Daniele Rege, et al.
Publicado: (2024)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
por: Lyu, Zhiheng, et al.
Publicado: (2025)
por: Lyu, Zhiheng, et al.
Publicado: (2025)
PEFT A2Z: Parameter-Efficient Fine-Tuning Survey for Large Language and Vision Models
por: Prottasha, Nusrat Jahan, et al.
Publicado: (2025)
por: Prottasha, Nusrat Jahan, et al.
Publicado: (2025)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
por: Tan, Zhiyu, et al.
Publicado: (2024)
por: Tan, Zhiyu, et al.
Publicado: (2024)
Rethinking Weight Decay for Robust Fine-Tuning of Foundation Models
por: Tian, Junjiao, et al.
Publicado: (2024)
por: Tian, Junjiao, et al.
Publicado: (2024)
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
por: Ding, Yi, et al.
Publicado: (2025)
por: Ding, Yi, et al.
Publicado: (2025)
Vision Language Models are Confused Tourists
por: Irawan, Patrick Amadeus, et al.
Publicado: (2025)
por: Irawan, Patrick Amadeus, et al.
Publicado: (2025)
Pixel Sentence Representation Learning
por: Xiao, Chenghao, et al.
Publicado: (2024)
por: Xiao, Chenghao, et al.
Publicado: (2024)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
por: Wei, Xinyu, et al.
Publicado: (2025)
por: Wei, Xinyu, et al.
Publicado: (2025)
SemiHVision: Enhancing Medical Multimodal Models with a Semi-Human Annotated Dataset and Fine-Tuned Instruction Generation
por: Wang, Junda, et al.
Publicado: (2024)
por: Wang, Junda, et al.
Publicado: (2024)
Autoregressive Pre-Training on Pixels and Texts
por: Chai, Yekun, et al.
Publicado: (2024)
por: Chai, Yekun, et al.
Publicado: (2024)
Talking Points: Describing and Localizing Pixels
por: Rusanovsky, Matan, et al.
Publicado: (2025)
por: Rusanovsky, Matan, et al.
Publicado: (2025)
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
por: Sun, Haoyuan, et al.
Publicado: (2025)
por: Sun, Haoyuan, et al.
Publicado: (2025)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
por: Yanuka, Moran, et al.
Publicado: (2024)
por: Yanuka, Moran, et al.
Publicado: (2024)
From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Models
por: Huang, Kung-Hsiang, et al.
Publicado: (2024)
por: Huang, Kung-Hsiang, et al.
Publicado: (2024)
Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
por: Masri, Sari, et al.
Publicado: (2025)
por: Masri, Sari, et al.
Publicado: (2025)
Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On
por: Goldshmidt, Roni
Publicado: (2025)
por: Goldshmidt, Roni
Publicado: (2025)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
por: Li, Xiujun, et al.
Publicado: (2023)
por: Li, Xiujun, et al.
Publicado: (2023)
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models
por: Zhu, Wenhui, et al.
Publicado: (2025)
por: Zhu, Wenhui, et al.
Publicado: (2025)
From Pixels to Prose: A Large Dataset of Dense Image Captions
por: Singla, Vasu, et al.
Publicado: (2024)
por: Singla, Vasu, et al.
Publicado: (2024)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
por: Gondal, Moazzam Umer, et al.
Publicado: (2025)
por: Gondal, Moazzam Umer, et al.
Publicado: (2025)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
por: Zhang, Wanpeng, et al.
Publicado: (2024)
por: Zhang, Wanpeng, et al.
Publicado: (2024)
TuneVLSeg: Prompt Tuning Benchmark for Vision-Language Segmentation Models
por: Adhikari, Rabin, et al.
Publicado: (2024)
por: Adhikari, Rabin, et al.
Publicado: (2024)
Recovering the Pre-Fine-Tuning Weights of Generative Models
por: Horwitz, Eliahu, et al.
Publicado: (2024)
por: Horwitz, Eliahu, et al.
Publicado: (2024)
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
por: Cai, Dexian, et al.
Publicado: (2025)
por: Cai, Dexian, et al.
Publicado: (2025)
Directional Gradient Projection for Robust Fine-Tuning of Foundation Models
por: Huang, Chengyue, et al.
Publicado: (2025)
por: Huang, Chengyue, et al.
Publicado: (2025)
Ejemplares similares
-
An Inpainting-Infused Pipeline for Attire and Background Replacement
por: Perche-Mahlow, Felipe Rodrigues, et al.
Publicado: (2024) -
Image captioning for Brazilian Portuguese using GRIT model
por: de Alencar, Rafael Silva, et al.
Publicado: (2024) -
Evaluation Metrics for Text Data Augmentation in NLP
por: Amadeus, Marcellus, et al.
Publicado: (2024) -
Amadeus-Verbo Technical Report: The powerful Qwen2.5 family models trained in Portuguese
por: Cruz-Castañeda, William Alberto, et al.
Publicado: (2025) -
Phonetically rich corpus construction for a low-resourced language
por: Amadeus, Marcellus, et al.
Publicado: (2024)