An Inpainting-Infused Pipeline for Attire and Background Replacement
Fuente:
arXiv
Guardado en:
| Autores principales: | Perche-Mahlow, Felipe Rodrigues, Felipe-Zanella, André, Cruz-Castañeda, William Alberto, Amadeus, Marcellus |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Pampas to Pixels: Fine-Tuning Diffusion Models for Gaúcho Heritage
por: Amadeus, Marcellus, et al.
Publicado: (2024)
por: Amadeus, Marcellus, et al.
Publicado: (2024)
Image captioning for Brazilian Portuguese using GRIT model
por: de Alencar, Rafael Silva, et al.
Publicado: (2024)
por: de Alencar, Rafael Silva, et al.
Publicado: (2024)
Evaluation Metrics for Text Data Augmentation in NLP
por: Amadeus, Marcellus, et al.
Publicado: (2024)
por: Amadeus, Marcellus, et al.
Publicado: (2024)
Amadeus-Verbo Technical Report: The powerful Qwen2.5 family models trained in Portuguese
por: Cruz-Castañeda, William Alberto, et al.
Publicado: (2025)
por: Cruz-Castañeda, William Alberto, et al.
Publicado: (2025)
Phonetically rich corpus construction for a low-resourced language
por: Amadeus, Marcellus, et al.
Publicado: (2024)
por: Amadeus, Marcellus, et al.
Publicado: (2024)
A Culturally-Aware Benchmark for Person Re-Identification in Modest Attire
por: Moghaddam, Alireza Sedighi, et al.
Publicado: (2024)
por: Moghaddam, Alireza Sedighi, et al.
Publicado: (2024)
Fill in the ____ (a Diffusion-based Image Inpainting Pipeline)
por: Gebre, Eyoel, et al.
Publicado: (2024)
por: Gebre, Eyoel, et al.
Publicado: (2024)
DreamPainter: Image Background Inpainting for E-commerce Scenarios
por: Zhao, Sijie, et al.
Publicado: (2025)
por: Zhao, Sijie, et al.
Publicado: (2025)
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
por: Anugraha, David, et al.
Publicado: (2025)
por: Anugraha, David, et al.
Publicado: (2025)
Evaluating the Clinical Impact of Generative Inpainting on Bone Age Estimation
por: Matsuoka, Felipe Akio, et al.
Publicado: (2025)
por: Matsuoka, Felipe Akio, et al.
Publicado: (2025)
An Evaluation Framework for Product Images Background Inpainting based on Human Feedback and Product Consistency
por: Liang, Yuqi, et al.
Publicado: (2024)
por: Liang, Yuqi, et al.
Publicado: (2024)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
por: Cho, Jaemin, et al.
Publicado: (2023)
por: Cho, Jaemin, et al.
Publicado: (2023)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
por: Han, Donghoon, et al.
Publicado: (2024)
por: Han, Donghoon, et al.
Publicado: (2024)
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
por: Zeng, Ziyun, et al.
Publicado: (2026)
por: Zeng, Ziyun, et al.
Publicado: (2026)
AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any Scenario
por: Li, Yuhan, et al.
Publicado: (2024)
por: Li, Yuhan, et al.
Publicado: (2024)
Assessing Image Inpainting via Re-Inpainting Self-Consistency Evaluation
por: Chen, Tianyi, et al.
Publicado: (2024)
por: Chen, Tianyi, et al.
Publicado: (2024)
DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines
por: Cardoso, Gabriel Pimenta de Freitas, et al.
Publicado: (2026)
por: Cardoso, Gabriel Pimenta de Freitas, et al.
Publicado: (2026)
COCO-Inpaint: A Benchmark for Detecting and Localizing Inpainting-Based Image Manipulations
por: Yan, Haozhen, et al.
Publicado: (2025)
por: Yan, Haozhen, et al.
Publicado: (2025)
InpaintSLat: Inpainting Structured 3D Latents via Initial Noise Optimization
por: Chung, Jaeyoung, et al.
Publicado: (2026)
por: Chung, Jaeyoung, et al.
Publicado: (2026)
Semantic Matters: Multimodal Features for Affective Analysis
por: Hallmen, Tobias, et al.
Publicado: (2025)
por: Hallmen, Tobias, et al.
Publicado: (2025)
MINT: Memory-Infused Prompt Tuning at Test-time for CLIP
por: Yi, Jiaming, et al.
Publicado: (2025)
por: Yi, Jiaming, et al.
Publicado: (2025)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
por: Sharma, Aditya, et al.
Publicado: (2024)
por: Sharma, Aditya, et al.
Publicado: (2024)
Sample-efficient Integration of New Modalities into Large Language Models
por: İnce, Osman Batur, et al.
Publicado: (2025)
por: İnce, Osman Batur, et al.
Publicado: (2025)
DAOVI: Distortion-Aware Omnidirectional Video Inpainting
por: Seshimo, Ryosuke, et al.
Publicado: (2025)
por: Seshimo, Ryosuke, et al.
Publicado: (2025)
How Do Inpainting Artifacts Propagate to Language?
por: Yashwante, Pratham, et al.
Publicado: (2026)
por: Yashwante, Pratham, et al.
Publicado: (2026)
VITED: Video Temporal Evidence Distillation
por: Lu, Yujie, et al.
Publicado: (2025)
por: Lu, Yujie, et al.
Publicado: (2025)
Audio-Infused Automatic Image Colorization by Exploiting Audio Scene Semantics
por: Zhao, Pengcheng, et al.
Publicado: (2024)
por: Zhao, Pengcheng, et al.
Publicado: (2024)
Automatic Layout Planning for Visually-Rich Documents with Instruction-Following Models
por: Zhu, Wanrong, et al.
Publicado: (2024)
por: Zhu, Wanrong, et al.
Publicado: (2024)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
por: Fu, Xingyu, et al.
Publicado: (2024)
por: Fu, Xingyu, et al.
Publicado: (2024)
Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)
por: Saxon, Michael, et al.
Publicado: (2024)
por: Saxon, Michael, et al.
Publicado: (2024)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
por: Li, Xiujun, et al.
Publicado: (2023)
por: Li, Xiujun, et al.
Publicado: (2023)
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
por: Winata, Genta Indra, et al.
Publicado: (2024)
por: Winata, Genta Indra, et al.
Publicado: (2024)
Raformer: Redundancy-Aware Transformer for Video Wire Inpainting
por: Ji, Zhong, et al.
Publicado: (2024)
por: Ji, Zhong, et al.
Publicado: (2024)
3D-Consistent Image Inpainting with Diffusion Models
por: Antsfeld, Leonid, et al.
Publicado: (2024)
por: Antsfeld, Leonid, et al.
Publicado: (2024)
BrushEdit: All-In-One Image Inpainting and Editing
por: Li, Yaowei, et al.
Publicado: (2024)
por: Li, Yaowei, et al.
Publicado: (2024)
Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation
por: Buburuzan, Alexandru
Publicado: (2025)
por: Buburuzan, Alexandru
Publicado: (2025)
Photorealistic Inpainting for Perturbation-based Explanations in Ecological Monitoring
por: Aghakishiyeva, Günel, et al.
Publicado: (2025)
por: Aghakishiyeva, Günel, et al.
Publicado: (2025)
VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
por: Schumann, Raphael, et al.
Publicado: (2023)
por: Schumann, Raphael, et al.
Publicado: (2023)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
por: Lu, Yujie, et al.
Publicado: (2024)
por: Lu, Yujie, et al.
Publicado: (2024)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
por: Feng, Weixi, et al.
Publicado: (2024)
por: Feng, Weixi, et al.
Publicado: (2024)
Ejemplares similares
-
From Pampas to Pixels: Fine-Tuning Diffusion Models for Gaúcho Heritage
por: Amadeus, Marcellus, et al.
Publicado: (2024) -
Image captioning for Brazilian Portuguese using GRIT model
por: de Alencar, Rafael Silva, et al.
Publicado: (2024) -
Evaluation Metrics for Text Data Augmentation in NLP
por: Amadeus, Marcellus, et al.
Publicado: (2024) -
Amadeus-Verbo Technical Report: The powerful Qwen2.5 family models trained in Portuguese
por: Cruz-Castañeda, William Alberto, et al.
Publicado: (2025) -
Phonetically rich corpus construction for a low-resourced language
por: Amadeus, Marcellus, et al.
Publicado: (2024)