TexTailor: Customized Text-aligned Texturing via Effective Resampling
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Suin, Kim, Dae-Shik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based Customization
di: Liu, Yisu, et al.
Pubblicazione: (2024)
di: Liu, Yisu, et al.
Pubblicazione: (2024)
Sora as a World Model? A Complete Survey on Text-to-Video Generation
di: Puspitasari, Fachrina Dewi, et al.
Pubblicazione: (2024)
di: Puspitasari, Fachrina Dewi, et al.
Pubblicazione: (2024)
From Prompt to Production:Automating Brand-Safe Marketing Imagery with Text-to-Image Models
di: Atighehchian, Parmida, et al.
Pubblicazione: (2026)
di: Atighehchian, Parmida, et al.
Pubblicazione: (2026)
Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation
di: Agarwal, Rachit, et al.
Pubblicazione: (2026)
di: Agarwal, Rachit, et al.
Pubblicazione: (2026)
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2025)
3D Adaptive Structural Convolution Network for Domain-Invariant Point Cloud Recognition
di: Kim, Younggun, et al.
Pubblicazione: (2024)
di: Kim, Younggun, et al.
Pubblicazione: (2024)
Invariant Representation via Decoupling Style and Spurious Features from Images
di: Li, Ruimeng, et al.
Pubblicazione: (2023)
di: Li, Ruimeng, et al.
Pubblicazione: (2023)
SITUATE -- Synthetic Object Counting Dataset for VLM training
di: Peinl, René, et al.
Pubblicazione: (2026)
di: Peinl, René, et al.
Pubblicazione: (2026)
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
Image Segmentation and Classification of E-waste for Training Robots for Waste Segregation
di: Tripathi, Prakriti
Pubblicazione: (2025)
di: Tripathi, Prakriti
Pubblicazione: (2025)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
di: Holm, Felix, et al.
Pubblicazione: (2025)
di: Holm, Felix, et al.
Pubblicazione: (2025)
Siamese Networks for Cat Re-Identification: Exploring Neural Models for Cat Instance Recognition
di: Trein, Tobias, et al.
Pubblicazione: (2025)
di: Trein, Tobias, et al.
Pubblicazione: (2025)
Evaluation of Environmental Conditions on Object Detection using Oriented Bounding Boxes for AR Applications
di: Li, Vladislav, et al.
Pubblicazione: (2023)
di: Li, Vladislav, et al.
Pubblicazione: (2023)
Appearance-based gaze estimation enhanced with synthetic images using deep neural networks
di: Herashchenko, Dmytro, et al.
Pubblicazione: (2023)
di: Herashchenko, Dmytro, et al.
Pubblicazione: (2023)
MFTF: Mask-free Training-free Object Level Layout Control Diffusion Model
di: Yang, Shan
Pubblicazione: (2024)
di: Yang, Shan
Pubblicazione: (2024)
Attentive VQ-VAE
di: Hoyos, Angello, et al.
Pubblicazione: (2023)
di: Hoyos, Angello, et al.
Pubblicazione: (2023)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
di: Chen, Zhangquan, et al.
Pubblicazione: (2026)
di: Chen, Zhangquan, et al.
Pubblicazione: (2026)
CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier
di: Ou, Ziyang
Pubblicazione: (2025)
di: Ou, Ziyang
Pubblicazione: (2025)
Rethinking Multimodal Point Cloud Completion: A Completion-by-Correction Perspective
di: Luo, Wang, et al.
Pubblicazione: (2025)
di: Luo, Wang, et al.
Pubblicazione: (2025)
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
di: Safdar, Aon, et al.
Pubblicazione: (2025)
di: Safdar, Aon, et al.
Pubblicazione: (2025)
Next-Generation License Plate Detection and Recognition System using YOLOv8
di: Amin, Arslan, et al.
Pubblicazione: (2025)
di: Amin, Arslan, et al.
Pubblicazione: (2025)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
di: Chen, Yiping, et al.
Pubblicazione: (2026)
di: Chen, Yiping, et al.
Pubblicazione: (2026)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
di: He, Jianxiang, et al.
Pubblicazione: (2025)
di: He, Jianxiang, et al.
Pubblicazione: (2025)
Instruction-based Image Editing with Planning, Reasoning, and Generation
di: Ji, Liya, et al.
Pubblicazione: (2026)
di: Ji, Liya, et al.
Pubblicazione: (2026)
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
di: Feng, Yigui, et al.
Pubblicazione: (2026)
di: Feng, Yigui, et al.
Pubblicazione: (2026)
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
di: Ma, Jie, et al.
Pubblicazione: (2023)
di: Ma, Jie, et al.
Pubblicazione: (2023)
Unified Auto-Encoding with Masked Diffusion
di: Hansen-Estruch, Philippe, et al.
Pubblicazione: (2024)
di: Hansen-Estruch, Philippe, et al.
Pubblicazione: (2024)
GraphTEN: Graph Enhanced Texture Encoding Network
di: Peng, Bo, et al.
Pubblicazione: (2025)
di: Peng, Bo, et al.
Pubblicazione: (2025)
Supervised Contrastive Learning for Few-Shot AI-Generated Image Detection and Attribution
di: Urueña, Jaime Álvarez, et al.
Pubblicazione: (2025)
di: Urueña, Jaime Álvarez, et al.
Pubblicazione: (2025)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
di: Ge, Shiping, et al.
Pubblicazione: (2024)
di: Ge, Shiping, et al.
Pubblicazione: (2024)
Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities
di: Mobbs, Rebecca, et al.
Pubblicazione: (2025)
di: Mobbs, Rebecca, et al.
Pubblicazione: (2025)
Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
di: He, Mengqi, et al.
Pubblicazione: (2026)
di: He, Mengqi, et al.
Pubblicazione: (2026)
HAAP: Vision-context Hierarchical Attention Autoregressive with Adaptive Permutation for Scene Text Recognition
di: Chen, Honghui, et al.
Pubblicazione: (2024)
di: Chen, Honghui, et al.
Pubblicazione: (2024)
Deformation-Free Cross-Domain Image Registration via Position-Encoded Temporal Attention
di: Wang, Yiwen, et al.
Pubblicazione: (2026)
di: Wang, Yiwen, et al.
Pubblicazione: (2026)
THIRDEYE: Cue-Aware Monocular Depth Estimation via Brain-Inspired Multi-Stage Fusion
di: Ioan, Calin Teodor
Pubblicazione: (2025)
di: Ioan, Calin Teodor
Pubblicazione: (2025)
Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments
di: Gonzalez, Laura Alejandra Encinar, et al.
Pubblicazione: (2025)
di: Gonzalez, Laura Alejandra Encinar, et al.
Pubblicazione: (2025)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
di: Wang, Zhaohui, et al.
Pubblicazione: (2025)
di: Wang, Zhaohui, et al.
Pubblicazione: (2025)
Deep Learning methodology for the identification of wood species using high-resolution macroscopic images
di: Herrera-Poyatos, David, et al.
Pubblicazione: (2024)
di: Herrera-Poyatos, David, et al.
Pubblicazione: (2024)
Vectra: A New Metric, Dataset, and Model for Visual Quality Assessment in E-Commerce In-Image Machine Translation
di: Wu, Qingyu, et al.
Pubblicazione: (2026)
di: Wu, Qingyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based Customization
di: Liu, Yisu, et al.
Pubblicazione: (2024) -
Sora as a World Model? A Complete Survey on Text-to-Video Generation
di: Puspitasari, Fachrina Dewi, et al.
Pubblicazione: (2024) -
From Prompt to Production:Automating Brand-Safe Marketing Imagery with Text-to-Image Models
di: Atighehchian, Parmida, et al.
Pubblicazione: (2026) -
Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation
di: Agarwal, Rachit, et al.
Pubblicazione: (2026) -
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
di: Ramakrishnan, Aashish Anantha, et al.
Pubblicazione: (2025)