Salvato in:
| Autori principali: | Xie, Xudong, Li, Yuzhe, Liu, Yang, Zhang, Zhifei, Wang, Zhaowen, Xiong, Wei, Bai, Xiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2408.00106 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Intelligent Artistic Typography: A Comprehensive Review of Artistic Text Design and Generation
di: Bai, Yuhang, et al.
Pubblicazione: (2024)
di: Bai, Yuhang, et al.
Pubblicazione: (2024)
DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models
di: Dong, Zhe, et al.
Pubblicazione: (2025)
di: Dong, Zhe, et al.
Pubblicazione: (2025)
Learning to Manipulate Artistic Images
di: Guo, Wei, et al.
Pubblicazione: (2024)
di: Guo, Wei, et al.
Pubblicazione: (2024)
Interact-Custom: Customized Human Object Interaction Image Generation
di: Xu, Zhu, et al.
Pubblicazione: (2025)
di: Xu, Zhu, et al.
Pubblicazione: (2025)
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
di: Liu, Yuliang, et al.
Pubblicazione: (2024)
di: Liu, Yuliang, et al.
Pubblicazione: (2024)
Rethinking Artistic Copyright Infringements in the Era of Text-to-Image Generative Models
di: Moayeri, Mazda, et al.
Pubblicazione: (2024)
di: Moayeri, Mazda, et al.
Pubblicazione: (2024)
VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models
di: Feng, Kailai, et al.
Pubblicazione: (2024)
di: Feng, Kailai, et al.
Pubblicazione: (2024)
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
di: Xie, Xudong, et al.
Pubblicazione: (2024)
di: Xie, Xudong, et al.
Pubblicazione: (2024)
GCA-SUNet: A Gated Context-Aware Swin-UNet for Exemplar-Free Counting
di: Wu, Yuzhe, et al.
Pubblicazione: (2024)
di: Wu, Yuzhe, et al.
Pubblicazione: (2024)
StyleForge: Enhancing Text-to-Image Synthesis for Any Artistic Styles with Dual Binding
di: Park, Junseo, et al.
Pubblicazione: (2024)
di: Park, Junseo, et al.
Pubblicazione: (2024)
Enhanced Semantic Segmentation Pipeline for WeatherProof Dataset Challenge
di: Zhang, Nan, et al.
Pubblicazione: (2024)
di: Zhang, Nan, et al.
Pubblicazione: (2024)
Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation
di: Dong, Zhe, et al.
Pubblicazione: (2024)
di: Dong, Zhe, et al.
Pubblicazione: (2024)
CroBIM-U: Uncertainty-Driven Referring Remote Sensing Image Segmentation
di: Sun, Yuzhe, et al.
Pubblicazione: (2026)
di: Sun, Yuzhe, et al.
Pubblicazione: (2026)
Coarse-to-Fine Structure-Aware Artistic Style Transfer
di: Liu, Kunxiao, et al.
Pubblicazione: (2025)
di: Liu, Kunxiao, et al.
Pubblicazione: (2025)
ALPS: An Auto-Labeling and Pre-training Scheme for Remote Sensing Segmentation With Segment Anything Model
di: Zhang, Song, et al.
Pubblicazione: (2024)
di: Zhang, Song, et al.
Pubblicazione: (2024)
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
di: Li, Zhang, et al.
Pubblicazione: (2025)
di: Li, Zhang, et al.
Pubblicazione: (2025)
15M Multimodal Facial Image-Text Dataset
di: Dai, Dawei, et al.
Pubblicazione: (2024)
di: Dai, Dawei, et al.
Pubblicazione: (2024)
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
di: Fu, Ling, et al.
Pubblicazione: (2024)
di: Fu, Ling, et al.
Pubblicazione: (2024)
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
di: Li, Zhang, et al.
Pubblicazione: (2023)
di: Li, Zhang, et al.
Pubblicazione: (2023)
ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
A Text-Image Fusion Method with Data Augmentation Capabilities for Referring Medical Image Segmentation
di: Chai, Shurong, et al.
Pubblicazione: (2025)
di: Chai, Shurong, et al.
Pubblicazione: (2025)
An Exceptional Dataset For Rare Pancreatic Tumor Segmentation
di: Li, Wenqi, et al.
Pubblicazione: (2025)
di: Li, Wenqi, et al.
Pubblicazione: (2025)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
di: Chen, Haoyu, et al.
Pubblicazione: (2026)
di: Chen, Haoyu, et al.
Pubblicazione: (2026)
POSTA: A Go-to Framework for Customized Artistic Poster Generation
di: Chen, Haoyu, et al.
Pubblicazione: (2025)
di: Chen, Haoyu, et al.
Pubblicazione: (2025)
CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis with Multimodal Diffusion
di: Huang, Nisha, et al.
Pubblicazione: (2024)
di: Huang, Nisha, et al.
Pubblicazione: (2024)
SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing
di: Gu, Jing, et al.
Pubblicazione: (2024)
di: Gu, Jing, et al.
Pubblicazione: (2024)
Enhancing Boundary Segmentation for Topological Accuracy with Skeleton-based Methods
di: Liu, Chuni, et al.
Pubblicazione: (2024)
di: Liu, Chuni, et al.
Pubblicazione: (2024)
StageDesigner: Artistic Stage Generation for Scenography via Theater Scripts
di: Gan, Zhaoxing, et al.
Pubblicazione: (2025)
di: Gan, Zhaoxing, et al.
Pubblicazione: (2025)
Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
di: Jiang, Chen, et al.
Pubblicazione: (2023)
di: Jiang, Chen, et al.
Pubblicazione: (2023)
Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation
di: Xie, Dian, et al.
Pubblicazione: (2026)
di: Xie, Dian, et al.
Pubblicazione: (2026)
TGC-Net: A Structure-Aware and Semantically-Aligned Framework for Text-Guided Medical Image Segmentation
di: Lin, Gaoren, et al.
Pubblicazione: (2025)
di: Lin, Gaoren, et al.
Pubblicazione: (2025)
MASTER: Multimodal Segmentation with Text Prompts
di: Liu, Fuyang, et al.
Pubblicazione: (2025)
di: Liu, Fuyang, et al.
Pubblicazione: (2025)
Ideal Registration? Segmentation is All You Need
di: Chen, Xiang, et al.
Pubblicazione: (2025)
di: Chen, Xiang, et al.
Pubblicazione: (2025)
Graph Relation Distillation for Efficient Biomedical Instance Segmentation
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection
di: Jia, Zexi, et al.
Pubblicazione: (2025)
di: Jia, Zexi, et al.
Pubblicazione: (2025)
Towards Universal Text-driven CT Image Segmentation
di: Li, Yuheng, et al.
Pubblicazione: (2025)
di: Li, Yuheng, et al.
Pubblicazione: (2025)
Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation
di: Ji, Haoyu, et al.
Pubblicazione: (2025)
di: Ji, Haoyu, et al.
Pubblicazione: (2025)
Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
di: Jisheng, Dang, et al.
Pubblicazione: (2025)
di: Jisheng, Dang, et al.
Pubblicazione: (2025)
Harnessing Group-Oriented Consistency Constraints for Semi-Supervised Semantic Segmentation in CdZnTe Semiconductors
di: Li, Peihao, et al.
Pubblicazione: (2025)
di: Li, Peihao, et al.
Pubblicazione: (2025)
Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models
di: Wang, Xiaoyan, et al.
Pubblicazione: (2025)
di: Wang, Xiaoyan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Intelligent Artistic Typography: A Comprehensive Review of Artistic Text Design and Generation
di: Bai, Yuhang, et al.
Pubblicazione: (2024) -
DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models
di: Dong, Zhe, et al.
Pubblicazione: (2025) -
Learning to Manipulate Artistic Images
di: Guo, Wei, et al.
Pubblicazione: (2024) -
Interact-Custom: Customized Human Object Interaction Image Generation
di: Xu, Zhu, et al.
Pubblicazione: (2025) -
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
di: Liu, Yuliang, et al.
Pubblicazione: (2024)