Saved in:
| Main Author: | Zhang, Shiwen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.21703 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
by: Cong, Yuren, et al.
Published: (2023)
by: Cong, Yuren, et al.
Published: (2023)
TurboEdit: Instant text-based image editing
by: Wu, Zongze, et al.
Published: (2024)
by: Wu, Zongze, et al.
Published: (2024)
Action-based image editing guided by human instructions
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
Can language-guided unsupervised adaptation improve medical image classification using unpaired images and texts?
by: Rahman, Umaima, et al.
Published: (2024)
by: Rahman, Umaima, et al.
Published: (2024)
Semantic-guided Fine-tuning of Foundation Model for Long-tailed Visual Recognition
by: Peng, Yufei, et al.
Published: (2025)
by: Peng, Yufei, et al.
Published: (2025)
Borrowing from anything: A generalizable framework for reference-guided instance editing
by: Zhou, Shengxiao, et al.
Published: (2025)
by: Zhou, Shengxiao, et al.
Published: (2025)
Exploring text-to-image generation for historical document image retrieval
by: Cote, Melissa, et al.
Published: (2025)
by: Cote, Melissa, et al.
Published: (2025)
Consistent text-to-image generation via scene de-contextualization
by: Tang, Song, et al.
Published: (2025)
by: Tang, Song, et al.
Published: (2025)
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
by: Sinha, Sanchit, et al.
Published: (2026)
by: Sinha, Sanchit, et al.
Published: (2026)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
LEAST: "Local" text-conditioned image style transfer
by: Singh, Silky, et al.
Published: (2024)
by: Singh, Silky, et al.
Published: (2024)
Forgedit: Text Guided Image Editing via Learning and Forgetting
by: Zhang, Shiwen, et al.
Published: (2023)
by: Zhang, Shiwen, et al.
Published: (2023)
EvRainDrop: HyperGraph-guided Completion for Effective Frame and Event Stream Aggregation
by: Wang, Futian, et al.
Published: (2025)
by: Wang, Futian, et al.
Published: (2025)
Exploring scalable medical image encoders beyond text supervision
by: Pérez-García, Fernando, et al.
Published: (2024)
by: Pérez-García, Fernando, et al.
Published: (2024)
Efficient scene text image super-resolution with semantic guidance
by: TomyEnrique, LeoWu, et al.
Published: (2024)
by: TomyEnrique, LeoWu, et al.
Published: (2024)
Ranking-aware adapter for text-driven image ordering with CLIP
by: Yu, Wei-Hsiang, et al.
Published: (2024)
by: Yu, Wei-Hsiang, et al.
Published: (2024)
IMMA: Immunizing text-to-image Models against Malicious Adaptation
by: Zheng, Amber Yijia, et al.
Published: (2023)
by: Zheng, Amber Yijia, et al.
Published: (2023)
HyperGCT: A Dynamic Hyper-GNN-Learned Geometric Constraint for 3D Registration
by: Zhang, Xiyu, et al.
Published: (2025)
by: Zhang, Xiyu, et al.
Published: (2025)
HyperDet: Generalizable Detection of Synthesized Images by Generating and Merging A Mixture of Hyper LoRAs
by: Cao, Huangsen, et al.
Published: (2024)
by: Cao, Huangsen, et al.
Published: (2024)
QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit
by: Zhang, Shiwen, et al.
Published: (2026)
by: Zhang, Shiwen, et al.
Published: (2026)
Visual question answering based evaluation metrics for text-to-image generation
by: Miyamoto, Mizuki, et al.
Published: (2024)
by: Miyamoto, Mizuki, et al.
Published: (2024)
Concept Corrector: Erase concepts on the fly for text-to-image diffusion models
by: Meng, Zheling, et al.
Published: (2025)
by: Meng, Zheling, et al.
Published: (2025)
IMAGE-ALCHEMY: Advancing subject fidelity in personalised text-to-image generation
by: Tiwari, Amritanshu, et al.
Published: (2025)
by: Tiwari, Amritanshu, et al.
Published: (2025)
Enriched text-guided variational multimodal knowledge distillation network (VMD) for automated diagnosis of plaque vulnerability in 3D carotid artery MRI
by: Cao, Bo, et al.
Published: (2025)
by: Cao, Bo, et al.
Published: (2025)
Rethink Sparse Signals for Pose-guided Text-to-image Generation
by: Xuan, Wenjie, et al.
Published: (2025)
by: Xuan, Wenjie, et al.
Published: (2025)
Trinity Detector:text-assisted and attention mechanisms based spectral fusion for diffusion generation image detection
by: Song, Jiawei, et al.
Published: (2024)
by: Song, Jiawei, et al.
Published: (2024)
HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation
by: Nguyen, Trong-Thuan, et al.
Published: (2024)
by: Nguyen, Trong-Thuan, et al.
Published: (2024)
HyperSpaceX: Radial and Angular Exploration of HyperSpherical Dimensions
by: Chiranjeev, Chiranjeev, et al.
Published: (2024)
by: Chiranjeev, Chiranjeev, et al.
Published: (2024)
HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion
by: Liu, Xian, et al.
Published: (2023)
by: Liu, Xian, et al.
Published: (2023)
CDST: Color Disentangled Style Transfer for Universal Style Reference Customization
by: Zhang, Shiwen, et al.
Published: (2025)
by: Zhang, Shiwen, et al.
Published: (2025)
Dark Miner: Defend against undesirable generation for text-to-image diffusion models
by: Meng, Zheling, et al.
Published: (2024)
by: Meng, Zheling, et al.
Published: (2024)
PointT2I: LLM-based text-to-image generation via keypoints
by: Lee, Taekyung, et al.
Published: (2025)
by: Lee, Taekyung, et al.
Published: (2025)
HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion
by: He, Yu, et al.
Published: (2026)
by: He, Yu, et al.
Published: (2026)
A large-scale image-text dataset benchmark for farmland segmentation
by: Tao, Chao, et al.
Published: (2025)
by: Tao, Chao, et al.
Published: (2025)
HyperDefect-YOLO: Enhance YOLO with HyperGraph Computation for Industrial Defect Detection
by: Zuo, Zuo, et al.
Published: (2024)
by: Zuo, Zuo, et al.
Published: (2024)
HyperSLICE: HyperBand optimized Spiral for Low-latency Interactive Cardiac Examination
by: Jaubert, Olivier, et al.
Published: (2023)
by: Jaubert, Olivier, et al.
Published: (2023)
Gaussian highpass guided image filtering
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
Bio-inspired fine-tuning for selective transfer learning in image classification
by: Davila, Ana, et al.
Published: (2026)
by: Davila, Ana, et al.
Published: (2026)
HyperSpace: Hypernetworks for spacing-adaptive image segmentation
by: Joutard, Samuel, et al.
Published: (2024)
by: Joutard, Samuel, et al.
Published: (2024)
Hyper-Transformer for Amodal Completion
by: Gao, Jianxiong, et al.
Published: (2024)
by: Gao, Jianxiong, et al.
Published: (2024)
Similar Items
-
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
by: Cong, Yuren, et al.
Published: (2023) -
TurboEdit: Instant text-based image editing
by: Wu, Zongze, et al.
Published: (2024) -
Action-based image editing guided by human instructions
by: Trusca, Maria Mihaela, et al.
Published: (2024) -
Can language-guided unsupervised adaptation improve medical image classification using unpaired images and texts?
by: Rahman, Umaima, et al.
Published: (2024) -
Semantic-guided Fine-tuning of Foundation Model for Long-tailed Visual Recognition
by: Peng, Yufei, et al.
Published: (2025)