VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zhipeng, Yang, Lan, Qi, Yonggang, Zhang, Honggang, Pang, Kaiyue, Li, Ke, Song, Yi-Zhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neural Fields for 3D Tracking of Anatomy and Surgical Instruments in Monocular Laparoscopic Video Clips
by: Gerats, Beerend G. A., et al.
Published: (2024)
by: Gerats, Beerend G. A., et al.
Published: (2024)
Learning to Expand Images for Efficient Visual Autoregressive Modeling
by: Yang, Ruiqing, et al.
Published: (2025)
by: Yang, Ruiqing, et al.
Published: (2025)
Improving Visual Object Tracking through Visual Prompting
by: Chen, Shih-Fang, et al.
Published: (2024)
by: Chen, Shih-Fang, et al.
Published: (2024)
Denoising VAE as an Explainable Feature Reduction and Diagnostic Pipeline for Autism Based on Resting state fMRI
by: Zheng, Xinyuan, et al.
Published: (2024)
by: Zheng, Xinyuan, et al.
Published: (2024)
EDSNet: Efficient-DSNet for Video Summarization
by: Prasad, Ashish, et al.
Published: (2024)
by: Prasad, Ashish, et al.
Published: (2024)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
by: Shahin, Nada, et al.
Published: (2025)
by: Shahin, Nada, et al.
Published: (2025)
BG-YOLO: A Bidirectional-Guided Method for Underwater Object Detection
by: Zhang, Jian, et al.
Published: (2024)
by: Zhang, Jian, et al.
Published: (2024)
AVadCLIP: Audio-Visual Collaboration for Robust Video Anomaly Detection
by: Wu, Peng, et al.
Published: (2025)
by: Wu, Peng, et al.
Published: (2025)
Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection
by: Maity, Subhajit, et al.
Published: (2025)
by: Maity, Subhajit, et al.
Published: (2025)
Multimodal-to-Text Prompt Engineering in Large Language Models Using Feature Embeddings for GNSS Interference Characterization
by: Manjunath, Harshith, et al.
Published: (2025)
by: Manjunath, Harshith, et al.
Published: (2025)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
by: Shahin, Nada, et al.
Published: (2025)
by: Shahin, Nada, et al.
Published: (2025)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
by: Radwan, Ahmed, et al.
Published: (2024)
by: Radwan, Ahmed, et al.
Published: (2024)
IAMAP: Unlocking Deep Learning in QGIS for non-coders and limited computing resources
by: Tresson, Paul, et al.
Published: (2025)
by: Tresson, Paul, et al.
Published: (2025)
Optimizing Multi-Scale Representations to Detect Effect Heterogeneity Using Earth Observation and Computer Vision: Applications to Two Anti-Poverty RCTs
by: Zhu, Fucheng Warren, et al.
Published: (2024)
by: Zhu, Fucheng Warren, et al.
Published: (2024)
Chat-Driven Text Generation and Interaction for Person Retrieval
by: Xie, Zequn, et al.
Published: (2025)
by: Xie, Zequn, et al.
Published: (2025)
MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics
by: Li, Jing, et al.
Published: (2025)
by: Li, Jing, et al.
Published: (2025)
Facial Attribute Based Text Guided Face Anonymization
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
FOCUS on Contamination: Hydrology-Informed Noise-Aware Learning for Geospatial PFAS Mapping
by: Khan, Jowaria, et al.
Published: (2025)
by: Khan, Jowaria, et al.
Published: (2025)
A Guide to Structureless Visual Localization
by: Panek, Vojtech, et al.
Published: (2025)
by: Panek, Vojtech, et al.
Published: (2025)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
by: Tourani, Ali, et al.
Published: (2023)
by: Tourani, Ali, et al.
Published: (2023)
Estimating optical vegetation indices and biophysical variables for temperate forests with Sentinel-1 SAR data using machine learning techniques: A case study for Czechia
by: Paluba, Daniel, et al.
Published: (2023)
by: Paluba, Daniel, et al.
Published: (2023)
Combining Absolute and Semi-Generalized Relative Poses for Visual Localization
by: Panek, Vojtech, et al.
Published: (2024)
by: Panek, Vojtech, et al.
Published: (2024)
Privacy-Preserving Structureless Visual Localization via Image Obfuscation
by: Panek, Vojtech, et al.
Published: (2026)
by: Panek, Vojtech, et al.
Published: (2026)
Investigation of cardinality classification for bacterial colony counting using explainable artificial intelligence
by: Zheng, Minghua, et al.
Published: (2026)
by: Zheng, Minghua, et al.
Published: (2026)
Learning to count small and clustered objects with application to bacterial colonies
by: Zheng, Minghua, et al.
Published: (2026)
by: Zheng, Minghua, et al.
Published: (2026)
WPDA: Frequency-based Backdoor Attack with Wavelet Packet Decomposition
by: Song, Zhengyao, et al.
Published: (2024)
by: Song, Zhengyao, et al.
Published: (2024)
Data Augmentation in Earth Observation: A Diffusion Model Approach
by: Sousa, Tiago, et al.
Published: (2024)
by: Sousa, Tiago, et al.
Published: (2024)
Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations
by: Wang, Junjue, et al.
Published: (2026)
by: Wang, Junjue, et al.
Published: (2026)
Traffic Scene Small Target Detection Method Based on YOLOv8n-SPTS Model for Autonomous Driving
by: Wu, Songhan
Published: (2025)
by: Wu, Songhan
Published: (2025)
Visual Style Prompt Learning Using Diffusion Models for Blind Face Restoration
by: Lu, Wanglong, et al.
Published: (2024)
by: Lu, Wanglong, et al.
Published: (2024)
Low-Cost Tree Crown Dieback Estimation Using Deep Learning-Based Segmentation
by: Allen, M. J., et al.
Published: (2024)
by: Allen, M. J., et al.
Published: (2024)
WaveMix: A Resource-efficient Neural Network for Image Analysis
by: Jeevan, Pranav, et al.
Published: (2022)
by: Jeevan, Pranav, et al.
Published: (2022)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
HATL: Hierarchical Adaptive-Transfer Learning Framework for Sign Language Machine Translation
by: Shahin, Nada, et al.
Published: (2026)
by: Shahin, Nada, et al.
Published: (2026)
Model Agnostic Defense against Adversarial Patch Attacks on Object Detection in Unmanned Aerial Vehicles
by: Pathak, Saurabh, et al.
Published: (2024)
by: Pathak, Saurabh, et al.
Published: (2024)
Advanced Long-term Earth System Forecasting
by: Wu, Hao, et al.
Published: (2025)
by: Wu, Hao, et al.
Published: (2025)
DIsoN: Decentralized Isolation Networks for Out-of-Distribution Detection in Medical Imaging
by: Wagner, Felix, et al.
Published: (2025)
by: Wagner, Felix, et al.
Published: (2025)
FLD+: Data-efficient Evaluation Metric for Generative Models
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
Normalizing Flow-Based Metric for Image Generation
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
Similar Items
-
Neural Fields for 3D Tracking of Anatomy and Surgical Instruments in Monocular Laparoscopic Video Clips
by: Gerats, Beerend G. A., et al.
Published: (2024) -
Learning to Expand Images for Efficient Visual Autoregressive Modeling
by: Yang, Ruiqing, et al.
Published: (2025) -
Improving Visual Object Tracking through Visual Prompting
by: Chen, Shih-Fang, et al.
Published: (2024) -
Denoising VAE as an Explainable Feature Reduction and Diagnostic Pipeline for Autism Based on Resting state fMRI
by: Zheng, Xinyuan, et al.
Published: (2024) -
EDSNet: Efficient-DSNet for Video Summarization
by: Prasad, Ashish, et al.
Published: (2024)