CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, YuXin, Lu, Yu, Sun, Haoyuan, Yao, Huanjin, Liu, Fanglong, Sun, Yifan, Feng, Haocheng, Zhou, Hang, Wang, Jingdong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dense Connector for MLLMs
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
Automated Multi-level Preference for MLLMs
von: Zhang, Mengxi, et al.
Veröffentlicht: (2024)
von: Zhang, Mengxi, et al.
Veröffentlicht: (2024)
RefAlign: Representation Alignment for Reference-to-Video Generation
von: Wang, Lei, et al.
Veröffentlicht: (2026)
von: Wang, Lei, et al.
Veröffentlicht: (2026)
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
von: Song, Yuxin, et al.
Veröffentlicht: (2025)
von: Song, Yuxin, et al.
Veröffentlicht: (2025)
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
von: Wu, Wenhao, et al.
Veröffentlicht: (2023)
von: Wu, Wenhao, et al.
Veröffentlicht: (2023)
SFTformer: A Spatial-Frequency-Temporal Correlation-Decoupling Transformer for Radar Echo Extrapolation
von: Xu, Liangyu, et al.
Veröffentlicht: (2024)
von: Xu, Liangyu, et al.
Veröffentlicht: (2024)
Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
Preserving Hamiltonian Locality in Real-Space Coarse-Graining via Kernel Projection
von: Haoyuan, Sun
Veröffentlicht: (2026)
von: Haoyuan, Sun
Veröffentlicht: (2026)
MonoFormer: One Transformer for Both Diffusion and Autoregression
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection
von: Zhang, Yanran, et al.
Veröffentlicht: (2026)
von: Zhang, Yanran, et al.
Veröffentlicht: (2026)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
von: Li, YuXin, et al.
Veröffentlicht: (2025)
von: Li, YuXin, et al.
Veröffentlicht: (2025)
AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models
von: Yao, Fanglong, et al.
Veröffentlicht: (2024)
von: Yao, Fanglong, et al.
Veröffentlicht: (2024)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
von: Sun, Yasheng, et al.
Veröffentlicht: (2025)
von: Sun, Yasheng, et al.
Veröffentlicht: (2025)
Beyond Attribution: Unified Concept-Level Explanations
von: Liu, Junhao, et al.
Veröffentlicht: (2024)
von: Liu, Junhao, et al.
Veröffentlicht: (2024)
GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection
von: Huang, Xuan, et al.
Veröffentlicht: (2026)
von: Huang, Xuan, et al.
Veröffentlicht: (2026)
Towards Efficient Multimodal Unified Reasoning Model via Model Merging
von: Yin, Qixiang, et al.
Veröffentlicht: (2025)
von: Yin, Qixiang, et al.
Veröffentlicht: (2025)
3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
von: Song, Xinyang, et al.
Veröffentlicht: (2025)
von: Song, Xinyang, et al.
Veröffentlicht: (2025)
UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
von: Song, Xinyang, et al.
Veröffentlicht: (2025)
von: Song, Xinyang, et al.
Veröffentlicht: (2025)
Gromov-Hausdorff limits of the Chern-Ricci flow on smooth Hermitian minimal models of general type
von: Sun, Haoyuan
Veröffentlicht: (2026)
von: Sun, Haoyuan
Veröffentlicht: (2026)
Mixed Hessian inequalities on Hermitian manifolds and applications
von: Sun, Haoyuan
Veröffentlicht: (2025)
von: Sun, Haoyuan
Veröffentlicht: (2025)
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
von: Fan, Yingying, et al.
Veröffentlicht: (2025)
von: Fan, Yingying, et al.
Veröffentlicht: (2025)
LesionGen: A Concept-Guided Diffusion Model for Dermatology Image Synthesis
von: Fayyad, Jamil, et al.
Veröffentlicht: (2025)
von: Fayyad, Jamil, et al.
Veröffentlicht: (2025)
Microscopic and Macroscopic Analysis of Purple Sweet Potato Dried Products Following Vacuum Steam Pulsation Blanching Pretreatment
von: Dong Wang, et al.
Veröffentlicht: (2024)
von: Dong Wang, et al.
Veröffentlicht: (2024)
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
Generic Local Duality and Purity Exponents
von: Hochster, Melvin, et al.
Veröffentlicht: (2025)
von: Hochster, Melvin, et al.
Veröffentlicht: (2025)
Assessing Model Generalization in Vicinity
von: Liu, Yuchi, et al.
Veröffentlicht: (2024)
von: Liu, Yuchi, et al.
Veröffentlicht: (2024)
NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation
von: Liu, Youzhi, et al.
Veröffentlicht: (2024)
von: Liu, Youzhi, et al.
Veröffentlicht: (2024)
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
A Unified Approach to Controlling Implicit Regularization via Mirror Descent
von: Sun, Haoyuan, et al.
Veröffentlicht: (2023)
von: Sun, Haoyuan, et al.
Veröffentlicht: (2023)
NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization
von: Li, Yifei, et al.
Veröffentlicht: (2025)
von: Li, Yifei, et al.
Veröffentlicht: (2025)
Visible‐Light Driven Radical Cyclization Strategy for C(sp 3 )─CH 2 CF 3 ─Functional Tetrahydroquinoline Derivatives
von: Liwen Lu, et al.
Veröffentlicht: (2025)
von: Liwen Lu, et al.
Veröffentlicht: (2025)
Nitrogen/Oxygen Co‐Doped Carbon Quantum Dots with Efficient High Color‐Purity Red Emission for Bright Electroluminescent LEDs
von: Chenhao Li, et al.
Veröffentlicht: (2025)
von: Chenhao Li, et al.
Veröffentlicht: (2025)
Nitrogen/Oxygen Co‐Doped Carbon Quantum Dots with Efficient High Color‐Purity Red Emission for Bright Electroluminescent LEDs
von: Chenhao Li, et al.
Veröffentlicht: (2025)
von: Chenhao Li, et al.
Veröffentlicht: (2025)
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
OmniGen: Unified Image Generation
von: Xiao, Shitao, et al.
Veröffentlicht: (2024)
von: Xiao, Shitao, et al.
Veröffentlicht: (2024)
P3Net: Progressive and Periodic Perturbation for Semi-Supervised Medical Image Segmentation
von: Yao, Zhenyan, et al.
Veröffentlicht: (2025)
von: Yao, Zhenyan, et al.
Veröffentlicht: (2025)
Robust Latent Representation Tuning for Image-text Classification
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
AI for Mathematics: Progress, Challenges, and Prospects
von: Ju, Haocheng, et al.
Veröffentlicht: (2026)
von: Ju, Haocheng, et al.
Veröffentlicht: (2026)
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dense Connector for MLLMs
von: Yao, Huanjin, et al.
Veröffentlicht: (2024) -
Automated Multi-level Preference for MLLMs
von: Zhang, Mengxi, et al.
Veröffentlicht: (2024) -
RefAlign: Representation Alignment for Reference-to-Video Generation
von: Wang, Lei, et al.
Veröffentlicht: (2026) -
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
von: Song, Yuxin, et al.
Veröffentlicht: (2025) -
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
von: Wu, Wenhao, et al.
Veröffentlicht: (2023)