SIGMA-GEN: Structure and Identity Guided Multi-subject Assembly for Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Saha, Oindrila, Krs, Vojtech, Mech, Radomir, Maji, Subhransu, Blackburn-Matzen, Kevin, Gadelha, Matheus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
3D Space as a Scratchpad for Editable Text-to-Image Generation
by: Saha, Oindrila, et al.
Published: (2026)
by: Saha, Oindrila, et al.
Published: (2026)
Not All Birds Look The Same: Identity-Preserving Generation For Birds
by: Sun, Aaron, et al.
Published: (2025)
by: Sun, Aaron, et al.
Published: (2025)
Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions
by: Saha, Oindrila, et al.
Published: (2024)
by: Saha, Oindrila, et al.
Published: (2024)
Generate, Transduct, Adapt: Iterative Transduction with VLMs
by: Saha, Oindrila, et al.
Published: (2025)
by: Saha, Oindrila, et al.
Published: (2025)
Seeing Through Clutter: Structured 3D Scene Reconstruction via Iterative Object Removal
by: Aguina-Kang, Rio, et al.
Published: (2026)
by: Aguina-Kang, Rio, et al.
Published: (2026)
Learning Continuous 3D Words for Text-to-Image Generation
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
by: Lawrence, Logan, et al.
Published: (2025)
by: Lawrence, Logan, et al.
Published: (2025)
Removing Reflections from RAW Photos
by: Kee, Eric, et al.
Published: (2024)
by: Kee, Eric, et al.
Published: (2024)
Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships
by: Daroya, Rangel, et al.
Published: (2024)
by: Daroya, Rangel, et al.
Published: (2024)
Active Measurement of Two-Point Correlations
by: Hamilton, Max, et al.
Published: (2026)
by: Hamilton, Max, et al.
Published: (2026)
Merlin L48 Spectrogram Dataset
by: Sun, Aaron, et al.
Published: (2025)
by: Sun, Aaron, et al.
Published: (2025)
RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs
by: Lawrence, Logan, et al.
Published: (2026)
by: Lawrence, Logan, et al.
Published: (2026)
YouDream: Generating Anatomically Controllable Consistent Text-to-3D Animals
by: Mishra, Sandeep, et al.
Published: (2024)
by: Mishra, Sandeep, et al.
Published: (2024)
C3DAG: Controlled 3D Animal Generation using 3D pose guidance
by: Mishra, Sandeep, et al.
Published: (2024)
by: Mishra, Sandeep, et al.
Published: (2024)
Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
by: Wang, Boyang, et al.
Published: (2025)
by: Wang, Boyang, et al.
Published: (2025)
Human-in-the-Loop Visual Re-ID for Population Size Estimation
by: Perez, Gustavo, et al.
Published: (2023)
by: Perez, Gustavo, et al.
Published: (2023)
SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens
by: Zhang, Xiaoyan, et al.
Published: (2026)
by: Zhang, Xiaoyan, et al.
Published: (2026)
VIDMP3: Video Editing by Representing Motion with Pose and Position Priors
by: Mishra, Sandeep, et al.
Published: (2025)
by: Mishra, Sandeep, et al.
Published: (2025)
SIGMA: Sinkhorn-Guided Masked Video Modeling
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image Sets
by: Decatur, Dale, et al.
Published: (2025)
by: Decatur, Dale, et al.
Published: (2025)
DISN: Deep Implicit Surface Network for High-quality Single-view 3D Reconstruction
by: Xu, Qiangeng, et al.
Published: (2019)
by: Xu, Qiangeng, et al.
Published: (2019)
Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study
by: Liu, Wuao, et al.
Published: (2026)
by: Liu, Wuao, et al.
Published: (2026)
Fast View Synthesis of Casual Videos with Soup-of-Planes
by: Lee, Yao-Chih, et al.
Published: (2023)
by: Lee, Yao-Chih, et al.
Published: (2023)
WildSAT: Learning Satellite Image Representations from Wildlife Observations
by: Daroya, Rangel, et al.
Published: (2024)
by: Daroya, Rangel, et al.
Published: (2024)
Active Measurement: Efficient Estimation at Scale
by: Hamilton, Max, et al.
Published: (2025)
by: Hamilton, Max, et al.
Published: (2025)
DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
by: Ngo, Tuan Duc, et al.
Published: (2026)
by: Ngo, Tuan Duc, et al.
Published: (2026)
Pattern Analogies: Learning to Perform Programmatic Image Edits by Analogy
by: Ganeshan, Aditya, et al.
Published: (2024)
by: Ganeshan, Aditya, et al.
Published: (2024)
CATRF: Codec-Adaptive TriPlane Radiance Fields for Volumetric Content Delivery
by: Chen, Tung-I, et al.
Published: (2026)
by: Chen, Tung-I, et al.
Published: (2026)
Improving Satellite Imagery Masking using Multi-task and Transfer Learning
by: Daroya, Rangel, et al.
Published: (2024)
by: Daroya, Rangel, et al.
Published: (2024)
SIGMA: Bridging Structural and Distributional Gaps for Vision Foundation Model Adaptation
by: Xiong, Lingyu, et al.
Published: (2026)
by: Xiong, Lingyu, et al.
Published: (2026)
EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
by: Vandersanden, Jente, et al.
Published: (2026)
by: Vandersanden, Jente, et al.
Published: (2026)
Reproducibility Study of "ITI-GEN: Inclusive Text-to-Image Generation"
by: Fernández, Daniel Gallo, et al.
Published: (2024)
by: Fernández, Daniel Gallo, et al.
Published: (2024)
LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation
by: Liu, Mushui, et al.
Published: (2024)
by: Liu, Mushui, et al.
Published: (2024)
Consensus-Driven Active Model Selection
by: Kay, Justin, et al.
Published: (2025)
by: Kay, Justin, et al.
Published: (2025)
SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images
by: Truong, Bao, et al.
Published: (2026)
by: Truong, Bao, et al.
Published: (2026)
StdGEN: Semantic-Decomposed 3D Character Generation from Single Images
by: He, Yuze, et al.
Published: (2024)
by: He, Yuze, et al.
Published: (2024)
Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
by: Shen, Liao, et al.
Published: (2025)
by: Shen, Liao, et al.
Published: (2025)
GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
by: Fortier-Chouinard, Frédéric, et al.
Published: (2025)
by: Fortier-Chouinard, Frédéric, et al.
Published: (2025)
DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation
by: Lin, Hongbin, et al.
Published: (2025)
by: Lin, Hongbin, et al.
Published: (2025)
Unmasking Deep Fakes: Leveraging Deep Learning for Video Authenticity Detection
by: Hasan, Mahmudul, et al.
Published: (2025)
by: Hasan, Mahmudul, et al.
Published: (2025)
Similar Items
-
3D Space as a Scratchpad for Editable Text-to-Image Generation
by: Saha, Oindrila, et al.
Published: (2026) -
Not All Birds Look The Same: Identity-Preserving Generation For Birds
by: Sun, Aaron, et al.
Published: (2025) -
Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions
by: Saha, Oindrila, et al.
Published: (2024) -
Generate, Transduct, Adapt: Iterative Transduction with VLMs
by: Saha, Oindrila, et al.
Published: (2025) -
Seeing Through Clutter: Structured 3D Scene Reconstruction via Iterative Object Removal
by: Aguina-Kang, Rio, et al.
Published: (2026)