SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xiaoyan, Bai, Zechen, Wang, Haofan, Song, Yiren |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Loom: Diffusion-Transformer for Interleaved Generation
von: Ye, Mingcheng, et al.
Veröffentlicht: (2025)
von: Ye, Mingcheng, et al.
Veröffentlicht: (2025)
OmniPSD: Layered PSD Generation with Diffusion Transformer
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs
von: Zheng, Yiren, et al.
Veröffentlicht: (2026)
von: Zheng, Yiren, et al.
Veröffentlicht: (2026)
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
von: Lu, Runnan, et al.
Veröffentlicht: (2025)
von: Lu, Runnan, et al.
Veröffentlicht: (2025)
EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2025)
VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
Factorized Visual Tokenization and Generation
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
TokenPure: Watermark Removal through Tokenized Appearance and Structural Guidance
von: Yang, Pei, et al.
Veröffentlicht: (2025)
von: Yang, Pei, et al.
Veröffentlicht: (2025)
TokenTrace: Multi-Concept Attribution through Watermarked Token Recovery
von: Zhang, Li, et al.
Veröffentlicht: (2026)
von: Zhang, Li, et al.
Veröffentlicht: (2026)
GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
von: Wang, Chun, et al.
Veröffentlicht: (2025)
von: Wang, Chun, et al.
Veröffentlicht: (2025)
SIGMA-GEN: Structure and Identity Guided Multi-subject Assembly for Image Generation
von: Saha, Oindrila, et al.
Veröffentlicht: (2025)
von: Saha, Oindrila, et al.
Veröffentlicht: (2025)
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
von: Liao, Chao, et al.
Veröffentlicht: (2025)
von: Liao, Chao, et al.
Veröffentlicht: (2025)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
ChatUMM: Robust Context Tracking for Conversational Interleaved Generation
von: Dai, Wenxun, et al.
Veröffentlicht: (2026)
von: Dai, Wenxun, et al.
Veröffentlicht: (2026)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
TransAnimate: Taming Layer Diffusion to Generate RGBA Video
von: Chen, Xuewei, et al.
Veröffentlicht: (2025)
von: Chen, Xuewei, et al.
Veröffentlicht: (2025)
InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation
von: Wang, Haofan, et al.
Veröffentlicht: (2024)
von: Wang, Haofan, et al.
Veröffentlicht: (2024)
InstantStyle-Plus: Style Transfer with Content-Preserving in Text-to-Image Generation
von: Wang, Haofan, et al.
Veröffentlicht: (2024)
von: Wang, Haofan, et al.
Veröffentlicht: (2024)
Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training
von: Xing, Jinbo, et al.
Veröffentlicht: (2026)
von: Xing, Jinbo, et al.
Veröffentlicht: (2026)
SIGMA: Sinkhorn-Guided Masked Video Modeling
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
Impossible Videos
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
von: Zhang, Yabo, et al.
Veröffentlicht: (2026)
von: Zhang, Yabo, et al.
Veröffentlicht: (2026)
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
SIGMA: Scale-Invariant Global Sparse Shape Matching
von: Gao, Maolin, et al.
Veröffentlicht: (2023)
von: Gao, Maolin, et al.
Veröffentlicht: (2023)
SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2023)
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2023)
iFlame: Interleaving Full and Linear Attention for Efficient Mesh Generation
von: Wang, Hanxiao, et al.
Veröffentlicht: (2025)
von: Wang, Hanxiao, et al.
Veröffentlicht: (2025)
MultiBind: A Benchmark for Attribute Misbinding in Multi-Subject Generation
von: Tian, Wenqing, et al.
Veröffentlicht: (2026)
von: Tian, Wenqing, et al.
Veröffentlicht: (2026)
Image Watermarks are Removable Using Controllable Regeneration from Clean Noise
von: Liu, Yepeng, et al.
Veröffentlicht: (2024)
von: Liu, Yepeng, et al.
Veröffentlicht: (2024)
CSGO: Content-Style Composition in Text-to-Image Generation
von: Xing, Peng, et al.
Veröffentlicht: (2024)
von: Xing, Peng, et al.
Veröffentlicht: (2024)
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026)
DuoGen: Towards General Purpose Interleaved Multimodal Generation
von: Shi, Min, et al.
Veröffentlicht: (2026)
von: Shi, Min, et al.
Veröffentlicht: (2026)
RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-Key Identification
von: Ci, Hai, et al.
Veröffentlicht: (2024)
von: Ci, Hai, et al.
Veröffentlicht: (2024)
Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
von: Nie, Ming, et al.
Veröffentlicht: (2026)
von: Nie, Ming, et al.
Veröffentlicht: (2026)
TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets
von: Liu, Zhixuan, et al.
Veröffentlicht: (2026)
von: Liu, Zhixuan, et al.
Veröffentlicht: (2026)
Mitty: Diffusion-based Human-to-Robot Video Generation
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Loom: Diffusion-Transformer for Interleaved Generation
von: Ye, Mingcheng, et al.
Veröffentlicht: (2025) -
OmniPSD: Layered PSD Generation with Diffusion Transformer
von: Liu, Cheng, et al.
Veröffentlicht: (2025) -
Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs
von: Zheng, Yiren, et al.
Veröffentlicht: (2026) -
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
von: Lu, Runnan, et al.
Veröffentlicht: (2025) -
EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2025)