Controllable Image Generation with Composed Parallel Token Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Stirling, Jamie, Al-Moubayed, Noura, Willcocks, Chris G., Shum, Hubert P. H. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Controllable Image Generation with Composed Parallel Token Prediction
by: Stirling, Jamie, et al.
Published: (2026)
by: Stirling, Jamie, et al.
Published: (2026)
Investigating Permutation-Invariant Discrete Representation Learning for Spatially Aligned Images
by: Stirling, Jamie S. J., et al.
Published: (2026)
by: Stirling, Jamie S. J., et al.
Published: (2026)
Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
by: Slack, Dean L, et al.
Published: (2025)
by: Slack, Dean L, et al.
Published: (2025)
Disentangling Racial Phenotypes: Fine-Grained Control of Race-related Facial Phenotype Characteristics
by: Yucer, Seyma, et al.
Published: (2024)
by: Yucer, Seyma, et al.
Published: (2024)
Everything is a Video: Unifying Modalities through Next-Frame Prediction
by: Hudson, G. Thomas, et al.
Published: (2024)
by: Hudson, G. Thomas, et al.
Published: (2024)
Repeat and Concatenate: 2D to 3D Image Translation with 3D to 3D Generative Modeling
by: Corona-Figueroa, Abril, et al.
Published: (2024)
by: Corona-Figueroa, Abril, et al.
Published: (2024)
Textual Localization: Decomposing Multi-concept Images for Subject-Driven Text-to-Image Generation
by: Shentu, Junjie, et al.
Published: (2024)
by: Shentu, Junjie, et al.
Published: (2024)
$\infty$-Diff: Infinite Resolution Diffusion with Subsampled Mollified States
by: Bond-Taylor, Sam, et al.
Published: (2023)
by: Bond-Taylor, Sam, et al.
Published: (2023)
AttenCraft: Attention-guided Disentanglement of Multiple Concepts for Text-to-Image Customization
by: Shentu, Junjie, et al.
Published: (2024)
by: Shentu, Junjie, et al.
Published: (2024)
Data Augmentation via Mixed Class Interpolation using Cycle-Consistent Generative Adversarial Networks Applied to Cross-Domain Imagery
by: Sasaki, Hiroshi, et al.
Published: (2020)
by: Sasaki, Hiroshi, et al.
Published: (2020)
RAPiD-Seg: Range-Aware Pointwise Distance Distribution Networks for 3D LiDAR Segmentation
by: Li, Li, et al.
Published: (2024)
by: Li, Li, et al.
Published: (2024)
MIEB: Massive Image Embedding Benchmark
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
One-Index Vector Quantization Based Adversarial Attack on Image Classification
by: Fan, Haiju, et al.
Published: (2024)
by: Fan, Haiju, et al.
Published: (2024)
TraIL-Det: Transformation-Invariant Local Feature Networks for 3D LiDAR Object Detection with Unsupervised Pre-Training
by: Li, Li, et al.
Published: (2024)
by: Li, Li, et al.
Published: (2024)
The Power of Next-Frame Prediction for Learning Physical Laws
by: Winterbottom, Thomas, et al.
Published: (2024)
by: Winterbottom, Thomas, et al.
Published: (2024)
On the Design Fundamentals of Diffusion Models: A Survey
by: Chang, Ziyi, et al.
Published: (2023)
by: Chang, Ziyi, et al.
Published: (2023)
VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations
by: Patel, Maitreya, et al.
Published: (2026)
by: Patel, Maitreya, et al.
Published: (2026)
Semi-Supervised Crowd Counting from Unlabeled Data
by: Duan, Haoran, et al.
Published: (2021)
by: Duan, Haoran, et al.
Published: (2021)
One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression
by: Miwa, Keita, et al.
Published: (2025)
by: Miwa, Keita, et al.
Published: (2025)
FIRE-CIR: Fine-grained Reasoning for Composed Fashion Image Retrieval
by: Gardères, François, et al.
Published: (2026)
by: Gardères, François, et al.
Published: (2026)
Interaction Asymmetry: A General Principle for Learning Composable Abstractions
by: Brady, Jack, et al.
Published: (2024)
by: Brady, Jack, et al.
Published: (2024)
Spectral Image Tokenizer
by: Esteves, Carlos, et al.
Published: (2024)
by: Esteves, Carlos, et al.
Published: (2024)
ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval
by: Xing, Eric, et al.
Published: (2025)
by: Xing, Eric, et al.
Published: (2025)
CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization
by: Zhao, Zelin, et al.
Published: (2025)
by: Zhao, Zelin, et al.
Published: (2025)
MaskBit: Embedding-free Image Generation via Bit Tokens
by: Weber, Mark, et al.
Published: (2024)
by: Weber, Mark, et al.
Published: (2024)
End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
by: Chu, Wenda, et al.
Published: (2026)
by: Chu, Wenda, et al.
Published: (2026)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
by: Ghasemi, Narges, et al.
Published: (2025)
by: Ghasemi, Narges, et al.
Published: (2025)
HyperTokens: Controlling Token Dynamics for Continual Video-Language Understanding
by: Nguyen, Toan, et al.
Published: (2026)
by: Nguyen, Toan, et al.
Published: (2026)
Language-Guided Image Tokenization for Generation
by: Zha, Kaiwen, et al.
Published: (2024)
by: Zha, Kaiwen, et al.
Published: (2024)
Triplet Synthesis For Enhancing Composed Image Retrieval via Counterfactual Image Generation
by: Uesugi, Kenta, et al.
Published: (2025)
by: Uesugi, Kenta, et al.
Published: (2025)
Composing Parts for Expressive Object Generation
by: Rangwani, Harsh, et al.
Published: (2024)
by: Rangwani, Harsh, et al.
Published: (2024)
DreamComposer: Controllable 3D Object Generation via Multi-View Conditions
by: Yang, Yunhan, et al.
Published: (2023)
by: Yang, Yunhan, et al.
Published: (2023)
DivControl: Knowledge Diversion for Controllable Image Generation
by: Xie, Yucheng, et al.
Published: (2025)
by: Xie, Yucheng, et al.
Published: (2025)
Semantic Prompting with Image-Token for Continual Learning
by: Han, Jisu, et al.
Published: (2024)
by: Han, Jisu, et al.
Published: (2024)
High-Resolution Image Synthesis via Next-Token Prediction
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
BATR-FST: Bi-Level Adaptive Token Refinement for Few-Shot Transformers
by: Al-Habib, Mohammed, et al.
Published: (2025)
by: Al-Habib, Mohammed, et al.
Published: (2025)
Controlling False Positives in Image Segmentation via Conformal Prediction
by: Mossina, Luca, et al.
Published: (2025)
by: Mossina, Luca, et al.
Published: (2025)
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
LGQ: Learning Discretization Geometry for Scalable and Stable Image Tokenization
by: Altun, Idil Bilge, et al.
Published: (2026)
by: Altun, Idil Bilge, et al.
Published: (2026)
Similar Items
-
Controllable Image Generation with Composed Parallel Token Prediction
by: Stirling, Jamie, et al.
Published: (2026) -
Investigating Permutation-Invariant Discrete Representation Learning for Spatially Aligned Images
by: Stirling, Jamie S. J., et al.
Published: (2026) -
Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
by: Slack, Dean L, et al.
Published: (2025) -
Disentangling Racial Phenotypes: Fine-Grained Control of Race-related Facial Phenotype Characteristics
by: Yucer, Seyma, et al.
Published: (2024) -
Everything is a Video: Unifying Modalities through Next-Frame Prediction
by: Hudson, G. Thomas, et al.
Published: (2024)