Halton Scheduler For Masked Generative Image Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Besnier, Victor, Chen, Mickael, Hurych, David, Valle, Eduardo, Cord, Matthieu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BIGFix: Bidirectional Image Generation with Token Fixing
by: Besnier, Victor, et al.
Published: (2025)
by: Besnier, Victor, et al.
Published: (2025)
Supervised Anomaly Detection for Complex Industrial Images
by: Baitieva, Aimira, et al.
Published: (2024)
by: Baitieva, Aimira, et al.
Published: (2024)
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
by: Bartoccioni, Florent, et al.
Published: (2025)
by: Bartoccioni, Florent, et al.
Published: (2025)
Annealed Winner-Takes-All for Motion Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
GaussRender: Learning 3D Occupancy with Gaussian Rendering
by: Chambon, Loïck, et al.
Published: (2025)
by: Chambon, Loïck, et al.
Published: (2025)
PointBeV: A Sparse Approach to BeV Predictions
by: Chambon, Loick, et al.
Published: (2023)
by: Chambon, Loick, et al.
Published: (2023)
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
by: Vallaeys, Théophane, et al.
Published: (2025)
by: Vallaeys, Théophane, et al.
Published: (2025)
Reliability in Semantic Segmentation: Can We Use Synthetic Data?
by: Loiseau, Thibaut, et al.
Published: (2023)
by: Loiseau, Thibaut, et al.
Published: (2023)
GIFT: A Framework Towards Global Interpretable Faithful Textual Explanations of Vision Classifiers
by: Zablocki, Éloi, et al.
Published: (2024)
by: Zablocki, Éloi, et al.
Published: (2024)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?
by: Xu, Yihong, et al.
Published: (2023)
by: Xu, Yihong, et al.
Published: (2023)
MaskMamba: A Hybrid Mamba-Transformer Model for Masked Image Generation
by: Chen, Wenchao, et al.
Published: (2024)
by: Chen, Wenchao, et al.
Published: (2024)
R3DPA: Leveraging 3D Representation Alignment and RGB Pretrained Priors for LiDAR Scene Generation
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting
by: Messaoud, Kaouther, et al.
Published: (2025)
by: Messaoud, Kaouther, et al.
Published: (2025)
Don't drop your samples! Coherence-aware training benefits Conditional diffusion
by: Dufour, Nicolas, et al.
Published: (2024)
by: Dufour, Nicolas, et al.
Published: (2024)
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Let-It-Flow: Simultaneous Optimization of 3D Flow and Object Clustering
by: Vacek, Patrik, et al.
Published: (2024)
by: Vacek, Patrik, et al.
Published: (2024)
ManiPose: Manifold-Constrained Multi-Hypothesis 3D Human Pose Estimation
by: Rommel, Cédric, et al.
Published: (2023)
by: Rommel, Cédric, et al.
Published: (2023)
MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale
by: Tang, Zhicong, et al.
Published: (2026)
by: Tang, Zhicong, et al.
Published: (2026)
Test-Time Conditioning with Representation-Aligned Visual Features
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
by: Vallaeys, Théophane, et al.
Published: (2024)
by: Vallaeys, Théophane, et al.
Published: (2024)
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
by: Vobecky, Antonin, et al.
Published: (2024)
by: Vobecky, Antonin, et al.
Published: (2024)
Valeo4Cast: A Modular Approach to End-to-End Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
Regularizing Self-supervised 3D Scene Flows with Surface Awareness and Cyclic Consistency
by: Vacek, Patrik, et al.
Published: (2023)
by: Vacek, Patrik, et al.
Published: (2023)
Autoregressive Image Generation with Masked Bit Modeling
by: Yu, Qihang, et al.
Published: (2026)
by: Yu, Qihang, et al.
Published: (2026)
FocusDiT: Masking Queries in Diffusion Transformers for Fine-grained Image Generation
by: Fang, Xueji, et al.
Published: (2026)
by: Fang, Xueji, et al.
Published: (2026)
Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
by: Luo, Yifu, et al.
Published: (2025)
by: Luo, Yifu, et al.
Published: (2025)
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
by: Shukor, Mustafa, et al.
Published: (2023)
by: Shukor, Mustafa, et al.
Published: (2023)
ScanMove: Motion Prediction and Transfer for Unregistered Body Meshes
by: Besnier, Thomas, et al.
Published: (2025)
by: Besnier, Thomas, et al.
Published: (2025)
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
by: Bai, Jinbin, et al.
Published: (2024)
by: Bai, Jinbin, et al.
Published: (2024)
LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression Comprehension
by: Cardiel, Amaia, et al.
Published: (2024)
by: Cardiel, Amaia, et al.
Published: (2024)
MAD: Motion Appearance Decoupling for efficient Driving World Models
by: Rahimi, Ahmad, et al.
Published: (2026)
by: Rahimi, Ahmad, et al.
Published: (2026)
Few-shot Image Generation via Masked Discrimination
by: Zhu, Jingyuan, et al.
Published: (2022)
by: Zhu, Jingyuan, et al.
Published: (2022)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
by: Gidaris, Spyros, et al.
Published: (2023)
by: Gidaris, Spyros, et al.
Published: (2023)
Masked Generative Transformer Is What You Need for Image Editing
by: Chow, Wei, et al.
Published: (2026)
by: Chow, Wei, et al.
Published: (2026)
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
by: Couairon, Paul, et al.
Published: (2024)
by: Couairon, Paul, et al.
Published: (2024)
Cross-view Masked Diffusion Transformers for Person Image Synthesis
by: Pham, Trung X., et al.
Published: (2024)
by: Pham, Trung X., et al.
Published: (2024)
Similar Items
-
BIGFix: Bidirectional Image Generation with Token Fixing
by: Besnier, Victor, et al.
Published: (2025) -
Supervised Anomaly Detection for Complex Industrial Images
by: Baitieva, Aimira, et al.
Published: (2024) -
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
by: Bartoccioni, Florent, et al.
Published: (2025) -
Annealed Winner-Takes-All for Motion Forecasting
by: Xu, Yihong, et al.
Published: (2024) -
GaussRender: Learning 3D Occupancy with Gaussian Rendering
by: Chambon, Loïck, et al.
Published: (2025)