Masked Generative Nested Transformers with Decode Time Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Goyal, Sahil, Tula, Debapriya, Jain, Gagan, Shenoy, Pradeep, Jain, Prateek, Paul, Sujoy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
by: Jain, Gagan, et al.
Published: (2024)
by: Jain, Gagan, et al.
Published: (2024)
LookupViT: Compressing visual information to a limited number of tokens
by: Koner, Rajat, et al.
Published: (2024)
by: Koner, Rajat, et al.
Published: (2024)
Improving Generalization via Meta-Learning on Hard Samples
by: Jain, Nishant, et al.
Published: (2024)
by: Jain, Nishant, et al.
Published: (2024)
ELT: Elastic Looped Transformers for Visual Generation
by: Goyal, Sahil, et al.
Published: (2026)
by: Goyal, Sahil, et al.
Published: (2026)
Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification
by: Mittal, Prateek, et al.
Published: (2023)
by: Mittal, Prateek, et al.
Published: (2023)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
by: Kišš, Martin, et al.
Published: (2025)
by: Kišš, Martin, et al.
Published: (2025)
PLUM: Improving Inference Efficiency By Leveraging Repetition-Sparsity Trade-Off
by: Kuhar, Sachit, et al.
Published: (2023)
by: Kuhar, Sachit, et al.
Published: (2023)
Fast Training of Diffusion Models with Masked Transformers
by: Zheng, Hongkai, et al.
Published: (2023)
by: Zheng, Hongkai, et al.
Published: (2023)
Accurate and Efficient World Modeling with Masked Latent Transformers
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
MaskOpt: A Large-Scale Mask Optimization Dataset to Advance AI in Integrated Circuit Manufacturing
by: Hu, Yuting, et al.
Published: (2025)
by: Hu, Yuting, et al.
Published: (2025)
Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation
by: Yariv, Guy, et al.
Published: (2025)
by: Yariv, Guy, et al.
Published: (2025)
MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild
by: Saleem, Muhammad Usama, et al.
Published: (2024)
by: Saleem, Muhammad Usama, et al.
Published: (2024)
Improved Algorithm for Deep Active Learning under Imbalance via Optimal Separation
by: Nuggehalli, Shyam, et al.
Published: (2023)
by: Nuggehalli, Shyam, et al.
Published: (2023)
MMM: Generative Masked Motion Model
by: Pinyoanuntapong, Ekkasit, et al.
Published: (2023)
by: Pinyoanuntapong, Ekkasit, et al.
Published: (2023)
MatFormer: Nested Transformer for Elastic Inference
by: Devvrit, et al.
Published: (2023)
by: Devvrit, et al.
Published: (2023)
Masked Face Recognition with Generative-to-Discriminative Representations
by: Ge, Shiming, et al.
Published: (2024)
by: Ge, Shiming, et al.
Published: (2024)
Semi-Supervised Masked Autoencoders: Unlocking Vision Transformer Potential with Limited Data
by: Faysal, Atik, et al.
Published: (2026)
by: Faysal, Atik, et al.
Published: (2026)
Leaf Angle Estimation using Mask R-CNN and LETR Vision Transformer
by: Margapuri, Venkat, et al.
Published: (2024)
by: Margapuri, Venkat, et al.
Published: (2024)
Inference-Time Scaling of Diffusion Models for Infrared Data Generation
by: Horstmann, Kai A., et al.
Published: (2025)
by: Horstmann, Kai A., et al.
Published: (2025)
DocSynthv2: A Practical Autoregressive Modeling for Document Generation
by: Biswas, Sanket, et al.
Published: (2024)
by: Biswas, Sanket, et al.
Published: (2024)
ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery
by: Shrivastava, Ayush, et al.
Published: (2026)
by: Shrivastava, Ayush, et al.
Published: (2026)
MTCNET: Multi-task Learning Paradigm for Crowd Count Estimation
by: Kumar, Abhay, et al.
Published: (2019)
by: Kumar, Abhay, et al.
Published: (2019)
SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
by: Barrett, Samuel J., et al.
Published: (2025)
by: Barrett, Samuel J., et al.
Published: (2025)
Scaling Image and Video Generation via Test-Time Evolutionary Search
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
by: Kumbong, Hermann, et al.
Published: (2025)
by: Kumbong, Hermann, et al.
Published: (2025)
LLM Augmented LLMs: Expanding Capabilities through Composition
by: Bansal, Rachit, et al.
Published: (2024)
by: Bansal, Rachit, et al.
Published: (2024)
IMTS is Worth Time $\times$ Channel Patches: Visual Masked Autoencoders for Irregular Multivariate Time Series Prediction
by: Hu, Zhangyi, et al.
Published: (2025)
by: Hu, Zhangyi, et al.
Published: (2025)
LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines
by: Cherian, Anoop, et al.
Published: (2024)
by: Cherian, Anoop, et al.
Published: (2024)
Beyond Human Vision: The Role of Large Vision Language Models in Microscope Image Analysis
by: Verma, Prateek, et al.
Published: (2024)
by: Verma, Prateek, et al.
Published: (2024)
VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection
by: Chen, PengYu, et al.
Published: (2026)
by: Chen, PengYu, et al.
Published: (2026)
LookSync: Large-Scale Visual Product Search System for AI-Generated Fashion Looks
by: M, Pradeep, et al.
Published: (2025)
by: M, Pradeep, et al.
Published: (2025)
MixMask: Revisiting Masking Strategy for Siamese ConvNets
by: Vishniakov, Kirill, et al.
Published: (2022)
by: Vishniakov, Kirill, et al.
Published: (2022)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters
by: Chen, Mouxiang, et al.
Published: (2024)
by: Chen, Mouxiang, et al.
Published: (2024)
Mechanisms of Non-Monotonic Scaling in Vision Transformers
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
MedSegFactory: Text-Guided Generation of Medical Image-Mask Pairs
by: Mao, Jiawei, et al.
Published: (2025)
by: Mao, Jiawei, et al.
Published: (2025)
Decoding Defensive Coverage Responsibilities in American Football Using Factorized Attention Based Transformer Models
by: Song, Kevin, et al.
Published: (2026)
by: Song, Kevin, et al.
Published: (2026)
REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
by: Zhang, Yitian, et al.
Published: (2025)
by: Zhang, Yitian, et al.
Published: (2025)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
by: Li, Xingyao, et al.
Published: (2026)
by: Li, Xingyao, et al.
Published: (2026)
Similar Items
-
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
by: Jain, Gagan, et al.
Published: (2024) -
LookupViT: Compressing visual information to a limited number of tokens
by: Koner, Rajat, et al.
Published: (2024) -
Improving Generalization via Meta-Learning on Hard Samples
by: Jain, Nishant, et al.
Published: (2024) -
ELT: Elastic Looped Transformers for Visual Generation
by: Goyal, Sahil, et al.
Published: (2026) -
Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification
by: Mittal, Prateek, et al.
Published: (2023)