ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Ni, Zanlin, Wang, Yulin, Zhou, Renping, Han, Yizeng, Guo, Jiayi, Liu, Zhiyuan, Yao, Yuan, Huang, Gao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaNAT: Exploring Adaptive Policy for Token-Based Image Generation
by: Ni, Zanlin, et al.
Published: (2024)
by: Ni, Zanlin, et al.
Published: (2024)
Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
by: Ni, Zanlin, et al.
Published: (2024)
by: Ni, Zanlin, et al.
Published: (2024)
AdaGen: Learning Adaptive Policy for Image Synthesis
by: Ni, Zanlin, et al.
Published: (2026)
by: Ni, Zanlin, et al.
Published: (2026)
UniTTA: Unified Benchmark and Versatile Framework Towards Realistic Test-Time Adaptation
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
by: Zhou, Renping, et al.
Published: (2025)
by: Zhou, Renping, et al.
Published: (2025)
Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment
by: Guo, Jiayi, et al.
Published: (2024)
by: Guo, Jiayi, et al.
Published: (2024)
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
by: Yue, Yang, et al.
Published: (2026)
by: Yue, Yang, et al.
Published: (2026)
LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
by: Xu, Ruyi, et al.
Published: (2024)
by: Xu, Ruyi, et al.
Published: (2024)
CODA: Repurposing Continuous VAEs for Discrete Tokenization
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
SimPro: A Simple Probabilistic Framework Towards Realistic Long-Tailed Semi-Supervised Learning
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
GRA: Detecting Oriented Objects through Group-wise Rotating and Attention
by: Wang, Jiangshan, et al.
Published: (2024)
by: Wang, Jiangshan, et al.
Published: (2024)
Mask Grounding for Referring Image Segmentation
by: Chng, Yong Xien, et al.
Published: (2023)
by: Chng, Yong Xien, et al.
Published: (2023)
Latency-aware Unified Dynamic Networks for Efficient Image Recognition
by: Han, Yizeng, et al.
Published: (2023)
by: Han, Yizeng, et al.
Published: (2023)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
Few-Step Distillation for Text-to-Image Generation: A Practical Guide
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
Meta-Semi: A Meta-learning Approach for Semi-supervised Learning
by: Wang, Yulin, et al.
Published: (2020)
by: Wang, Yulin, et al.
Published: (2020)
Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
by: Zheng, Peng, et al.
Published: (2025)
by: Zheng, Peng, et al.
Published: (2025)
GSVA: Generalized Segmentation via Multimodal Large Language Models
by: Xia, Zhuofan, et al.
Published: (2023)
by: Xia, Zhuofan, et al.
Published: (2023)
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
by: Zhao, Wangbo, et al.
Published: (2025)
by: Zhao, Wangbo, et al.
Published: (2025)
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
by: Pu, Yifan, et al.
Published: (2024)
by: Pu, Yifan, et al.
Published: (2024)
Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception
by: Wang, Yulin, et al.
Published: (2025)
by: Wang, Yulin, et al.
Published: (2025)
Spatial Gram Alignment for Ultra-High-Resolution Image Synthesis
by: Zhang, Jinjin, et al.
Published: (2026)
by: Zhang, Jinjin, et al.
Published: (2026)
Cross-Modal Adapter for Vision-Language Retrieval
by: Jiang, Haojun, et al.
Published: (2022)
by: Jiang, Haojun, et al.
Published: (2022)
Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
Improved Masked Image Generation with Knowledge-Augmented Token Representations
by: Liang, Guotao, et al.
Published: (2025)
by: Liang, Guotao, et al.
Published: (2025)
Adapting Vision-Language Model with Fine-grained Semantics for Open-Vocabulary Segmentation
by: Chng, Yong Xien, et al.
Published: (2024)
by: Chng, Yong Xien, et al.
Published: (2024)
Exploring contextual modeling with linear complexity for point cloud segmentation
by: Chng, Yong Xien, et al.
Published: (2024)
by: Chng, Yong Xien, et al.
Published: (2024)
Rethinking Point Clouds as Sequences: A Causal Next-Token Predictive Learning Framework
by: Yao, Yumeng, et al.
Published: (2026)
by: Yao, Yumeng, et al.
Published: (2026)
Relational Retrieval: Leveraging Known-Novel Interactions for Generalized Category Discovery
by: Xu, Yulin, et al.
Published: (2026)
by: Xu, Yulin, et al.
Published: (2026)
Rethinking the Architecture Design for Efficient Generic Event Boundary Detection
by: Zheng, Ziwei, et al.
Published: (2024)
by: Zheng, Ziwei, et al.
Published: (2024)
TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization
by: Pan, Liang, et al.
Published: (2025)
by: Pan, Liang, et al.
Published: (2025)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026)
by: Liu, Zeyu, et al.
Published: (2026)
4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models
by: Li, Wanhua, et al.
Published: (2025)
by: Li, Wanhua, et al.
Published: (2025)
DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
by: Yang, Le, et al.
Published: (2024)
by: Yang, Le, et al.
Published: (2024)
Agent Attention: On the Integration of Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2023)
by: Han, Dongchen, et al.
Published: (2023)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
by: Mao, Junzhu, et al.
Published: (2025)
by: Mao, Junzhu, et al.
Published: (2025)
AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
PICS: Pairwise Image Compositing with Spatial Interactions
by: Zhou, Hang, et al.
Published: (2026)
by: Zhou, Hang, et al.
Published: (2026)
Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization
by: Liang, Jingyun, et al.
Published: (2026)
by: Liang, Jingyun, et al.
Published: (2026)
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration
by: Han, Jiayi, et al.
Published: (2025)
by: Han, Jiayi, et al.
Published: (2025)
Similar Items
-
AdaNAT: Exploring Adaptive Policy for Token-Based Image Generation
by: Ni, Zanlin, et al.
Published: (2024) -
Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
by: Ni, Zanlin, et al.
Published: (2024) -
AdaGen: Learning Adaptive Policy for Image Synthesis
by: Ni, Zanlin, et al.
Published: (2026) -
UniTTA: Unified Benchmark and Versatile Framework Towards Realistic Test-Time Adaptation
by: Du, Chaoqun, et al.
Published: (2024) -
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
by: Zhou, Renping, et al.
Published: (2025)