Semantic Context Matters: Improving Conditioning for Autoregressive Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Dongyang, Xu, Ryan, Zeng, Jianhao, Lan, Rui, Bai, Yancheng, Sun, Lei, Chu, Xiangxiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCALAR: Scale-wise Controllable Visual Autoregressive Learning
by: Xu, Ryan, et al.
Published: (2025)
by: Xu, Ryan, et al.
Published: (2025)
FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
by: Lan, Rui, et al.
Published: (2025)
by: Lan, Rui, et al.
Published: (2025)
Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
by: Zeng, Jianhao, et al.
Published: (2025)
by: Zeng, Jianhao, et al.
Published: (2025)
TEXTS-Diff: TEXTS-Aware Diffusion Model for Real-World Text Image Super-Resolution
by: He, Haodong, et al.
Published: (2026)
by: He, Haodong, et al.
Published: (2026)
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
by: Li, Mingxing, et al.
Published: (2025)
by: Li, Mingxing, et al.
Published: (2025)
RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
by: He, Haodong, et al.
Published: (2025)
by: He, Haodong, et al.
Published: (2025)
Elucidating the SNR-t Bias of Diffusion Probabilistic Models
by: Yu, Meng, et al.
Published: (2026)
by: Yu, Meng, et al.
Published: (2026)
Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers
by: Chen, Ruidong, et al.
Published: (2026)
by: Chen, Ruidong, et al.
Published: (2026)
Context-Aware Autoregressive Models for Multi-Conditional Image Generation
by: Chen, Yixiao, et al.
Published: (2025)
by: Chen, Yixiao, et al.
Published: (2025)
Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
by: Gu, Jiatao, et al.
Published: (2024)
by: Gu, Jiatao, et al.
Published: (2024)
Hierarchical Context Alignment with Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
FingER: Content Aware Fine-grained Evaluation with Reasoning for AI-Generated Videos
by: Chen, Rui, et al.
Published: (2025)
by: Chen, Rui, et al.
Published: (2025)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
AHMF: Adaptive Hybrid-Memory-Fusion Model for Driver Attention Prediction
by: Xu, Dongyang, et al.
Published: (2024)
by: Xu, Dongyang, et al.
Published: (2024)
SCTNet: Single-Branch CNN with Transformer Semantic Information for Real-Time Segmentation
by: Xu, Zhengze, et al.
Published: (2023)
by: Xu, Zhengze, et al.
Published: (2023)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
by: Chen, Shuo, et al.
Published: (2026)
by: Chen, Shuo, et al.
Published: (2026)
Efficient Conditional Generation on Scale-based Visual Autoregressive Models
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis
by: Chen, Jintao, et al.
Published: (2026)
by: Chen, Jintao, et al.
Published: (2026)
ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer
by: Hu, Jinyi, et al.
Published: (2024)
by: Hu, Jinyi, et al.
Published: (2024)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
by: Bai, Sule, et al.
Published: (2025)
by: Bai, Sule, et al.
Published: (2025)
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
by: Fan, Jiaqi, et al.
Published: (2024)
by: Fan, Jiaqi, et al.
Published: (2024)
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
by: Liu, Zhaochen, et al.
Published: (2024)
by: Liu, Zhaochen, et al.
Published: (2024)
Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
by: Shen, Cuifeng, et al.
Published: (2025)
by: Shen, Cuifeng, et al.
Published: (2025)
EditAR: Unified Conditional Generation with Autoregressive Models
by: Mu, Jiteng, et al.
Published: (2025)
by: Mu, Jiteng, et al.
Published: (2025)
USP: Unified Self-Supervised Pretraining for Image Generation and Understanding
by: Chu, Xiangxiang, et al.
Published: (2025)
by: Chu, Xiangxiang, et al.
Published: (2025)
IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction
by: Yi, Ran, et al.
Published: (2025)
by: Yi, Ran, et al.
Published: (2025)
Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models
by: Phunyaphibarn, Prin, et al.
Published: (2025)
by: Phunyaphibarn, Prin, et al.
Published: (2025)
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
by: Su, Qile, et al.
Published: (2026)
by: Su, Qile, et al.
Published: (2026)
Generative Feature Imputing -- A Technique for Error-resilient Semantic Communication
by: Huang, Jianhao, et al.
Published: (2025)
by: Huang, Jianhao, et al.
Published: (2025)
Urban Socio-Semantic Segmentation with Vision-Language Reasoning
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Structure Matters: Tackling the Semantic Discrepancy in Diffusion Models for Image Inpainting
by: Liu, Haipeng, et al.
Published: (2024)
by: Liu, Haipeng, et al.
Published: (2024)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
by: Gu, Yuchao, et al.
Published: (2025)
by: Gu, Yuchao, et al.
Published: (2025)
MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian Conditioning
by: Zhang, Jinhua, et al.
Published: (2025)
by: Zhang, Jinhua, et al.
Published: (2025)
Context Matters: Learning Global Semantics via Object-Centric Representation
by: Zhong, Jike, et al.
Published: (2025)
by: Zhong, Jike, et al.
Published: (2025)
Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
by: Zheng, Peng, et al.
Published: (2025)
by: Zheng, Peng, et al.
Published: (2025)
Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Hierarchical Temporal Context Learning for Camera-based Semantic Scene Completion
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
by: Lei, Jiachen, et al.
Published: (2025)
by: Lei, Jiachen, et al.
Published: (2025)
HairGPT: Strand-as-Language Autoregressive Modeling for Realistic 3D Hairstyle Synthesis
by: Luo, Haimin, et al.
Published: (2026)
by: Luo, Haimin, et al.
Published: (2026)
Self-control: A Better Conditional Mechanism for Masked Autoregressive Model
by: Qu, Qiaoying, et al.
Published: (2024)
by: Qu, Qiaoying, et al.
Published: (2024)
Similar Items
-
SCALAR: Scale-wise Controllable Visual Autoregressive Learning
by: Xu, Ryan, et al.
Published: (2025) -
FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
by: Lan, Rui, et al.
Published: (2025) -
Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
by: Zeng, Jianhao, et al.
Published: (2025) -
TEXTS-Diff: TEXTS-Aware Diffusion Model for Real-World Text Image Super-Resolution
by: He, Haodong, et al.
Published: (2026) -
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
by: Li, Mingxing, et al.
Published: (2025)