Not All Tokens are Guided Equal: Improving Guidance in Visual Autoregressive Models
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Ky Dan, Tran, Hoang Lam, Dinh, Anh-Dung, Liu, Daochang, Cai, Weidong, Wang, Xiuying, Xu, Chang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compress Guidance in Conditional Diffusion Sampling
by: Dinh, Anh-Dung, et al.
Published: (2024)
by: Dinh, Anh-Dung, et al.
Published: (2024)
Improving Generalization in Visual Reasoning via Self-Ensemble
by: Nguyen, Tien-Huy, et al.
Published: (2024)
by: Nguyen, Tien-Huy, et al.
Published: (2024)
KOPPA: Improving Prompt-based Continual Learning with Key-Query Orthogonal Projection and Prototype-based One-Versus-All
by: Tran, Quyen, et al.
Published: (2023)
by: Tran, Quyen, et al.
Published: (2023)
A Survey on Vietnamese Document Analysis and Recognition: Challenges and Future Directions
by: Le, Anh, et al.
Published: (2025)
by: Le, Anh, et al.
Published: (2025)
Generative Physical AI in Vision: A Survey
by: Liu, Daochang, et al.
Published: (2025)
by: Liu, Daochang, et al.
Published: (2025)
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
by: Nguyen, Quang-Binh, et al.
Published: (2025)
by: Nguyen, Quang-Binh, et al.
Published: (2025)
SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation
by: Nguyen, Thuan Hoang, et al.
Published: (2023)
by: Nguyen, Thuan Hoang, et al.
Published: (2023)
Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control
by: Le, Nhat, et al.
Published: (2026)
by: Le, Nhat, et al.
Published: (2026)
TP-GMOT: Tracking Generic Multiple Object by Textual Prompt with Motion-Appearance Cost (MAC) SORT
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Improved Training Technique for Shortcut Models
by: Nguyen, Anh, et al.
Published: (2025)
by: Nguyen, Anh, et al.
Published: (2025)
Towards Memorization-Free Diffusion Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
by: Nguyen, Hong, et al.
Published: (2025)
by: Nguyen, Hong, et al.
Published: (2025)
CAKE: Real-time Action Detection via Motion Distillation and Background-aware Contrastive Learning
by: Hoang, Hieu, et al.
Published: (2026)
by: Hoang, Hieu, et al.
Published: (2026)
SATURN: Autoregressive Image Generation Guided by Scene Graphs
by: Vo, Thanh-Nhan, et al.
Published: (2025)
by: Vo, Thanh-Nhan, et al.
Published: (2025)
Generalization Bounds for Robust Contrastive Learning: From Theory to Practice
by: Tran, Ngoc N., et al.
Published: (2023)
by: Tran, Ngoc N., et al.
Published: (2023)
Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model Guidance
by: Pham, Duc-Hai, et al.
Published: (2024)
by: Pham, Duc-Hai, et al.
Published: (2024)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)
by: Xu, Siyu, et al.
Published: (2024)
More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning
by: Tran, Luong, et al.
Published: (2025)
by: Tran, Luong, et al.
Published: (2025)
SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
by: Xu, Dongli, et al.
Published: (2025)
by: Xu, Dongli, et al.
Published: (2025)
Not All Pixels Are Equal: Confidence-Guided Attention for Feature Matching
by: Li, Dongyue
Published: (2025)
by: Li, Dongyue
Published: (2025)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction
by: Yi, Ran, et al.
Published: (2025)
by: Yi, Ran, et al.
Published: (2025)
Semise: Semi-supervised learning for severity representation in medical image
by: Tran, Dung T., et al.
Published: (2025)
by: Tran, Dung T., et al.
Published: (2025)
Rest2Visual: Predicting Visually Evoked fMRI from Resting-State Scans
by: Zhou, Chuyang, et al.
Published: (2025)
by: Zhou, Chuyang, et al.
Published: (2025)
Low-Light Enhancement via Encoder-Decoder Network with Illumination Guidance
by: Tran, Le-Anh, et al.
Published: (2025)
by: Tran, Le-Anh, et al.
Published: (2025)
Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Exploring Local Memorization in Diffusion Models via Bright Ending Attention
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
by: Le, Minh, et al.
Published: (2025)
by: Le, Minh, et al.
Published: (2025)
VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation
by: Liao, Xinyao, et al.
Published: (2026)
by: Liao, Xinyao, et al.
Published: (2026)
TinySense: Effective CSI Compression for Scalable and Accurate Wi-Fi Sensing
by: Gian, Toan, et al.
Published: (2026)
by: Gian, Toan, et al.
Published: (2026)
PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
by: Vo, Dinh-Khoi, et al.
Published: (2026)
by: Vo, Dinh-Khoi, et al.
Published: (2026)
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
by: Wang, Bohan, et al.
Published: (2025)
by: Wang, Bohan, et al.
Published: (2025)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Improving Flexible Image Tokenizers for Autoregressive Image Generation
by: Fu, Zixuan, et al.
Published: (2026)
by: Fu, Zixuan, et al.
Published: (2026)
Not All Regions Are Equal: Attention-Guided Perturbation Network for Industrial Anomaly Detection
by: Huang, Tingfeng, et al.
Published: (2024)
by: Huang, Tingfeng, et al.
Published: (2024)
Visual Autoregressive Modeling for Instruction-Guided Image Editing
by: Mao, Qingyang, et al.
Published: (2025)
by: Mao, Qingyang, et al.
Published: (2025)
MetaAug: Meta-Data Augmentation for Post-Training Quantization
by: Pham, Cuong, et al.
Published: (2024)
by: Pham, Cuong, et al.
Published: (2024)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
by: Nguyen, Trong-Tung, et al.
Published: (2024)
by: Nguyen, Trong-Tung, et al.
Published: (2024)
Similar Items
-
Compress Guidance in Conditional Diffusion Sampling
by: Dinh, Anh-Dung, et al.
Published: (2024) -
Improving Generalization in Visual Reasoning via Self-Ensemble
by: Nguyen, Tien-Huy, et al.
Published: (2024) -
KOPPA: Improving Prompt-based Continual Learning with Key-Query Orthogonal Projection and Prototype-based One-Versus-All
by: Tran, Quyen, et al.
Published: (2023) -
A Survey on Vietnamese Document Analysis and Recognition: Challenges and Future Directions
by: Le, Anh, et al.
Published: (2025) -
Generative Physical AI in Vision: A Survey
by: Liu, Daochang, et al.
Published: (2025)