CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Quang-Binh, Luu, Minh, Nguyen, Quang, Tran, Anh, Nguyen, Khoi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
by: Nguyen, Hung, et al.
Published: (2024)
by: Nguyen, Hung, et al.
Published: (2024)
OE3DIS: Open-Ended 3D Point Cloud Instance Segmentation
by: Nguyen, Phuc D. A., et al.
Published: (2024)
by: Nguyen, Phuc D. A., et al.
Published: (2024)
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments
by: Nguyen, Van Quang
Published: (2026)
by: Nguyen, Van Quang
Published: (2026)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
by: Nguyen, Toan, et al.
Published: (2025)
by: Nguyen, Toan, et al.
Published: (2025)
TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning
by: Dinh, Quang Minh, et al.
Published: (2024)
by: Dinh, Quang Minh, et al.
Published: (2024)
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
by: Nguyen, Cong-Duy, et al.
Published: (2025)
by: Nguyen, Cong-Duy, et al.
Published: (2025)
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
by: Nguyen, Trong-Tung, et al.
Published: (2024)
by: Nguyen, Trong-Tung, et al.
Published: (2024)
Stable Messenger: Steganography for Message-Concealed Image Generation
by: Nguyen, Quang, et al.
Published: (2023)
by: Nguyen, Quang, et al.
Published: (2023)
MambaU-Lite: A Lightweight Model based on Mamba and Integrated Channel-Spatial Attention for Skin Lesion Segmentation
by: Nguyen, Thi-Nhu-Quynh, et al.
Published: (2024)
by: Nguyen, Thi-Nhu-Quynh, et al.
Published: (2024)
Semi-Supervised Semantic Segmentation using Redesigned Self-Training for White Blood Cells
by: Luu, Vinh Quoc, et al.
Published: (2024)
by: Luu, Vinh Quoc, et al.
Published: (2024)
Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask Tracking
by: Nguyen, Phuc, et al.
Published: (2024)
by: Nguyen, Phuc, et al.
Published: (2024)
StyleVAR: Controllable Image Style Transfer via Visual Autoregressive Modeling
by: Jing, Liqi, et al.
Published: (2026)
by: Jing, Liqi, et al.
Published: (2026)
InverFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting
by: Vu, Duc, et al.
Published: (2026)
by: Vu, Duc, et al.
Published: (2026)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
by: Pham, Anh-Cuong, et al.
Published: (2024)
by: Pham, Anh-Cuong, et al.
Published: (2024)
PAT: Pixel-wise Adaptive Training for Long-tailed Segmentation
by: Do, Khoi, et al.
Published: (2024)
by: Do, Khoi, et al.
Published: (2024)
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
by: Quang, Ngoc Bui Lam, et al.
Published: (2025)
by: Quang, Ngoc Bui Lam, et al.
Published: (2025)
Lightweight Temporal Transformer Decomposition for Federated Autonomous Driving
by: Do, Tuong, et al.
Published: (2025)
by: Do, Tuong, et al.
Published: (2025)
DiverseDream: Diverse Text-to-3D Synthesis with Augmented Text Embedding
by: Tran, Uy Dieu, et al.
Published: (2023)
by: Tran, Uy Dieu, et al.
Published: (2023)
ModeDreamer: Mode Guiding Score Distillation for Text-to-3D Generation using Reference Image Prompts
by: Tran, Uy Dieu, et al.
Published: (2024)
by: Tran, Uy Dieu, et al.
Published: (2024)
SwiftBrush v2: Make Your One-step Diffusion Model Better Than Its Teacher
by: Dao, Trung, et al.
Published: (2024)
by: Dao, Trung, et al.
Published: (2024)
Improving Generalization in Visual Reasoning via Self-Ensemble
by: Nguyen, Tien-Huy, et al.
Published: (2024)
by: Nguyen, Tien-Huy, et al.
Published: (2024)
Point Cloud Compression with Bits-back Coding
by: Hieu, Nguyen Quang, et al.
Published: (2024)
by: Hieu, Nguyen Quang, et al.
Published: (2024)
Not All Tokens are Guided Equal: Improving Guidance in Visual Autoregressive Models
by: Nguyen, Ky Dan, et al.
Published: (2025)
by: Nguyen, Ky Dan, et al.
Published: (2025)
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding
by: Nguyen-Truong, Hai, et al.
Published: (2024)
by: Nguyen-Truong, Hai, et al.
Published: (2024)
TF-SASM: Training-free Spatial-aware Sparse Memory for Multi-object Tracking
by: Nguyen-Quang, Thuc, et al.
Published: (2024)
by: Nguyen-Quang, Thuc, et al.
Published: (2024)
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
by: Tran, Quoc-Khang, et al.
Published: (2026)
by: Tran, Quoc-Khang, et al.
Published: (2026)
ConstStyle: Robust Domain Generalization with Unified Style Transformation
by: Tran, Nam Duong, et al.
Published: (2025)
by: Tran, Nam Duong, et al.
Published: (2025)
Anti-I2V: Safeguarding your photos from malicious image-to-video generation
by: Vu, Duc, et al.
Published: (2026)
by: Vu, Duc, et al.
Published: (2026)
ControlVAR: Exploring Controllable Visual Autoregressive Modeling
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
V-Math: An Agentic Approach to the Vietnamese National High School Graduation Mathematics Exams
by: Nguyen, Duong Q., et al.
Published: (2025)
by: Nguyen, Duong Q., et al.
Published: (2025)
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
by: Pham, Huy Quang, et al.
Published: (2024)
by: Pham, Huy Quang, et al.
Published: (2024)
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
by: Tran, Huyen T. T., et al.
Published: (2026)
by: Tran, Huyen T. T., et al.
Published: (2026)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization
by: Nguyen, Quang Vinh, et al.
Published: (2024)
by: Nguyen, Quang Vinh, et al.
Published: (2024)
Training Deep Visual Networks Beyond Loss and Accuracy Through a Dynamical Systems Approach
by: La Quang, Hai, et al.
Published: (2026)
by: La Quang, Hai, et al.
Published: (2026)
Contrastive Integrated Gradients: A Feature Attribution-Based Method for Explaining Whole Slide Image Classification
by: Vu, Anh Mai, et al.
Published: (2025)
by: Vu, Anh Mai, et al.
Published: (2025)
Symmetric masking strategy enhances the performance of Masked Image Modeling
by: Nguyen, Khanh-Binh, et al.
Published: (2024)
by: Nguyen, Khanh-Binh, et al.
Published: (2024)
Retro: Reusing teacher projection head for efficient embedding distillation on Lightweight Models via Self-supervised Learning
by: Nguyen, Khanh-Binh, et al.
Published: (2024)
by: Nguyen, Khanh-Binh, et al.
Published: (2024)
Leveraging Model Soups to Classify Intangible Cultural Heritage Images from the Mekong Delta
by: Tran, Quoc-Khang, et al.
Published: (2026)
by: Tran, Quoc-Khang, et al.
Published: (2026)
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
by: Nguyen, Hung Huy, et al.
Published: (2025)
by: Nguyen, Hung Huy, et al.
Published: (2025)
Similar Items
-
SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
by: Nguyen, Hung, et al.
Published: (2024) -
OE3DIS: Open-Ended 3D Point Cloud Instance Segmentation
by: Nguyen, Phuc D. A., et al.
Published: (2024) -
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments
by: Nguyen, Van Quang
Published: (2026) -
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
by: Nguyen, Toan, et al.
Published: (2025) -
TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning
by: Dinh, Quang Minh, et al.
Published: (2024)