DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Tianhang, Chen, Yitong, Song, Wei, Wu, Zuxuan, Li, Min, Wang, Jiaqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Channel-wise Vector Quantization
by: Song, Wei, et al.
Published: (2026)
by: Song, Wei, et al.
Published: (2026)
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization
by: Liu, Zhuohan, et al.
Published: (2026)
by: Liu, Zhuohan, et al.
Published: (2026)
VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
by: Du, Sinan, et al.
Published: (2025)
by: Du, Sinan, et al.
Published: (2025)
Improving Reconstruction of Representation Autoencoder
by: Liu, Siyu, et al.
Published: (2026)
by: Liu, Siyu, et al.
Published: (2026)
3D Surface Reconstruction with Enhanced High-Frequency Details
by: Zhang, Shikun, et al.
Published: (2025)
by: Zhang, Shikun, et al.
Published: (2025)
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
by: Chen, Yitong, et al.
Published: (2026)
by: Chen, Yitong, et al.
Published: (2026)
Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval
by: Zou, Zichen, et al.
Published: (2026)
by: Zou, Zichen, et al.
Published: (2026)
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
by: Zhang, Hui, et al.
Published: (2024)
by: Zhang, Hui, et al.
Published: (2024)
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
by: Liu, Huaize, et al.
Published: (2025)
by: Liu, Huaize, et al.
Published: (2025)
Monocular Online Reconstruction with Enhanced Detail Preservation
by: Wu, Songyin, et al.
Published: (2025)
by: Wu, Songyin, et al.
Published: (2025)
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
by: Chen, Yitong, et al.
Published: (2025)
by: Chen, Yitong, et al.
Published: (2025)
RPiAE: A Representation-Pivoted Autoencoder Enhancing Both Image Generation and Editing
by: Gong, Yue, et al.
Published: (2026)
by: Gong, Yue, et al.
Published: (2026)
Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection
by: Chen, Yitong, et al.
Published: (2024)
by: Chen, Yitong, et al.
Published: (2024)
Indoor Scene Reconstruction with Fine-Grained Details Using Hybrid Representation and Normal Prior Enhancement
by: Ye, Sheng, et al.
Published: (2023)
by: Ye, Sheng, et al.
Published: (2023)
Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE
by: Bao, Wei, et al.
Published: (2026)
by: Bao, Wei, et al.
Published: (2026)
UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing
by: Wang, Dianyi, et al.
Published: (2026)
by: Wang, Dianyi, et al.
Published: (2026)
Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control
by: Li, Danfeng, et al.
Published: (2025)
by: Li, Danfeng, et al.
Published: (2025)
GeoQuery: Geometry-Query Diffusion for Sparse-View Reconstruction
by: Cao, Xiao, et al.
Published: (2026)
by: Cao, Xiao, et al.
Published: (2026)
Long-LRM++: Preserving Fine Details in Feed-Forward Wide-Coverage Reconstruction
by: Ziwen, Chen, et al.
Published: (2025)
by: Ziwen, Chen, et al.
Published: (2025)
DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
by: Liu, Yiheng, et al.
Published: (2025)
by: Liu, Yiheng, et al.
Published: (2025)
Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging
by: Li, Yitong, et al.
Published: (2025)
by: Li, Yitong, et al.
Published: (2025)
DeRA: Decoupled Representation Alignment for Video Tokenization
by: Guo, Pengbo, et al.
Published: (2025)
by: Guo, Pengbo, et al.
Published: (2025)
Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
by: You, Zuyao, et al.
Published: (2025)
by: You, Zuyao, et al.
Published: (2025)
OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
by: Gao, Sensen, et al.
Published: (2026)
by: Gao, Sensen, et al.
Published: (2026)
DSDRNet: Disentangling Representation and Reconstruct Network for Domain Generalization
by: Yang, Juncheng, et al.
Published: (2024)
by: Yang, Juncheng, et al.
Published: (2024)
SARMAE: Masked Autoencoder for SAR Representation Learning
by: Liu, Danxu, et al.
Published: (2025)
by: Liu, Danxu, et al.
Published: (2025)
ReTracing: An Archaeological Approach Through Body, Machine, and Generative Systems
by: Wang, Yitong, et al.
Published: (2026)
by: Wang, Yitong, et al.
Published: (2026)
GeoGS3D: Single-view 3D Reconstruction via Geometric-aware Diffusion Model and Gaussian Splatting
by: Feng, Qijun, et al.
Published: (2024)
by: Feng, Qijun, et al.
Published: (2024)
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
by: Qian, Chengxuan, et al.
Published: (2025)
by: Qian, Chengxuan, et al.
Published: (2025)
QEMesh: Employing A Quadric Error Metrics-Based Representation for Mesh Generation
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
Hybrid Mesh-Gaussian Representation for Efficient Indoor Scene Reconstruction
by: Huang, Binxiao, et al.
Published: (2025)
by: Huang, Binxiao, et al.
Published: (2025)
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
by: Yang, Zhenhao, et al.
Published: (2026)
by: Yang, Zhenhao, et al.
Published: (2026)
Unified Scene Representation and Reconstruction for 3D Large Language Models
by: Chu, Tao, et al.
Published: (2024)
by: Chu, Tao, et al.
Published: (2024)
UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph Generation
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
Eliminating VAE for Fast and High-Resolution Generative Detail Restoration
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization
by: Zheng, Peng, et al.
Published: (2025)
by: Zheng, Peng, et al.
Published: (2025)
Preference Score Distillation: Leveraging 2D Rewards to Align Text-to-3D Generation with Human Preference
by: Leng, Jiaqi, et al.
Published: (2026)
by: Leng, Jiaqi, et al.
Published: (2026)
Sketch-1-to-3: One Single Sketch to 3D Detailed Face Reconstruction
by: Wen, Liting, et al.
Published: (2025)
by: Wen, Liting, et al.
Published: (2025)
MaxQ: Multi-Axis Query for N:M Sparsity Network
by: Xiang, Jingyang, et al.
Published: (2023)
by: Xiang, Jingyang, et al.
Published: (2023)
GeoMVD: Geometry-Enhanced Multi-View Generation Model Based on Geometric Information Extraction
by: Wu, Jiaqi, et al.
Published: (2025)
by: Wu, Jiaqi, et al.
Published: (2025)
Similar Items
-
Channel-wise Vector Quantization
by: Song, Wei, et al.
Published: (2026) -
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization
by: Liu, Zhuohan, et al.
Published: (2026) -
VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
by: Du, Sinan, et al.
Published: (2025) -
Improving Reconstruction of Representation Autoencoder
by: Liu, Siyu, et al.
Published: (2026) -
3D Surface Reconstruction with Enhanced High-Frequency Details
by: Zhang, Shikun, et al.
Published: (2025)