Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
Fuente:
arXiv
Saved in:
| Main Authors: | Xiang, Xunzhi, Tian, Xingye, Zhang, Guiyu, Chen, Yabo, Zhang, Shaofeng, Wang, Xuebo, Tao, Xin, Fan, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
by: Zhang, Guiyu, et al.
Published: (2026)
by: Zhang, Guiyu, et al.
Published: (2026)
Pathwise Test-Time Correction for Autoregressive Long Video Generation
by: Xiang, Xunzhi, et al.
Published: (2026)
by: Xiang, Xunzhi, et al.
Published: (2026)
Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation
by: Ma, Junyuan, et al.
Published: (2026)
by: Ma, Junyuan, et al.
Published: (2026)
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
by: Xiang, Xunzhi, et al.
Published: (2025)
by: Xiang, Xunzhi, et al.
Published: (2025)
Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
by: Xiang, Xunzhi, et al.
Published: (2025)
by: Xiang, Xunzhi, et al.
Published: (2025)
PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models
by: Su, Yuanhao, et al.
Published: (2026)
by: Su, Yuanhao, et al.
Published: (2026)
Decoupling Complexity from Scale in Latent Diffusion Model
by: Zhong, Tianxiong, et al.
Published: (2025)
by: Zhong, Tianxiong, et al.
Published: (2025)
Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
by: Zhang, Guiyu, et al.
Published: (2025)
by: Zhang, Guiyu, et al.
Published: (2025)
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
by: Zhong, Tianxiong, et al.
Published: (2025)
by: Zhong, Tianxiong, et al.
Published: (2025)
Denoising Vision Transformers
by: Yang, Jiawei, et al.
Published: (2024)
by: Yang, Jiawei, et al.
Published: (2024)
Measuring and Controlling the Spectral Bias for Self-Supervised Image Denoising
by: Zhang, Wang, et al.
Published: (2025)
by: Zhang, Wang, et al.
Published: (2025)
Step-level Denoising-time Diffusion Alignment with Multiple Objectives
by: Zhang, Qi, et al.
Published: (2026)
by: Zhang, Qi, et al.
Published: (2026)
PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders
by: Zhang, Xiangdong, et al.
Published: (2024)
by: Zhang, Xiangdong, et al.
Published: (2024)
Fréchet Denoised Distance: Enhancing Plausibility Evaluation for Generated Designs with Denoising Autoencoder
by: Fan, Jiajie, et al.
Published: (2024)
by: Fan, Jiajie, et al.
Published: (2024)
Point-DAE: Denoising Autoencoders for Self-supervised Point Cloud Learning
by: Zhang, Yabin, et al.
Published: (2022)
by: Zhang, Yabin, et al.
Published: (2022)
Approximate Size Targets Are Sufficient for Accurate Semantic Segmentation
by: Fan, Xingye, et al.
Published: (2025)
by: Fan, Xingye, et al.
Published: (2025)
Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution
by: Wang, Jingkai, et al.
Published: (2026)
by: Wang, Jingkai, et al.
Published: (2026)
Flow-NeRF: Joint Learning of Geometry, Poses, and Dense Flow within Unified Neural Representations
by: Zheng, Xunzhi, et al.
Published: (2025)
by: Zheng, Xunzhi, et al.
Published: (2025)
Representing Noisy Image Without Denoising
by: Qi, Shuren, et al.
Published: (2023)
by: Qi, Shuren, et al.
Published: (2023)
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
by: Hu, Jiahao, et al.
Published: (2024)
by: Hu, Jiahao, et al.
Published: (2024)
Spatiotemporal Blind-Spot Network with Calibrated Flow Alignment for Self-Supervised Video Denoising
by: Chen, Zikang, et al.
Published: (2024)
by: Chen, Zikang, et al.
Published: (2024)
Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
by: Zhang, Xiangdong, et al.
Published: (2025)
by: Zhang, Xiangdong, et al.
Published: (2025)
SGDFormer: One-stage Transformer-based Architecture for Cross-Spectral Stereo Image Guided Denoising
by: Zhang, Runmin, et al.
Published: (2024)
by: Zhang, Runmin, et al.
Published: (2024)
TESSER: Transfer-Enhancing Adversarial Attacks from Vision Transformers via Spectral and Semantic Regularization
by: Guesmi, Amira, et al.
Published: (2025)
by: Guesmi, Amira, et al.
Published: (2025)
Pattern Integration and Enhancement Vision Transformer for Self-Supervised Learning in Remote Sensing
by: Lu, Kaixuan, et al.
Published: (2024)
by: Lu, Kaixuan, et al.
Published: (2024)
Hadamard Attention Recurrent Transformer: A Strong Baseline for Stereo Matching Transformer
by: Chen, Ziyang, et al.
Published: (2025)
by: Chen, Ziyang, et al.
Published: (2025)
SIEFormer: Spectral-Interpretable and -Enhanced Transformer for Generalized Category Discovery
by: Li, Chunming, et al.
Published: (2026)
by: Li, Chunming, et al.
Published: (2026)
Semantic Graph Consistency: Going Beyond Patches for Regularizing Self-Supervised Vision Transformers
by: Devaguptapu, Chaitanya, et al.
Published: (2024)
by: Devaguptapu, Chaitanya, et al.
Published: (2024)
Teacher-Guided Causal Interventions for Image Denoising: Orthogonal Content-Noise Disentanglement in Vision Transformers
by: Jiang, Kuai, et al.
Published: (2026)
by: Jiang, Kuai, et al.
Published: (2026)
DRACO: A Denoising-Reconstruction Autoencoder for Cryo-EM
by: Shen, Yingjun, et al.
Published: (2024)
by: Shen, Yingjun, et al.
Published: (2024)
Self-Supervised Vision Transformer for Enhanced Virtual Clothes Try-On
by: Lu, Lingxiao, et al.
Published: (2024)
by: Lu, Lingxiao, et al.
Published: (2024)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
Vision Transformer with Sparse Scan Prior
by: Zhang, Yuguang, et al.
Published: (2024)
by: Zhang, Yuguang, et al.
Published: (2024)
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
by: Ge, Wenhang, et al.
Published: (2026)
by: Ge, Wenhang, et al.
Published: (2026)
Cascade-Zero123: One Image to Highly Consistent 3D with Self-Prompted Nearby Views
by: Chen, Yabo, et al.
Published: (2023)
by: Chen, Yabo, et al.
Published: (2023)
Metric-Solver: Sliding Anchored Metric Depth Estimation from a Single Image
by: Wen, Tao, et al.
Published: (2025)
by: Wen, Tao, et al.
Published: (2025)
Latent Denoising Makes Good Tokenizers
by: Yang, Jiawei, et al.
Published: (2025)
by: Yang, Jiawei, et al.
Published: (2025)
Entropy Reveals Block Importance in Masked Self-Supervised Vision Transformers
by: Xiang, Peihao, et al.
Published: (2026)
by: Xiang, Peihao, et al.
Published: (2026)
Xformer: Hybrid X-Shaped Transformer for Image Denoising
by: Zhang, Jiale, et al.
Published: (2023)
by: Zhang, Jiale, et al.
Published: (2023)
Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection
by: Zhu, Wenjie, et al.
Published: (2025)
by: Zhu, Wenjie, et al.
Published: (2025)
Similar Items
-
SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
by: Zhang, Guiyu, et al.
Published: (2026) -
Pathwise Test-Time Correction for Autoregressive Long Video Generation
by: Xiang, Xunzhi, et al.
Published: (2026) -
Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation
by: Ma, Junyuan, et al.
Published: (2026) -
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
by: Xiang, Xunzhi, et al.
Published: (2025) -
Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
by: Xiang, Xunzhi, et al.
Published: (2025)