Frequency-Aware Flow Matching for High-Quality Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Ren, Sucheng, Yu, Qihang, He, Ju, Shen, Xiaohui, Yuille, Alan, Chen, Liang-Chieh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
di: Chen, Jieneng, et al.
Pubblicazione: (2024)
di: Chen, Jieneng, et al.
Pubblicazione: (2024)
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
di: He, Ju, et al.
Pubblicazione: (2023)
di: He, Ju, et al.
Pubblicazione: (2023)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
di: Yang, Timing, et al.
Pubblicazione: (2025)
di: Yang, Timing, et al.
Pubblicazione: (2025)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
di: He, Ju, et al.
Pubblicazione: (2025)
di: He, Ju, et al.
Pubblicazione: (2025)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
di: Kim, Dongwon, et al.
Pubblicazione: (2025)
di: Kim, Dongwon, et al.
Pubblicazione: (2025)
Randomized Autoregressive Visual Generation
di: Yu, Qihang, et al.
Pubblicazione: (2024)
di: Yu, Qihang, et al.
Pubblicazione: (2024)
Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization
di: Liu, Qihao, et al.
Pubblicazione: (2024)
di: Liu, Qihao, et al.
Pubblicazione: (2024)
Dictionary-based Framework for Interpretable and Consistent Object Parsing
di: Zhang, Tiezheng, et al.
Pubblicazione: (2025)
di: Zhang, Tiezheng, et al.
Pubblicazione: (2025)
Large Language Models are Universal Reasoners for Visual Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2026)
di: Ren, Sucheng, et al.
Pubblicazione: (2026)
An Image is Worth 32 Tokens for Reconstruction and Generation
di: Yu, Qihang, et al.
Pubblicazione: (2024)
di: Yu, Qihang, et al.
Pubblicazione: (2024)
Autoregressive Image Generation with Masked Bit Modeling
di: Yu, Qihang, et al.
Pubblicazione: (2026)
di: Yu, Qihang, et al.
Pubblicazione: (2026)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
COCONut: Modernizing COCO Segmentation
di: Deng, Xueqing, et al.
Pubblicazione: (2024)
di: Deng, Xueqing, et al.
Pubblicazione: (2024)
MaskBit: Embedding-free Image Generation via Bit Tokens
di: Weber, Mark, et al.
Pubblicazione: (2024)
di: Weber, Mark, et al.
Pubblicazione: (2024)
Autoregressive Video Generation beyond Next Frames Prediction
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
Rejuvenating image-GPT as Strong Visual Representation Learners
di: Ren, Sucheng, et al.
Pubblicazione: (2023)
di: Ren, Sucheng, et al.
Pubblicazione: (2023)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
HResFormer: Hybrid Residual Transformer for Volumetric Medical Image Segmentation
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
SPFormer: Enhancing Vision Transformer with Superpixel Representation
di: Mei, Jieru, et al.
Pubblicazione: (2024)
di: Mei, Jieru, et al.
Pubblicazione: (2024)
Quality Sentinel: Estimating Label Quality and Errors in Medical Segmentation Datasets
di: Chen, Yixiong, et al.
Pubblicazione: (2024)
di: Chen, Yixiong, et al.
Pubblicazione: (2024)
ViT-5: Vision Transformers for The Mid-2020s
di: Wang, Feng, et al.
Pubblicazione: (2026)
di: Wang, Feng, et al.
Pubblicazione: (2026)
MedFlowSeg: Flow Matching for Medical Image Segmentation with Frequency-Aware Attention
di: Chen, Zhi, et al.
Pubblicazione: (2026)
di: Chen, Zhi, et al.
Pubblicazione: (2026)
Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting
di: Shin, Inkyu, et al.
Pubblicazione: (2024)
di: Shin, Inkyu, et al.
Pubblicazione: (2024)
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
di: Deng, Xueqing, et al.
Pubblicazione: (2025)
di: Deng, Xueqing, et al.
Pubblicazione: (2025)
FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
di: Lin, Mingfeng, et al.
Pubblicazione: (2026)
di: Lin, Mingfeng, et al.
Pubblicazione: (2026)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
di: Liu, Haoyu, et al.
Pubblicazione: (2026)
di: Liu, Haoyu, et al.
Pubblicazione: (2026)
Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
di: Park, Dogyun, et al.
Pubblicazione: (2025)
di: Park, Dogyun, et al.
Pubblicazione: (2025)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
di: Kerssies, Tommie, et al.
Pubblicazione: (2026)
di: Kerssies, Tommie, et al.
Pubblicazione: (2026)
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
di: Lin, Wang, et al.
Pubblicazione: (2026)
di: Lin, Wang, et al.
Pubblicazione: (2026)
Mamba-R: Vision Mamba ALSO Needs Registers
di: Wang, Feng, et al.
Pubblicazione: (2024)
di: Wang, Feng, et al.
Pubblicazione: (2024)
Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency
di: Wang, Feng, et al.
Pubblicazione: (2024)
di: Wang, Feng, et al.
Pubblicazione: (2024)
Geometry-Aware Image Flow Matching
di: Lee, Junho, et al.
Pubblicazione: (2026)
di: Lee, Junho, et al.
Pubblicazione: (2026)
Time-reversed Flow Matching with Worst Transport in High-dimensional Latent Space for Image Anomaly Detection
di: Li, Liangwei, et al.
Pubblicazione: (2025)
di: Li, Liangwei, et al.
Pubblicazione: (2025)
Deeply Supervised Flow-Based Generative Models
di: Shin, Inkyu, et al.
Pubblicazione: (2025)
di: Shin, Inkyu, et al.
Pubblicazione: (2025)
Frequency-Aware Density Control via Reparameterization for High-Quality Rendering of 3D Gaussian Splatting
di: Zeng, Zhaojie, et al.
Pubblicazione: (2025)
di: Zeng, Zhaojie, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
di: Ren, Sucheng, et al.
Pubblicazione: (2024) -
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2025) -
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
di: Ren, Sucheng, et al.
Pubblicazione: (2025) -
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
di: Chen, Jieneng, et al.
Pubblicazione: (2024) -
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2024)