Native-Resolution Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zidong, Bai, Lei, Yue, Xiangyu, Ouyang, Wanli, Zhang, Yiyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transition Models: Rethinking the Generative Learning Objective
by: Wang, Zidong, et al.
Published: (2025)
by: Wang, Zidong, et al.
Published: (2025)
Scaling Up Your Kernels: Large Kernel Design in ConvNets towards Universal Representations
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
EMR-Merging: Tuning-Free High-Performance Model Merging
by: Huang, Chenyu, et al.
Published: (2024)
by: Huang, Chenyu, et al.
Published: (2024)
Dynamic Base model Shift for Delta Compression
by: Huang, Chenyu, et al.
Published: (2025)
by: Huang, Chenyu, et al.
Published: (2025)
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
by: Yue, Xiaoyu, et al.
Published: (2025)
by: Yue, Xiaoyu, et al.
Published: (2025)
PredBench: Benchmarking Spatio-Temporal Prediction across Diverse Disciplines
by: Wang, ZiDong, et al.
Published: (2024)
by: Wang, ZiDong, et al.
Published: (2024)
Diffusion Models Need Visual Priors for Image Generation
by: Yue, Xiaoyu, et al.
Published: (2024)
by: Yue, Xiaoyu, et al.
Published: (2024)
Explore the Limits of Omni-modal Pretraining at Scale
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
by: Zhang, Zhixin, et al.
Published: (2024)
by: Zhang, Zhixin, et al.
Published: (2024)
Multimodal Long Video Modeling Based on Temporal Dynamic Context
by: Hao, Haoran, et al.
Published: (2025)
by: Hao, Haoran, et al.
Published: (2025)
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
by: Ding, Xiaohan, et al.
Published: (2023)
by: Ding, Xiaohan, et al.
Published: (2023)
Seeing Delta Parameters as JPEG Images: Data-Free Delta Compression with Discrete Cosine Transform
by: Huang, Chenyu, et al.
Published: (2025)
by: Huang, Chenyu, et al.
Published: (2025)
Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression
by: Wang, Xiaohui, et al.
Published: (2025)
by: Wang, Xiaohui, et al.
Published: (2025)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
High-Resolution Image Synthesis via Next-Token Prediction
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
by: Chen, Zhuokun, et al.
Published: (2025)
by: Chen, Zhuokun, et al.
Published: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval
by: Ouyang, Pengxiang, et al.
Published: (2025)
by: Ouyang, Pengxiang, et al.
Published: (2025)
GUPNet++: Geometry Uncertainty Propagation Network for Monocular 3D Object Detection
by: Lu, Yan, et al.
Published: (2023)
by: Lu, Yan, et al.
Published: (2023)
FiT: Flexible Vision Transformer for Diffusion Model
by: Lu, Zeyu, et al.
Published: (2024)
by: Lu, Zeyu, et al.
Published: (2024)
Prompt Learning for Oriented Power Transmission Tower Detection in High-Resolution SAR Images
by: Li, Tianyang, et al.
Published: (2024)
by: Li, Tianyang, et al.
Published: (2024)
ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
by: Xue, Xiangyuan, et al.
Published: (2024)
by: Xue, Xiangyuan, et al.
Published: (2024)
Generative 3D Gaussian Splatting for Arbitrary-ResolutionAtmospheric Downscaling and Forecasting
by: Han, Tao, et al.
Published: (2026)
by: Han, Tao, et al.
Published: (2026)
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
by: Serra, Alessandro Pietro, et al.
Published: (2024)
by: Serra, Alessandro Pietro, et al.
Published: (2024)
CASISR: Circular Arbitrary-Scale Image Super-Resolution
by: Li, Honggui, et al.
Published: (2026)
by: Li, Honggui, et al.
Published: (2026)
Pretrained Image-Text Models are Secretly Video Captioners
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
BIFRÖST: 3D-Aware Image compositing with Language Instructions
by: Li, Lingxiao, et al.
Published: (2024)
by: Li, Lingxiao, et al.
Published: (2024)
Generalizable Geometric Image Caption Synthesis
by: Xin, Yue, et al.
Published: (2025)
by: Xin, Yue, et al.
Published: (2025)
WeatherGFM: Learning A Weather Generalist Foundation Model via In-context Learning
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning
by: Yang, Suorong, et al.
Published: (2025)
by: Yang, Suorong, et al.
Published: (2025)
An Adaptor for Triggering Semi-Supervised Learning to Out-of-Box Serve Deep Image Clustering
by: Duan, Yue, et al.
Published: (2025)
by: Duan, Yue, et al.
Published: (2025)
Native Segmentation Vision Transformers
by: Brasó, Guillem, et al.
Published: (2025)
by: Brasó, Guillem, et al.
Published: (2025)
Accessing Vision Foundation Models via ImageNet-1K
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Block Flow: Learning Straight Flow on Data Blocks
by: Wang, Zibin, et al.
Published: (2025)
by: Wang, Zibin, et al.
Published: (2025)
Approximating Signed Distance Fields With Sparse Ellipsoidal Radial Basis Function Networks: A Dynamic Multi-Objective Optimization Strategy
by: Lian, Bobo, et al.
Published: (2025)
by: Lian, Bobo, et al.
Published: (2025)
Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
by: Zhang, Mingyuan, et al.
Published: (2025)
by: Zhang, Mingyuan, et al.
Published: (2025)
Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers
by: Crowson, Katherine, et al.
Published: (2024)
by: Crowson, Katherine, et al.
Published: (2024)
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
Pixel Distillation: A New Knowledge Distillation Scheme for Low-Resolution Image Recognition
by: Guo, Guangyu, et al.
Published: (2021)
by: Guo, Guangyu, et al.
Published: (2021)
HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
Similar Items
-
Transition Models: Rethinking the Generative Learning Objective
by: Wang, Zidong, et al.
Published: (2025) -
Scaling Up Your Kernels: Large Kernel Design in ConvNets towards Universal Representations
by: Zhang, Yiyuan, et al.
Published: (2024) -
EMR-Merging: Tuning-Free High-Performance Model Merging
by: Huang, Chenyu, et al.
Published: (2024) -
Dynamic Base model Shift for Delta Compression
by: Huang, Chenyu, et al.
Published: (2025) -
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
by: Yue, Xiaoyu, et al.
Published: (2025)