Ultra-Resolution Adaptation with Ease
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Ruonan, Liu, Songhua, Tan, Zhenxiong, Wang, Xinchao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
by: Yu, Ruonan, et al.
Published: (2026)
by: Yu, Ruonan, et al.
Published: (2026)
CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up
by: Liu, Songhua, et al.
Published: (2024)
by: Liu, Songhua, et al.
Published: (2024)
LinFusion: 1 GPU, 1 Minute, 16K Image
by: Liu, Songhua, et al.
Published: (2024)
by: Liu, Songhua, et al.
Published: (2024)
MindBridge: A Cross-Subject Brain Decoding Framework
by: Wang, Shizun, et al.
Published: (2024)
by: Wang, Shizun, et al.
Published: (2024)
Video-Infinity: Distributed Long Video Generation
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
by: Liu, Yuhe, et al.
Published: (2026)
by: Liu, Yuhe, et al.
Published: (2026)
Image Editing As Programs with Diffusion Models
by: Hu, Yujia, et al.
Published: (2025)
by: Hu, Yujia, et al.
Published: (2025)
Teddy: Efficient Large-Scale Dataset Distillation via Taylor-Approximated Matching
by: Yu, Ruonan, et al.
Published: (2024)
by: Yu, Ruonan, et al.
Published: (2024)
FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation
by: Wu, Yunfeng, et al.
Published: (2025)
by: Wu, Yunfeng, et al.
Published: (2025)
SpotEdit: Selective Region Editing in Diffusion Transformers
by: Qin, Zhibin, et al.
Published: (2025)
by: Qin, Zhibin, et al.
Published: (2025)
Vision Bridge Transformer at Scale
by: Tan, Zhenxiong, et al.
Published: (2025)
by: Tan, Zhenxiong, et al.
Published: (2025)
OminiControl2: Efficient Conditioning for Diffusion Transformers
by: Tan, Zhenxiong, et al.
Published: (2025)
by: Tan, Zhenxiong, et al.
Published: (2025)
Distilled Datamodel with Reverse Gradient Matching
by: Ye, Jingwen, et al.
Published: (2024)
by: Ye, Jingwen, et al.
Published: (2024)
OminiControl: Minimal and Universal Control for Diffusion Transformer
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
Heavy Labels Out! Dataset Distillation with Label Space Lightening
by: Yu, Ruonan, et al.
Published: (2024)
by: Yu, Ruonan, et al.
Published: (2024)
CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation
by: Zhou, Letian, et al.
Published: (2025)
by: Zhou, Letian, et al.
Published: (2025)
Understanding Dataset Distillation via Spectral Filtering
by: Bo, Deyu, et al.
Published: (2025)
by: Bo, Deyu, et al.
Published: (2025)
One-shot Federated Learning via Synthetic Distiller-Distillate Communication
by: Zhang, Junyuan, et al.
Published: (2024)
by: Zhang, Junyuan, et al.
Published: (2024)
Control and Realism: Best of Both Worlds in Layout-to-Image without Training
by: Li, Bonan, et al.
Published: (2025)
by: Li, Bonan, et al.
Published: (2025)
Flash Sculptor: Modular 3D Worlds from Objects
by: Hu, Yujia, et al.
Published: (2025)
by: Hu, Yujia, et al.
Published: (2025)
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
by: li, Bonan, et al.
Published: (2025)
by: li, Bonan, et al.
Published: (2025)
Minute-Long Videos with Dual Parallelisms
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images
by: Wu, Yunfeng, et al.
Published: (2026)
by: Wu, Yunfeng, et al.
Published: (2026)
AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
by: Chen, Zigeng, et al.
Published: (2024)
by: Chen, Zigeng, et al.
Published: (2024)
Rethinking Large-scale Dataset Compression: Shifting Focus From Labels to Images
by: Xiao, Lingao, et al.
Published: (2025)
by: Xiao, Lingao, et al.
Published: (2025)
StyDeSty: Min-Max Stylization and Destylization for Single Domain Generalization
by: Liu, Songhua, et al.
Published: (2024)
by: Liu, Songhua, et al.
Published: (2024)
Neural Lineage
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
Dual-Path Adversarial Lifting for Domain Shift Correction in Online Test-time Adaptation
by: Tang, Yushun, et al.
Published: (2024)
by: Tang, Yushun, et al.
Published: (2024)
Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding
by: Yu, Runpeng, et al.
Published: (2025)
by: Yu, Runpeng, et al.
Published: (2025)
Encapsulating Knowledge in One Prompt
by: Li, Qi, et al.
Published: (2024)
by: Li, Qi, et al.
Published: (2024)
Hash3D: Training-free Acceleration for 3D Generation
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
Language Model as Visual Explainer
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic Reasoning
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Compositional Video Generation as Flow Equalization
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
Domain-Adaptive 2D Human Pose Estimation via Dual Teachers in Extremely Low-Light Conditions
by: Ai, Yihao, et al.
Published: (2024)
by: Ai, Yihao, et al.
Published: (2024)
UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion Space
by: Liu, Yong, et al.
Published: (2025)
by: Liu, Yong, et al.
Published: (2025)
PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
UltraGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
by: Zhang, Yuyao, et al.
Published: (2025)
by: Zhang, Yuyao, et al.
Published: (2025)
ViMU: Benchmarking Video Metaphorical Understanding
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Attention Prompting on Image for Large Vision-Language Models
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
Similar Items
-
ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
by: Yu, Ruonan, et al.
Published: (2026) -
CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up
by: Liu, Songhua, et al.
Published: (2024) -
LinFusion: 1 GPU, 1 Minute, 16K Image
by: Liu, Songhua, et al.
Published: (2024) -
MindBridge: A Cross-Subject Brain Decoding Framework
by: Wang, Shizun, et al.
Published: (2024) -
Video-Infinity: Distributed Long Video Generation
by: Tan, Zhenxiong, et al.
Published: (2024)