Exploring Scalable Unified Modeling for General Low-Level Vision
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xiangyu, Zhu, Kaiwen, Pu, Yuandong, Cao, Shuo, Li, Xiaohui, Zhang, Wenlong, Liu, Yihao, Qiao, Yu, Zhou, Jiantao, Dong, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning A Low-Level Vision Generalist via Visual Task Prompt
by: Chen, Xiangyu, et al.
Published: (2024)
by: Chen, Xiangyu, et al.
Published: (2024)
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
A Comparative Study of Image Restoration Networks for General Backbone Network Design
by: Chen, Xiangyu, et al.
Published: (2023)
by: Chen, Xiangyu, et al.
Published: (2023)
UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
by: Cao, Shuo, et al.
Published: (2025)
by: Cao, Shuo, et al.
Published: (2025)
Accelerating Masked Image Generation by Learning Latent Controlled Dynamics
by: Zhu, Kaiwen, et al.
Published: (2026)
by: Zhu, Kaiwen, et al.
Published: (2026)
GRIDS: Grouped Multiple-Degradation Restoration with Image Degradation Similarity
by: Cao, Shuo, et al.
Published: (2024)
by: Cao, Shuo, et al.
Published: (2024)
Unifying Image Processing as Visual Prompting Question Answering
by: Liu, Yihao, et al.
Published: (2023)
by: Liu, Yihao, et al.
Published: (2023)
Toward Generalizable Deblurring: Leveraging Massive Blur Priors with Linear Attention for Real-World Scenarios
by: Gao, Yuanting, et al.
Published: (2026)
by: Gao, Yuanting, et al.
Published: (2026)
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
by: Cao, Shuo, et al.
Published: (2025)
by: Cao, Shuo, et al.
Published: (2025)
HAT: Hybrid Attention Transformer for Image Restoration
by: Chen, Xiangyu, et al.
Published: (2023)
by: Chen, Xiangyu, et al.
Published: (2023)
A Preliminary Exploration Towards General Image Restoration
by: Kong, Xiangtao, et al.
Published: (2024)
by: Kong, Xiangtao, et al.
Published: (2024)
PICABench: How Far Are We from Physically Realistic Image Editing?
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
DualX-VSR: Dual Axial Spatial$\times$Temporal Transformer for Real-World Video Super-Resolution without Motion Compensation
by: Cao, Shuo, et al.
Published: (2025)
by: Cao, Shuo, et al.
Published: (2025)
Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and Generation
by: Zhou, Zhiwang, et al.
Published: (2025)
by: Zhou, Zhiwang, et al.
Published: (2025)
SEAL: A Framework for Systematic Evaluation of Real-World Super-Resolution
by: Zhang, Wenlong, et al.
Published: (2023)
by: Zhang, Wenlong, et al.
Published: (2023)
LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution
by: Li, Xiaohui, et al.
Published: (2025)
by: Li, Xiaohui, et al.
Published: (2025)
StableI2I: Spotting Unintended Changes in Image-to-Image Transition
by: Li, Jiayang, et al.
Published: (2026)
by: Li, Jiayang, et al.
Published: (2026)
Revisiting the Generalization Problem of Low-level Vision Models Through the Lens of Image Deraining
by: Hu, Jinfan, et al.
Published: (2025)
by: Hu, Jinfan, et al.
Published: (2025)
DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex Degradations
by: Li, Xiaohui, et al.
Published: (2025)
by: Li, Xiaohui, et al.
Published: (2025)
Towards Efficient SDRTV-to-HDRTV by Learning from Image Formation
by: Chen, Xiangyu, et al.
Published: (2023)
by: Chen, Xiangyu, et al.
Published: (2023)
SynWeather: Weather Observation Data Synthesis across Multiple Regions and Variables via a General Diffusion Transformer
by: Xu, Kaiyi, et al.
Published: (2025)
by: Xu, Kaiyi, et al.
Published: (2025)
UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
An Intelligent Agentic System for Complex Image Restoration Problems
by: Zhu, Kaiwen, et al.
Published: (2024)
by: Zhu, Kaiwen, et al.
Published: (2024)
MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?
by: Wang, Yuandong, et al.
Published: (2025)
by: Wang, Yuandong, et al.
Published: (2025)
A Unified Perspective on Adversarial Membership Manipulation in Vision Models
by: Gao, Ruize, et al.
Published: (2026)
by: Gao, Ruize, et al.
Published: (2026)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
by: Li, Teng, et al.
Published: (2025)
by: Li, Teng, et al.
Published: (2025)
GRN+: A Simplified Generative Reinforcement Network for Tissue Layer Analysis in 3D Ultrasound Images for Chronic Low-back Pain
by: Zeng, Zixue, et al.
Published: (2025)
by: Zeng, Zixue, et al.
Published: (2025)
Rethinking Low-Rank Adaptation in Vision: Exploring Head-Level Responsiveness across Diverse Tasks
by: Zhong, Yibo, et al.
Published: (2024)
by: Zhong, Yibo, et al.
Published: (2024)
L4P: Towards Unified Low-Level 4D Vision Perception
by: Badki, Abhishek, et al.
Published: (2025)
by: Badki, Abhishek, et al.
Published: (2025)
Exploring Position Encoding in Diffusion U-Net for Training-free High-resolution Image Generation
by: Zhou, Feng, et al.
Published: (2025)
by: Zhou, Feng, et al.
Published: (2025)
Factuality Matters: When Image Generation and Editing Meet Structured Visuals
by: Zhuo, Le, et al.
Published: (2025)
by: Zhuo, Le, et al.
Published: (2025)
Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models
by: Ding, Xin, et al.
Published: (2025)
by: Ding, Xin, et al.
Published: (2025)
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
by: Zhao, Min, et al.
Published: (2026)
by: Zhao, Min, et al.
Published: (2026)
MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
by: Xu, Wenbo, et al.
Published: (2026)
by: Xu, Wenbo, et al.
Published: (2026)
Seeing It Before It Happens: In-Generation NSFW Detection for Diffusion-Based Text-to-Image Models
by: Yang, Fan, et al.
Published: (2025)
by: Yang, Fan, et al.
Published: (2025)
FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder
by: Dong, Zeyu, et al.
Published: (2026)
by: Dong, Zeyu, et al.
Published: (2026)
UNIT: Unifying Image and Text Recognition in One Vision Encoder
by: Zhu, Yi, et al.
Published: (2024)
by: Zhu, Yi, et al.
Published: (2024)
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models
by: Liu, Lu, et al.
Published: (2026)
by: Liu, Lu, et al.
Published: (2026)
Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation
by: Ye, Chongjie, et al.
Published: (2026)
by: Ye, Chongjie, et al.
Published: (2026)
Similar Items
-
Learning A Low-Level Vision Generalist via Visual Task Prompt
by: Chen, Xiangyu, et al.
Published: (2024) -
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
by: Pu, Yuandong, et al.
Published: (2025) -
A Comparative Study of Image Restoration Networks for General Backbone Network Design
by: Chen, Xiangyu, et al.
Published: (2023) -
UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
by: Cao, Shuo, et al.
Published: (2025) -
Accelerating Masked Image Generation by Learning Latent Controlled Dynamics
by: Zhu, Kaiwen, et al.
Published: (2026)