OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Letian, Ren, Sucheng, Liu, Yanqing, Li, Xianhang, Wang, Zeyu, Zhou, Yuyin, Yao, Huaxiu, Zheng, Zeyu, Nie, Weili, Liu, Guilin, Yu, Zhiding, Xie, Cihang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
by: Liu, Yanqing, et al.
Published: (2025)
by: Liu, Yanqing, et al.
Published: (2025)
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
by: Li, Xianhang, et al.
Published: (2025)
by: Li, Xianhang, et al.
Published: (2025)
3D-TransUNet for Brain Metastases Segmentation in the BraTS2023 Challenge
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)
by: Liu, Yanqing, et al.
Published: (2024)
Rejuvenating image-GPT as Strong Visual Representation Learners
by: Ren, Sucheng, et al.
Published: (2023)
by: Ren, Sucheng, et al.
Published: (2023)
Revisiting Adversarial Training at Scale
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
What If We Recaption Billions of Web Images with LLaMA-3?
by: Li, Xianhang, et al.
Published: (2024)
by: Li, Xianhang, et al.
Published: (2024)
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
by: Yang, Jinrui, et al.
Published: (2026)
by: Yang, Jinrui, et al.
Published: (2026)
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
by: Mao, Jiawei, et al.
Published: (2026)
by: Mao, Jiawei, et al.
Published: (2026)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
by: Wang, Zijun, et al.
Published: (2026)
by: Wang, Zijun, et al.
Published: (2026)
UnPuzzle: A Unified Framework for Pathology Image Analysis
by: Liao, Dankai, et al.
Published: (2025)
by: Liao, Dankai, et al.
Published: (2025)
Scaling White-Box Transformers for Vision
by: Yang, Jinrui, et al.
Published: (2024)
by: Yang, Jinrui, et al.
Published: (2024)
I$^2$VC: A Unified Framework for Intra- & Inter-frame Video Compression
by: Liu, Meiqin, et al.
Published: (2024)
by: Liu, Meiqin, et al.
Published: (2024)
FDG-Diff: Frequency-Domain-Guided Diffusion Framework for Compressed Hazy Image Restoration
by: Zhang, Ruicheng, et al.
Published: (2025)
by: Zhang, Ruicheng, et al.
Published: (2025)
Mamba-based Light Field Super-Resolution with Efficient Subspace Scanning
by: Gao, Ruisheng, et al.
Published: (2024)
by: Gao, Ruisheng, et al.
Published: (2024)
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
by: Liu, Jiaqi, et al.
Published: (2026)
by: Liu, Jiaqi, et al.
Published: (2026)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
by: Liu, Haoyu, et al.
Published: (2026)
by: Liu, Haoyu, et al.
Published: (2026)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
Mamba-R: Vision Mamba ALSO Needs Registers
by: Wang, Feng, et al.
Published: (2024)
by: Wang, Feng, et al.
Published: (2024)
Optimal Transport Based Unsupervised Restoration Learning Exploiting Degradation Sparsity
by: Wen, Fei, et al.
Published: (2023)
by: Wen, Fei, et al.
Published: (2023)
Unifying Image Processing as Visual Prompting Question Answering
by: Liu, Yihao, et al.
Published: (2023)
by: Liu, Yihao, et al.
Published: (2023)
FairEnc: A Fair Vision-Language Model with Fair Vision and Text Encoders for Glaucoma Detection
by: Elhabebe, Mohamed, et al.
Published: (2026)
by: Elhabebe, Mohamed, et al.
Published: (2026)
Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding
by: Ma, Chenglong, et al.
Published: (2025)
by: Ma, Chenglong, et al.
Published: (2025)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
by: Gao, Yipeng, et al.
Published: (2023)
by: Gao, Yipeng, et al.
Published: (2023)
Enhancing Super-Resolution Networks through Realistic Thick-Slice CT Simulation
by: Tang, Zeyu, et al.
Published: (2023)
by: Tang, Zeyu, et al.
Published: (2023)
Generating Progressive Images from Pathological Transitions via Diffusion Model
by: Liu, Zeyu, et al.
Published: (2023)
by: Liu, Zeyu, et al.
Published: (2023)
AutoEncoder Convolutional Neural Network for Pneumonia Detection
by: Nosa-Omoruyi, Michael, et al.
Published: (2024)
by: Nosa-Omoruyi, Michael, et al.
Published: (2024)
Encoder-Quantization-Motion-based Video Quality Metrics
by: Chen, Yixu, et al.
Published: (2024)
by: Chen, Yixu, et al.
Published: (2024)
RAVEN: Radar Adaptive Vision Encoders for Efficient Chirp-wise Object Detection and Segmentation
by: Sen, Anuvab, et al.
Published: (2026)
by: Sen, Anuvab, et al.
Published: (2026)
PathRWKV: Enhancing Whole Slide Image Inference with Asymmetric Recurrent Modeling
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
UP-Diff: Latent Diffusion Model for Remote Sensing Urban Prediction
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
Modeling the HEVC Encoding Energy Using the Encoder Processing Time
by: Ramasubbu, Geetha, et al.
Published: (2022)
by: Ramasubbu, Geetha, et al.
Published: (2022)
Hibou: A Family of Foundational Vision Transformers for Pathology
by: Nechaev, Dmitry, et al.
Published: (2024)
by: Nechaev, Dmitry, et al.
Published: (2024)
Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency
by: Wang, Feng, et al.
Published: (2024)
by: Wang, Feng, et al.
Published: (2024)
You Only Train Once: A Unified Framework for Both Full-Reference and No-Reference Image Quality Assessment
by: Yun, Yi Ke, et al.
Published: (2023)
by: Yun, Yi Ke, et al.
Published: (2023)
FRAPPE: Full Input, Residual Output Autoencoding with Projection Pursuit Encoder
by: Jacobellis, Dan, et al.
Published: (2026)
by: Jacobellis, Dan, et al.
Published: (2026)
A CT Image Denoising Method with Residual Encoder-Decoder Network
by: Shawn, Helena, et al.
Published: (2024)
by: Shawn, Helena, et al.
Published: (2024)
SimpleMem: Efficient Lifelong Memory for LLM Agents
by: Liu, Jiaqi, et al.
Published: (2026)
by: Liu, Jiaqi, et al.
Published: (2026)
Make Both Ends Meet: A Synergistic Optimization Infrared Small Target Detection with Streamlined Computational Overhead
by: Jing, Yuxin, et al.
Published: (2025)
by: Jing, Yuxin, et al.
Published: (2025)
Similar Items
-
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
by: Liu, Yanqing, et al.
Published: (2025) -
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
by: Li, Xianhang, et al.
Published: (2025) -
3D-TransUNet for Brain Metastases Segmentation in the BraTS2023 Challenge
by: Yang, Siwei, et al.
Published: (2024) -
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
by: Ren, Sucheng, et al.
Published: (2024) -
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)