OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yanqing, Li, Xianhang, Zhang, Letian, Wang, Zirui, Zheng, Zeyu, Zhou, Yuyin, Xie, Cihang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation
von: Zhang, Letian, et al.
Veröffentlicht: (2026)
von: Zhang, Letian, et al.
Veröffentlicht: (2026)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
von: Liu, Yanqing, et al.
Veröffentlicht: (2024)
von: Liu, Yanqing, et al.
Veröffentlicht: (2024)
Scaling White-Box Transformers for Vision
von: Yang, Jinrui, et al.
Veröffentlicht: (2024)
von: Yang, Jinrui, et al.
Veröffentlicht: (2024)
Revisiting Adversarial Training at Scale
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
von: Yang, Jinrui, et al.
Veröffentlicht: (2026)
von: Yang, Jinrui, et al.
Veröffentlicht: (2026)
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
von: Mao, Jiawei, et al.
Veröffentlicht: (2024)
von: Mao, Jiawei, et al.
Veröffentlicht: (2024)
3D-TransUNet for Brain Metastases Segmentation in the BraTS2023 Challenge
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
von: Gao, Yipeng, et al.
Veröffentlicht: (2023)
von: Gao, Yipeng, et al.
Veröffentlicht: (2023)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
L2B: Learning to Bootstrap Robust Models for Combating Label Noise
von: Zhou, Yuyin, et al.
Veröffentlicht: (2022)
von: Zhou, Yuyin, et al.
Veröffentlicht: (2022)
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
von: Mao, Jiawei, et al.
Veröffentlicht: (2026)
von: Mao, Jiawei, et al.
Veröffentlicht: (2026)
MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
Autoregressive Pretraining with Mamba in Vision
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
von: Qian, Zhaofang, et al.
Veröffentlicht: (2025)
von: Qian, Zhaofang, et al.
Veröffentlicht: (2025)
What If We Recaption Billions of Web Images with LLaMA-3?
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
CAST: Modeling Visual State Transitions for Consistent Video Retrieval
von: Liu, Yanqing, et al.
Veröffentlicht: (2026)
von: Liu, Yanqing, et al.
Veröffentlicht: (2026)
Generative Image Layer Decomposition with Visual Effects
von: Yang, Jinrui, et al.
Veröffentlicht: (2024)
von: Yang, Jinrui, et al.
Veröffentlicht: (2024)
ARFlow: Autoregressive Flow with Hybrid Linear Attention
von: Hui, Mude, et al.
Veröffentlicht: (2025)
von: Hui, Mude, et al.
Veröffentlicht: (2025)
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
von: Yang, Siwei, et al.
Veröffentlicht: (2025)
von: Yang, Siwei, et al.
Veröffentlicht: (2025)
Mamba-R: Vision Mamba ALSO Needs Registers
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
Efficient Reinforcement Learning Through Adaptively Pretrained Visual Encoder
von: Zhang, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhan, et al.
Veröffentlicht: (2025)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
Rejuvenating image-GPT as Strong Visual Representation Learners
von: Ren, Sucheng, et al.
Veröffentlicht: (2023)
von: Ren, Sucheng, et al.
Veröffentlicht: (2023)
Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding
von: Xie, Jiangnan, et al.
Veröffentlicht: (2025)
von: Xie, Jiangnan, et al.
Veröffentlicht: (2025)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
von: Li, Zhuowan, et al.
Veröffentlicht: (2022)
von: Li, Zhuowan, et al.
Veröffentlicht: (2022)
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
von: Liu, Zhiheng, et al.
Veröffentlicht: (2026)
von: Liu, Zhiheng, et al.
Veröffentlicht: (2026)
On the Adversarial Robustness of Camera-based 3D Object Detection
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2023)
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2023)
One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
von: Peng, Cihang, et al.
Veröffentlicht: (2025)
von: Peng, Cihang, et al.
Veröffentlicht: (2025)
VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
von: Mei, Hefei, et al.
Veröffentlicht: (2025)
von: Mei, Hefei, et al.
Veröffentlicht: (2025)
BackdoorIDS: Zero-shot Backdoor Detection for Pretrained Vision Encoder
von: Huang, Siquan, et al.
Veröffentlicht: (2026)
von: Huang, Siquan, et al.
Veröffentlicht: (2026)
Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies
von: Qian, Jianing, et al.
Veröffentlicht: (2024)
von: Qian, Jianing, et al.
Veröffentlicht: (2024)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
von: Liu, Haoyu, et al.
Veröffentlicht: (2026)
von: Liu, Haoyu, et al.
Veröffentlicht: (2026)
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
von: Li, Xinpeng, et al.
Veröffentlicht: (2026)
von: Li, Xinpeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
von: Li, Xianhang, et al.
Veröffentlicht: (2025) -
OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation
von: Zhang, Letian, et al.
Veröffentlicht: (2026) -
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
von: Liu, Yanqing, et al.
Veröffentlicht: (2024) -
Scaling White-Box Transformers for Vision
von: Yang, Jinrui, et al.
Veröffentlicht: (2024) -
Revisiting Adversarial Training at Scale
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)