Ovis-U1 Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Guo-Hua, Zhao, Shanshan, Zhang, Xinjie, Cao, Liangfu, Zhan, Pengxin, Duan, Lunhao, Lu, Shiyin, Fu, Minghao, Chen, Xiaohao, Zhao, Jianshan, Li, Yang, Chen, Qing-Guo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ovis-Image Technical Report
by: Wang, Guo-Hua, et al.
Published: (2025)
by: Wang, Guo-Hua, et al.
Published: (2025)
Ovis2.5 Technical Report
by: Lu, Shiyin, et al.
Published: (2025)
by: Lu, Shiyin, et al.
Published: (2025)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
by: Lu, Shiyin, et al.
Published: (2024)
by: Lu, Shiyin, et al.
Published: (2024)
TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
Local-consistent Transformation Learning for Rotation-invariant Point Cloud Analysis
by: Chen, Yiyang, et al.
Published: (2024)
by: Chen, Yiyang, et al.
Published: (2024)
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning
by: Chen, Yiyang, et al.
Published: (2025)
by: Chen, Yiyang, et al.
Published: (2025)
High-quality Pseudo-labeling for Point Cloud Segmentation with Scene-level Annotation
by: Duan, Lunhao, et al.
Published: (2025)
by: Duan, Lunhao, et al.
Published: (2025)
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
by: Duan, Lunhao, et al.
Published: (2024)
by: Duan, Lunhao, et al.
Published: (2024)
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
by: Gao, Sensen, et al.
Published: (2025)
by: Gao, Sensen, et al.
Published: (2025)
Neural Stereo Video Compression with Hybrid Disparity Compensation
by: Jiang, Shiyin, et al.
Published: (2025)
by: Jiang, Shiyin, et al.
Published: (2025)
MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
by: Sun, Hui, et al.
Published: (2025)
by: Sun, Hui, et al.
Published: (2025)
MASSeg : 2nd Technical Report for 4th PVUW MOSE Track
by: Cao, Xuqiang, et al.
Published: (2025)
by: Cao, Xuqiang, et al.
Published: (2025)
Technical Report for ActivityNet Challenge 2022 -- Temporal Action Localization
by: Chen, Shimin, et al.
Published: (2024)
by: Chen, Shimin, et al.
Published: (2024)
Technical Report for SoccerNet Challenge 2022 -- Replay Grounding Task
by: Chen, Shimin, et al.
Published: (2024)
by: Chen, Shimin, et al.
Published: (2024)
Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
by: Yang, Wenhao, et al.
Published: (2025)
by: Yang, Wenhao, et al.
Published: (2025)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
by: Zhang, Yi-Kai, et al.
Published: (2024)
by: Zhang, Yi-Kai, et al.
Published: (2024)
Singpath-VL Technical Report
by: Qiu, Zhen, et al.
Published: (2026)
by: Qiu, Zhen, et al.
Published: (2026)
Controlling the Latent Diffusion Model for Generative Image Shadow Removal via Residual Generation
by: Li, Xinjie, et al.
Published: (2024)
by: Li, Xinjie, et al.
Published: (2024)
Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image Compression
by: Jiang, Shiyin, et al.
Published: (2026)
by: Jiang, Shiyin, et al.
Published: (2026)
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Kwai Keye-VL Technical Report
by: Kwai Keye Team, et al.
Published: (2025)
by: Kwai Keye Team, et al.
Published: (2025)
Mano Technical Report
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
AnchorRoute: Human Motion Synthesis with Interval-Routed Sparse Contro
by: Fang, Pengcheng, et al.
Published: (2026)
by: Fang, Pengcheng, et al.
Published: (2026)
PEMF-VTO: Point-Enhanced Video Virtual Try-on via Mask-free Paradigm
by: Chang, Tianyu, et al.
Published: (2024)
by: Chang, Tianyu, et al.
Published: (2024)
Better Fit: Accommodate Variations in Clothing Types for Virtual Try-on
by: Song, Dan, et al.
Published: (2024)
by: Song, Dan, et al.
Published: (2024)
iMedImage Technical Report
by: Wei, Ran, et al.
Published: (2025)
by: Wei, Ran, et al.
Published: (2025)
Logics-Parsing Technical Report
by: Chen, Xiangyang, et al.
Published: (2025)
by: Chen, Xiangyang, et al.
Published: (2025)
HunyuanVideo 1.5 Technical Report
by: Wu, Bing, et al.
Published: (2025)
by: Wu, Bing, et al.
Published: (2025)
Parrot: Multilingual Visual Instruction Tuning
by: Sun, Hai-Long, et al.
Published: (2024)
by: Sun, Hai-Long, et al.
Published: (2024)
Technical Report for Soccernet 2023 -- Dense Video Captioning
by: Ruan, Zheng, et al.
Published: (2024)
by: Ruan, Zheng, et al.
Published: (2024)
SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing
by: Ge, Yuying, et al.
Published: (2024)
by: Ge, Yuying, et al.
Published: (2024)
ERNIE-Image Technical Report
by: Liu, Jiaxiang, et al.
Published: (2026)
by: Liu, Jiaxiang, et al.
Published: (2026)
UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars
by: Zhan, Xiaoyu, et al.
Published: (2026)
by: Zhan, Xiaoyu, et al.
Published: (2026)
StreamingClaw Technical Report
by: Chen, Jiawei, et al.
Published: (2026)
by: Chen, Jiawei, et al.
Published: (2026)
FireRed-Image-Edit-1.0 Technical Report
by: Super Intelligence Team, et al.
Published: (2026)
by: Super Intelligence Team, et al.
Published: (2026)
Qwen-Image Technical Report
by: Wu, Chenfei, et al.
Published: (2025)
by: Wu, Chenfei, et al.
Published: (2025)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Similar Items
-
Ovis-Image Technical Report
by: Wang, Guo-Hua, et al.
Published: (2025) -
Ovis2.5 Technical Report
by: Lu, Shiyin, et al.
Published: (2025) -
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025) -
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
by: Fu, Minghao, et al.
Published: (2025) -
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
by: Lu, Shiyin, et al.
Published: (2024)