Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
Fuente:
arXiv
Guardado en:
| Autores principales: | Wei, Zhixiang, Li, Yi, Kan, Zhehan, Jiang, Xinghua, Long, Zuwei, Liu, Shifeng, Shen, Hongze, Liu, Wei, Tan, Xiaoyu, Lin, Haojia, Zhu, Yubo, Li, Qianyu, Yin, Di, Cao, Haoyu, Gu, Weibo, Li, Xin, Liu, Yinsong, Jiang, Deqiang, Sun, Xing, Wu, Yunsheng, Tang, Mingkong, Liu, Shuangyin, Tang, Lexiang, Lin, Haodong, Lu, Junru, Qin, Jiarui, Qiao, Lingfeng, Qiao, Ruizhi, Ke, Bo, He, Jianfeng, Li, Ke, Li, Yangning, Shen, Yunhang, Zhang, Mengdan, Chen, Peixian, Yin, Kun, Liu, Bing, Wu, Yunfei, Chen, Huang, Cai, Zhongpeng, Li, Xiaotian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding
por: Yin, Kun, et al.
Publicado: (2026)
por: Yin, Kun, et al.
Publicado: (2026)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
por: Li, Yi, et al.
Publicado: (2026)
por: Li, Yi, et al.
Publicado: (2026)
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
por: Lu, Junru, et al.
Publicado: (2025)
por: Lu, Junru, et al.
Publicado: (2025)
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
por: Kan, Zhehan, et al.
Publicado: (2025)
por: Kan, Zhehan, et al.
Publicado: (2025)
VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
por: Long, Zuwei, et al.
Publicado: (2025)
por: Long, Zuwei, et al.
Publicado: (2025)
HRVDA: High-Resolution Visual Document Assistant
por: Liu, Chaohu, et al.
Publicado: (2024)
por: Liu, Chaohu, et al.
Publicado: (2024)
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
por: Li, Xudong, et al.
Publicado: (2025)
por: Li, Xudong, et al.
Publicado: (2025)
Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization
por: Shi, Yuchen, et al.
Publicado: (2025)
por: Shi, Yuchen, et al.
Publicado: (2025)
Probability-density-aware Semi-supervised Learning
por: Liu, Shuyang, et al.
Publicado: (2024)
por: Liu, Shuyang, et al.
Publicado: (2024)
DeepOmni: Towards Seamless and Smart Speech Interaction with Adaptive Modality-Specific MoE
por: Shao, Hang, et al.
Publicado: (2025)
por: Shao, Hang, et al.
Publicado: (2025)
Evaluation of the effectiveness of using prednisolone, tacrolimus, and intravenous immunoglobulin combination therapy on immune‐mediated necrotizing myopathy—A non‐randomized, observational research
por: Mengyang Liu, et al.
Publicado: (2024)
por: Mengyang Liu, et al.
Publicado: (2024)
Competition between surficial and volumetric diffusion in sintering TiO 2 polymorphs by molecular dynamics simulation
por: Jiang Li, et al.
Publicado: (2024)
por: Jiang Li, et al.
Publicado: (2024)
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
por: Fu, Chaoyou, et al.
Publicado: (2025)
por: Fu, Chaoyou, et al.
Publicado: (2025)
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
por: Fu, Chaoyou, et al.
Publicado: (2023)
por: Fu, Chaoyou, et al.
Publicado: (2023)
Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language Models
por: Li, Xin, et al.
Publicado: (2024)
por: Li, Xin, et al.
Publicado: (2024)
Adaptive Global and Fine-Grained Perceptual Fusion for MLLM Embeddings Compatible with Hard Negative Amplification
por: Hu, Lexiang, et al.
Publicado: (2026)
por: Hu, Lexiang, et al.
Publicado: (2026)
A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?
por: Liu, Xinyu, et al.
Publicado: (2024)
por: Liu, Xinyu, et al.
Publicado: (2024)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
por: Zhou, Chenyu, et al.
Publicado: (2024)
por: Zhou, Chenyu, et al.
Publicado: (2024)
LensWalk: Agentic Video Understanding by Planning How You See in Videos
por: Li, Keliang, et al.
Publicado: (2026)
por: Li, Keliang, et al.
Publicado: (2026)
Lightweight single-image super-resolution network based on dual paths
por: Ke, Li, et al.
Publicado: (2024)
por: Ke, Li, et al.
Publicado: (2024)
Targeted Nanoprobes Enabled Precision Theranostics in Triple‐Negative Breast Cancer
por: Ke Ma, et al.
Publicado: (2025)
por: Ke Ma, et al.
Publicado: (2025)
Progressive Domain Adaptation for Thermal Infrared Object Tracking
por: Li, Qiao, et al.
Publicado: (2024)
por: Li, Qiao, et al.
Publicado: (2024)
Symmetry Discovery for Different Data Types
por: Hu, Lexiang, et al.
Publicado: (2024)
por: Hu, Lexiang, et al.
Publicado: (2024)
Explicit Discovery of Nonlinear Symmetries from Dynamic Data
por: Hu, Lexiang, et al.
Publicado: (2025)
por: Hu, Lexiang, et al.
Publicado: (2025)
Governing Equation Discovery from Data Based on Differential Invariants
por: Hu, Lexiang, et al.
Publicado: (2025)
por: Hu, Lexiang, et al.
Publicado: (2025)
No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
por: Dai, Zunkai, et al.
Publicado: (2026)
por: Dai, Zunkai, et al.
Publicado: (2026)
A novel dual-stage algorithm for capacitated arc routing problems with time-dependent service costs
por: Li, Qingya, et al.
Publicado: (2024)
por: Li, Qingya, et al.
Publicado: (2024)
Heavy‐Textured Rhizosphere Soils Enhance Microbial Nitrogen Fixation in a Desert Shrub Ecosystem
por: Chenhua Li, et al.
Publicado: (2025)
por: Chenhua Li, et al.
Publicado: (2025)
DiffVL: Diffusion-Based Visual Localization on 2D Maps via BEV-Conditioned GPS Denoising
por: Gao, Li, et al.
Publicado: (2025)
por: Gao, Li, et al.
Publicado: (2025)
Summit Vitals: Multi-Camera and Multi-Signal Biosensing at High Altitudes
por: Liu, Ke, et al.
Publicado: (2024)
por: Liu, Ke, et al.
Publicado: (2024)
Dynamic Contrastive Knowledge Distillation for Efficient Image Restoration
por: Zhou, Yunshuai, et al.
Publicado: (2024)
por: Zhou, Yunshuai, et al.
Publicado: (2024)
Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language Models
por: Liu, Ziche, et al.
Publicado: (2024)
por: Liu, Ziche, et al.
Publicado: (2024)
Robust Adaptive Control for High‐Order Nonlinear Systems With Unknown Upper Bound Uncertainties Based on Fully Actuated System Approaches and Multi‐Objective Optimization
por: Da‐Ke Gu, et al.
Publicado: (2025)
por: Da‐Ke Gu, et al.
Publicado: (2025)
DREAM: Document Reconstruction via End-to-end Autoregressive Model
por: Li, Xin, et al.
Publicado: (2025)
por: Li, Xin, et al.
Publicado: (2025)
The effects of shock process and terrestrial weathering on mercury isotopes in meteorites
por: Yan Fan, et al.
Publicado: (2026)
por: Yan Fan, et al.
Publicado: (2026)
SG-Reg: Generalizable and Efficient Scene Graph Registration
por: Liu, Chuhao, et al.
Publicado: (2025)
por: Liu, Chuhao, et al.
Publicado: (2025)
Microstructure and Wear Behavior of In Situ Multi‐Ceramic‐Reinforced Al0.5CoCrFeNi Coating Tuned by Laser Energy Density
por: Xingyi Liu, et al.
Publicado: (2026)
por: Xingyi Liu, et al.
Publicado: (2026)
Contrastive Local Manifold Learning for No-Reference Image Quality Assessment
por: Huang, Zihao, et al.
Publicado: (2024)
por: Huang, Zihao, et al.
Publicado: (2024)
FM-Fusion: Instance-aware Semantic Mapping Boosted by Vision-Language Foundation Models
por: Liu, Chuhao, et al.
Publicado: (2024)
por: Liu, Chuhao, et al.
Publicado: (2024)
JW-VL: A Vision-Language Model for Solar Physics
por: Shao, Mingfu, et al.
Publicado: (2026)
por: Shao, Mingfu, et al.
Publicado: (2026)
Ejemplares similares
-
Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding
por: Yin, Kun, et al.
Publicado: (2026) -
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
por: Li, Yi, et al.
Publicado: (2026) -
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
por: Lu, Junru, et al.
Publicado: (2025) -
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
por: Kan, Zhehan, et al.
Publicado: (2025) -
VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
por: Long, Zuwei, et al.
Publicado: (2025)