Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Kun, Wu, Yunfei, Liu, Bing, Cai, Zhongpeng, Li, Xiaotian, Chen, Huang, Li, Xin, Cao, Haoyu, Liu, Yinsong, Jiang, Deqiang, Sun, Xing, Wu, Yunsheng, Li, Qianyu, Guo, Antai, Liao, Yanzhen, Qu, Yanqiu, Lin, Haodong, He, Chengxu, Liu, Shuangyin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
by: Wei, Zhixiang, et al.
Published: (2026)
by: Wei, Zhixiang, et al.
Published: (2026)
DREAM: Document Reconstruction via End-to-end Autoregressive Model
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
by: Lu, Junru, et al.
Published: (2025)
by: Lu, Junru, et al.
Published: (2025)
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
by: Wang, Yubo, et al.
Published: (2024)
by: Wang, Yubo, et al.
Published: (2024)
Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language Models
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization
by: Shi, Yuchen, et al.
Published: (2025)
by: Shi, Yuchen, et al.
Published: (2025)
Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning
by: Dong, Junnan, et al.
Published: (2025)
by: Dong, Junnan, et al.
Published: (2025)
HRVDA: High-Resolution Visual Document Assistant
by: Liu, Chaohu, et al.
Published: (2024)
by: Liu, Chaohu, et al.
Published: (2024)
Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and Reaction
by: Yan, Haoqiu, et al.
Published: (2024)
by: Yan, Haoqiu, et al.
Published: (2024)
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
by: Liu, Wuyang, et al.
Published: (2023)
by: Liu, Wuyang, et al.
Published: (2023)
Multi-directional Safe Rectangle Corridor-Based MPC for Nonholonomic Robots Navigation in Cluttered Environment
by: Qu, Yinsong, et al.
Published: (2025)
by: Qu, Yinsong, et al.
Published: (2025)
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
by: Kan, Zhehan, et al.
Published: (2025)
by: Kan, Zhehan, et al.
Published: (2025)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
by: Li, Yi, et al.
Published: (2026)
by: Li, Yi, et al.
Published: (2026)
PSGait: Gait Recognition using Parsing Skeleton
by: Xu, Hangrui, et al.
Published: (2025)
by: Xu, Hangrui, et al.
Published: (2025)
Massive MIMO CSI Feedback with Spiking Neural Networks
by: Liu, Yanzhen, et al.
Published: (2026)
by: Liu, Yanzhen, et al.
Published: (2026)
LncRNA FAM30A Suppresses Proliferation and Metastasis of Colorectal Carcinoma by Blocking the JAK–STAT Signalling
by: Jin Liu, et al.
Published: (2025)
by: Jin Liu, et al.
Published: (2025)
Efficient Document Parsing via Parallel Token Prediction
by: Li, Lei, et al.
Published: (2026)
by: Li, Lei, et al.
Published: (2026)
Explore Human Parsing Modality for Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
Transformer-based EEG Decoding: A Survey
by: Zhang, Haodong, et al.
Published: (2025)
by: Zhang, Haodong, et al.
Published: (2025)
Latent Video Dataset Distillation
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU Utilization
by: Wu, Yize, et al.
Published: (2025)
by: Wu, Yize, et al.
Published: (2025)
Autonomous vehicle decision and control through reinforcement learning with traffic flow randomization
by: Lin, Yuan, et al.
Published: (2024)
by: Lin, Yuan, et al.
Published: (2024)
PEARL: Parallel Speculative Decoding with Adaptive Draft Length
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
by: An, Zihao, et al.
Published: (2026)
by: An, Zihao, et al.
Published: (2026)
DASSF: Dynamic-Attention Scale-Sequence Fusion for Aerial Object Detection
by: Li, Haodong, et al.
Published: (2024)
by: Li, Haodong, et al.
Published: (2024)
QUDSELECT: Selective Decoding for Questions Under Discussion Parsing
by: Suvarna, Ashima, et al.
Published: (2024)
by: Suvarna, Ashima, et al.
Published: (2024)
ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
by: Li, Jia-Nan, et al.
Published: (2025)
by: Li, Jia-Nan, et al.
Published: (2025)
Perception and Recognition of Proximity and Contact Processes via Parallel Double‐Transistor Configuration
by: Shixin Liu, et al.
Published: (2025)
by: Shixin Liu, et al.
Published: (2025)
PMSN: A Parallel Multi-compartment Spiking Neuron for Multi-scale Temporal Processing
by: Chen, Xinyi, et al.
Published: (2024)
by: Chen, Xinyi, et al.
Published: (2024)
Cerberus: Efficient Inference with Adaptive Parallel Decoding and Sequential Knowledge Enhancement
by: Liu, Yuxuan, et al.
Published: (2024)
by: Liu, Yuxuan, et al.
Published: (2024)
S$^{2}$-DMs:Skip-Step Diffusion Models
by: Wang, Yixuan, et al.
Published: (2024)
by: Wang, Yixuan, et al.
Published: (2024)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
by: Li, Zhang, et al.
Published: (2025)
by: Li, Zhang, et al.
Published: (2025)
Parallel Decoding via Hidden Transfer for Lossless Large Language Model Acceleration
by: Wu, Pengfei, et al.
Published: (2024)
by: Wu, Pengfei, et al.
Published: (2024)
How Much Parallelism Is "Free"? A Principle of Near-Free Parallelism for Parallel Decoding
by: He, Minghua, et al.
Published: (2026)
by: He, Minghua, et al.
Published: (2026)
Generation Meets Verification: Accelerating Large Language Model Inference with Smart Parallel Auto-Correct Decoding
by: Yi, Hanling, et al.
Published: (2024)
by: Yi, Hanling, et al.
Published: (2024)
Design of an all-facet illuminator for high NA EUV lithography exposure tool based on deep reinforcement learning
by: Li, Tong, et al.
Published: (2025)
by: Li, Tong, et al.
Published: (2025)
Logics-Parsing Technical Report
by: Chen, Xiangyang, et al.
Published: (2025)
by: Chen, Xiangyang, et al.
Published: (2025)
Large Model Enabled Embodied Intelligence for 6G Integrated Perception, Communication, and Computation Network
by: Li, Zhuoran, et al.
Published: (2025)
by: Li, Zhuoran, et al.
Published: (2025)
Cross-Validated Cross-Channel Self-Attention and Denoising for Automatic Modulation Classification
by: Suman, Prakash, et al.
Published: (2026)
by: Suman, Prakash, et al.
Published: (2026)
A Lightweight Deep Learning Model for Automatic Modulation Classification using Dual Path Deep Residual Shrinkage Network
by: Suman, Prakash, et al.
Published: (2025)
by: Suman, Prakash, et al.
Published: (2025)
Similar Items
-
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
by: Wei, Zhixiang, et al.
Published: (2026) -
DREAM: Document Reconstruction via End-to-end Autoregressive Model
by: Li, Xin, et al.
Published: (2025) -
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
by: Lu, Junru, et al.
Published: (2025) -
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
by: Wang, Yubo, et al.
Published: (2024) -
Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language Models
by: Li, Xin, et al.
Published: (2024)