Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Daxiang, Zheng, Mingming, Xu, Dong, Luo, Chunhua, Zhuang, Bairong, Li, Yuxuan, He, Ruoyun, Wang, Haoran, Zhang, Wenyu, Wang, Wenbo, Wang, Yicheng, Xiong, Xue, Zheng, Ayong, Zuo, Xiaoying, Ou, Ziwei, Gu, Jingnan, Guo, Quanhao, Wu, Jianmin, Yin, Dawei, Shen, Dou |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
by: Dong, Daxiang, et al.
Published: (2025)
by: Dong, Daxiang, et al.
Published: (2025)
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime
by: Zhu, Tianshu, et al.
Published: (2026)
by: Zhu, Tianshu, et al.
Published: (2026)
QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs
by: Li, Shupeng, et al.
Published: (2025)
by: Li, Shupeng, et al.
Published: (2025)
Tree-of-Code: A Tree-Structured Exploring Framework for End-to-End Code Generation and Execution in Complex Task Handling
by: Ni, Ziyi, et al.
Published: (2024)
by: Ni, Ziyi, et al.
Published: (2024)
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games
by: Zhang, Wenyu, et al.
Published: (2026)
by: Zhang, Wenyu, et al.
Published: (2026)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
by: Zhong, Yufeng, et al.
Published: (2026)
by: Zhong, Yufeng, et al.
Published: (2026)
Detached Skip-Links and $R$-Probe: Decoupling Feature Aggregation from Gradient Propagation for MLLM OCR
by: Yuan, Ziye, et al.
Published: (2026)
by: Yuan, Ziye, et al.
Published: (2026)
The End of Manual Decoding: Towards Truly End-to-End Language Models
by: Wang, Zhichao, et al.
Published: (2025)
by: Wang, Zhichao, et al.
Published: (2025)
Brightness of the Qianfan Satellites
by: Mallama, Anthony, et al.
Published: (2024)
by: Mallama, Anthony, et al.
Published: (2024)
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
by: You, Zhenghai, et al.
Published: (2025)
by: You, Zhenghai, et al.
Published: (2025)
Automated Data Quality Validation in an End-to-End GNN Framework
by: Dong, Sijie, et al.
Published: (2025)
by: Dong, Sijie, et al.
Published: (2025)
ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving
by: Zheng, Zhiyu, et al.
Published: (2025)
by: Zheng, Zhiyu, et al.
Published: (2025)
LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR
by: Taghadouini, Said, et al.
Published: (2026)
by: Taghadouini, Said, et al.
Published: (2026)
P$^{3}$Nav: End-to-End Perception, Prediction and Planning for Vision-and-Language Navigation
by: Li, Tianfu, et al.
Published: (2026)
by: Li, Tianfu, et al.
Published: (2026)
Improving End-to-End Training of Retrieval-Augmented Generation Models via Joint Stochastic Approximation
by: Cao, Hongyu, et al.
Published: (2025)
by: Cao, Hongyu, et al.
Published: (2025)
Countering Mainstream Bias via End-to-End Adaptive Local Learning
by: Pan, Jinhao, et al.
Published: (2024)
by: Pan, Jinhao, et al.
Published: (2024)
Embodied Cognition Augmented End2End Autonomous Driving
by: Niu, Ling, et al.
Published: (2025)
by: Niu, Ling, et al.
Published: (2025)
ContourFormer: Real-Time Contour-Based End-to-End Instance Segmentation Transformer
by: Yao, Weiwei, et al.
Published: (2025)
by: Yao, Weiwei, et al.
Published: (2025)
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
by: Sun, Wenchao, et al.
Published: (2024)
by: Sun, Wenchao, et al.
Published: (2024)
Morphological and molecular identification of Asterostroma roseoalbum sp. nov. (Peniophoraceae, Russulales), from southwestern China
by: Junhong Dong, et al.
Published: (2024)
by: Junhong Dong, et al.
Published: (2024)
PlatformX: An End-to-End Transferable Platform for Energy-Efficient Neural Architecture Search
by: Tu, Xiaolong, et al.
Published: (2025)
by: Tu, Xiaolong, et al.
Published: (2025)
End4: End-to-end Denoising Diffusion for Diffusion-Based Inpainting Detection
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
by: Lin, Wan, et al.
Published: (2024)
by: Lin, Wan, et al.
Published: (2024)
Stratos: An End-to-End Distillation Pipeline for Customized LLMs under Distributed Cloud Environments
by: Dai, Ziming, et al.
Published: (2025)
by: Dai, Ziming, et al.
Published: (2025)
EGA-V1: Unifying Online Advertising with End-to-End Learning
by: Qiu, Junyan, et al.
Published: (2025)
by: Qiu, Junyan, et al.
Published: (2025)
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning
by: Zhao, Haotian, et al.
Published: (2026)
by: Zhao, Haotian, et al.
Published: (2026)
LASER: An Efficient Target-Aware Segmented Attention Framework for End-to-End Long Sequence Modeling
by: Lin, Tianhe, et al.
Published: (2026)
by: Lin, Tianhe, et al.
Published: (2026)
PanguMotion: Continuous Driving Motion Forecasting with Pangu Transformers
by: Ren, Quanhao, et al.
Published: (2026)
by: Ren, Quanhao, et al.
Published: (2026)
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
by: Guo, Yinlin, et al.
Published: (2024)
by: Guo, Yinlin, et al.
Published: (2024)
End-to-End HOI Reconstruction Transformer with Graph-based Encoding
by: Wang, Zhenrong, et al.
Published: (2025)
by: Wang, Zhenrong, et al.
Published: (2025)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
by: Zhang, Yaolun, et al.
Published: (2026)
by: Zhang, Yaolun, et al.
Published: (2026)
Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models
by: Liu, Wenhan, et al.
Published: (2024)
by: Liu, Wenhan, et al.
Published: (2024)
PanopticSplatting: End-to-End Panoptic Gaussian Splatting
by: Xie, Yuxuan, et al.
Published: (2025)
by: Xie, Yuxuan, et al.
Published: (2025)
End-to-End Training for Back-Translation with Categorical Reparameterization Trick
by: Heo, DongNyeong, et al.
Published: (2022)
by: Heo, DongNyeong, et al.
Published: (2022)
SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
by: Li, Tianpeng, et al.
Published: (2025)
by: Li, Tianpeng, et al.
Published: (2025)
End-to-End Multi-Modal Diffusion Mamba
by: Lu, Chunhao, et al.
Published: (2025)
by: Lu, Chunhao, et al.
Published: (2025)
Similar Items
-
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
by: Dong, Daxiang, et al.
Published: (2025) -
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime
by: Zhu, Tianshu, et al.
Published: (2026) -
QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs
by: Li, Shupeng, et al.
Published: (2025) -
Tree-of-Code: A Tree-Structured Exploring Framework for End-to-End Code Generation and Execution in Complex Task Handling
by: Ni, Ziyi, et al.
Published: (2024) -
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games
by: Zhang, Wenyu, et al.
Published: (2026)