PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Xudong, Yan, Hao, Yin, Liang, Liu, Yang, Ding, Jing, Liao, Minghui, Liu, Yuliang, Chen, Wei, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
by: Luo, Dongliang, et al.
Published: (2025)
by: Luo, Dongliang, et al.
Published: (2025)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
by: Yin, Liang, et al.
Published: (2025)
by: Yin, Liang, et al.
Published: (2025)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
by: Liang, Jianxin, et al.
Published: (2024)
by: Liang, Jianxin, et al.
Published: (2024)
SparseAD: Sparse Query-Centric Paradigm for Efficient End-to-End Autonomous Driving
by: Zhang, Diankun, et al.
Published: (2024)
by: Zhang, Diankun, et al.
Published: (2024)
EMMA: End-to-End Multimodal Model for Autonomous Driving
by: Hwang, Jyh-Jing, et al.
Published: (2024)
by: Hwang, Jyh-Jing, et al.
Published: (2024)
End-to-End Graph Flattening Method for Large Language Models
by: Hong, Bin, et al.
Published: (2024)
by: Hong, Bin, et al.
Published: (2024)
Partial Scene Text Retrieval
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
End-to-End Simultaneous Dysarthric Speech Reconstruction with Frame-Level Adaptor and Multiple Wait-k Knowledge Distillation
by: Wu, Minghui, et al.
Published: (2026)
by: Wu, Minghui, et al.
Published: (2026)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
by: Sun, Mo, et al.
Published: (2024)
by: Sun, Mo, et al.
Published: (2024)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
by: Zhu, Jiaying, et al.
Published: (2025)
by: Zhu, Jiaying, et al.
Published: (2025)
End-to-End Long Document Summarization using Gradient Caching
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
Towards End-to-End Open Conversational Machine Reading
by: Zhou, Sizhe, et al.
Published: (2022)
by: Zhou, Sizhe, et al.
Published: (2022)
REMM:Rotation-Equivariant Framework for End-to-End Multimodal Image Matching
by: Nie, Han, et al.
Published: (2024)
by: Nie, Han, et al.
Published: (2024)
WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
by: Xu, Runsheng, et al.
Published: (2025)
by: Xu, Runsheng, et al.
Published: (2025)
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
by: Wang, Linhan, et al.
Published: (2026)
by: Wang, Linhan, et al.
Published: (2026)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
by: Baek, Jeonghun, et al.
Published: (2025)
by: Baek, Jeonghun, et al.
Published: (2025)
Leverage Cross-Attention for End-to-End Open-Vocabulary Panoptic Reconstruction
by: Yu, Xuan, et al.
Published: (2025)
by: Yu, Xuan, et al.
Published: (2025)
Progressive Evolution from Single-Point to Polygon for Scene Text
by: Deng, Linger, et al.
Published: (2023)
by: Deng, Linger, et al.
Published: (2023)
SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
by: Wang, Ruoyu, et al.
Published: (2026)
by: Wang, Ruoyu, et al.
Published: (2026)
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
by: You, Ling, et al.
Published: (2025)
by: You, Ling, et al.
Published: (2025)
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
by: Cheng, Dabing, et al.
Published: (2025)
by: Cheng, Dabing, et al.
Published: (2025)
VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
by: Yan, Hao, et al.
Published: (2025)
by: Yan, Hao, et al.
Published: (2025)
SDformer: Efficient End-to-End Transformer for Depth Completion
by: Qian, Jian, et al.
Published: (2024)
by: Qian, Jian, et al.
Published: (2024)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
by: Zhang, Yaolun, et al.
Published: (2026)
by: Zhang, Yaolun, et al.
Published: (2026)
SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
by: Sun, Wenchao, et al.
Published: (2024)
by: Sun, Wenchao, et al.
Published: (2024)
End-to-End Visual Autonomous Parking via Control-Aided Attention
by: Chen, Chao, et al.
Published: (2025)
by: Chen, Chao, et al.
Published: (2025)
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
An End-to-End Approach for Child Reading Assessment in the Xhosa Language
by: Chevtchenko, Sergio, et al.
Published: (2025)
by: Chevtchenko, Sergio, et al.
Published: (2025)
PDF: Point Diffusion Implicit Function for Large-scale Scene Neural Representation
by: Ding, Yuhan, et al.
Published: (2023)
by: Ding, Yuhan, et al.
Published: (2023)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving
by: Sun, Wenchao, et al.
Published: (2026)
by: Sun, Wenchao, et al.
Published: (2026)
Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
End-Cloud Collaboration Framework for Advanced AI Customer Service in E-commerce
by: Teng, Liangyu, et al.
Published: (2024)
by: Teng, Liangyu, et al.
Published: (2024)
Lightweight and Production-Ready PDF Visual Element Parsing
by: Liu, Meizhu, et al.
Published: (2026)
by: Liu, Meizhu, et al.
Published: (2026)
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
EndPrompt: Efficient Long-Context Extension via Terminal Anchoring
by: Tian, Han, et al.
Published: (2026)
by: Tian, Han, et al.
Published: (2026)
VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
Similar Items
-
SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
by: Luo, Dongliang, et al.
Published: (2025) -
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024) -
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
by: Yin, Liang, et al.
Published: (2025) -
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024) -
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
by: Liang, Jianxin, et al.
Published: (2024)