A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Yuxuan, Zhang, Yuanxing, Wang, Yushuo, Jin, Yichao, Ke, Kenneth Zhu, Zhao, Jingyuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models
by: Jin, Yichao, et al.
Published: (2025)
by: Jin, Yichao, et al.
Published: (2025)
Financial Table Extraction in Image Documents
by: Watson, William, et al.
Published: (2024)
by: Watson, William, et al.
Published: (2024)
Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
by: Jain, Chayan, et al.
Published: (2025)
by: Jain, Chayan, et al.
Published: (2025)
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
SynJAC: Synthetic-data-driven Joint-granular Adaptation and Calibration for Domain Specific Scanned Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
by: Chen, Jian, et al.
Published: (2025)
by: Chen, Jian, et al.
Published: (2025)
Handheld Video Document Scanning: A Robust On-Device Model for Multi-Page Document Scanning
by: Wigington, Curtis
Published: (2024)
by: Wigington, Curtis
Published: (2024)
Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers
by: Li, Bozhou, et al.
Published: (2026)
by: Li, Bozhou, et al.
Published: (2026)
HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios
by: Wang, Daming, et al.
Published: (2025)
by: Wang, Daming, et al.
Published: (2025)
Colour Extraction Pipeline for Odonates using Computer Vision
by: Rajaraman, Megan Mirnalini Sundaram, et al.
Published: (2026)
by: Rajaraman, Megan Mirnalini Sundaram, et al.
Published: (2026)
IPAD: Industrial Process Anomaly Detection Dataset
by: Liu, Jinfan, et al.
Published: (2024)
by: Liu, Jinfan, et al.
Published: (2024)
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
by: Lim, Ho Hung, et al.
Published: (2026)
by: Lim, Ho Hung, et al.
Published: (2026)
Can Vision-Language Models Handle Long-Context Code? An Empirical Study on Visual Compression
by: Zhong, Jianping, et al.
Published: (2026)
by: Zhong, Jianping, et al.
Published: (2026)
AuthSig: Safeguarding Scanned Signatures Against Unauthorized Reuse in Paperless Workflows
by: Zhang, RuiQiang, et al.
Published: (2025)
by: Zhang, RuiQiang, et al.
Published: (2025)
Empirical Studies of Large Scale Environment Scanning by Consumer Electronics
by: Wang, Mengyuan, et al.
Published: (2025)
by: Wang, Mengyuan, et al.
Published: (2025)
Autonomous AI-enabled Industrial Sorting Pipeline for Advanced Textile Recycling
by: Spyridis, Yannis, et al.
Published: (2024)
by: Spyridis, Yannis, et al.
Published: (2024)
MultiTaskDeltaNet: Change Detection-based Image Segmentation for Operando ETEM with Application to Carbon Gasification Kinetics
by: Niu, Yushuo, et al.
Published: (2025)
by: Niu, Yushuo, et al.
Published: (2025)
ChronoEarth-492K: A Large Scale and Long Horizon Spatiotemporal Hyperspectral Earth Observation Dataset and Benchmark
by: Si, Haozhe, et al.
Published: (2026)
by: Si, Haozhe, et al.
Published: (2026)
An Empirical Study of World Model Quantization
by: Fu, Zhongqian, et al.
Published: (2026)
by: Fu, Zhongqian, et al.
Published: (2026)
MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
by: Shen, Weixiang, et al.
Published: (2026)
by: Shen, Weixiang, et al.
Published: (2026)
Automating Video Thumbnails Selection and Generation with Multimodal and Multistage Analysis
by: Fantini, Elia
Published: (2024)
by: Fantini, Elia
Published: (2024)
Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
by: Wu, Chen, et al.
Published: (2026)
by: Wu, Chen, et al.
Published: (2026)
CenterMamba-SAM: Center-Prioritized Scanning and Temporal Prototypes for Brain Lesion Segmentation
by: Tian, Yu, et al.
Published: (2025)
by: Tian, Yu, et al.
Published: (2025)
Boosting AI Reliability with an FSM-Driven Streaming Inference Pipeline: An Industrial Case
by: Zhang, Yutian, et al.
Published: (2026)
by: Zhang, Yutian, et al.
Published: (2026)
An Empirical Study of Scaling Law for OCR
by: Rang, Miao, et al.
Published: (2023)
by: Rang, Miao, et al.
Published: (2023)
Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos
by: Pan, Yulin, et al.
Published: (2023)
by: Pan, Yulin, et al.
Published: (2023)
NovaLAD: A Fast, CPU-Optimized Document Extraction Pipeline for Generative AI and Data Intelligence
by: Ulla, Aman
Published: (2026)
by: Ulla, Aman
Published: (2026)
RibPull: Implicit Occupancy Fields and Medial Axis Extraction for CT Ribcage Scans
by: Nikolakakis, Emmanouil, et al.
Published: (2025)
by: Nikolakakis, Emmanouil, et al.
Published: (2025)
Integrated Pipeline for Monocular 3D Reconstruction and Finite Element Simulation in Industrial Applications
by: Zheng, Bowen
Published: (2025)
by: Zheng, Bowen
Published: (2025)
Relation-Rich Visual Document Generator for Visual Information Extraction
by: Jiang, Zi-Han, et al.
Published: (2025)
by: Jiang, Zi-Han, et al.
Published: (2025)
BIMStruct3D: A Fully Automated Hybrid Learning Scan-to-BIM Pipeline with Integrated Topology Refinement
by: Chamseddine, Mahdi, et al.
Published: (2026)
by: Chamseddine, Mahdi, et al.
Published: (2026)
ODGEN: Domain-specific Object Detection Data Generation with Diffusion Models
by: Zhu, Jingyuan, et al.
Published: (2024)
by: Zhu, Jingyuan, et al.
Published: (2024)
A Hierarchical Computer Vision Pipeline for Physiological Data Extraction from Bedside Monitors
by: Chau, Vinh, et al.
Published: (2025)
by: Chau, Vinh, et al.
Published: (2025)
Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities
by: Yang, Fan, et al.
Published: (2025)
by: Yang, Fan, et al.
Published: (2025)
SmartScan: An AI-based Interactive Framework for Automated Region Extraction from Satellite Images
by: Nagendra, Savinay, et al.
Published: (2025)
by: Nagendra, Savinay, et al.
Published: (2025)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
PyFi: Toward Pyramid-like Financial Image Understanding for VLMs via Adversarial Agents
by: Zhang, Yuqun, et al.
Published: (2025)
by: Zhang, Yuqun, et al.
Published: (2025)
DocRevive: A Unified Pipeline for Document Text Restoration
by: Purkayastha, Kunal, et al.
Published: (2026)
by: Purkayastha, Kunal, et al.
Published: (2026)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
by: Islam, Md Mohaiminul, et al.
Published: (2025)
by: Islam, Md Mohaiminul, et al.
Published: (2025)
Similar Items
-
Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models
by: Jin, Yichao, et al.
Published: (2025) -
Financial Table Extraction in Image Documents
by: Watson, William, et al.
Published: (2024) -
Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
by: Jain, Chayan, et al.
Published: (2025) -
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
by: Li, Bozhou, et al.
Published: (2025) -
SynJAC: Synthetic-data-driven Joint-granular Adaptation and Calibration for Domain Specific Scanned Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2024)