Up to 36x Speedup: Mask-based Parallel Inference Paradigm for Key Information Extraction in MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xinzhong, Guo, Ya, Li, Jing, Chen, Huan, Tu, Yi, Hong, Yijie, Liu, Gongshen, Zhu, Huijia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
InsightVision: A Comprehensive, Multi-Level Chinese-based Benchmark for Evaluating Implicit Visual Semantics in Large Vision Language Models
von: Yin, Xiaofei, et al.
Veröffentlicht: (2025)
von: Yin, Xiaofei, et al.
Veröffentlicht: (2025)
Keep the General, Inject the Specific: Structured Dialogue Fine-Tuning for Knowledge Injection without Catastrophic Forgetting
von: Hong, Yijie, et al.
Veröffentlicht: (2025)
von: Hong, Yijie, et al.
Veröffentlicht: (2025)
UNER: A Unified Prediction Head for Named Entity Recognition in Visually-rich Documents
von: Tu, Yi, et al.
Veröffentlicht: (2024)
von: Tu, Yi, et al.
Veröffentlicht: (2024)
Reinforcement Learning from Denoising Feedback
von: He, Qi, et al.
Veröffentlicht: (2026)
von: He, Qi, et al.
Veröffentlicht: (2026)
Fast Monte Carlo Tree Diffusion: 100x Speedup via Parallel Sparse Planning
von: Yoon, Jaesik, et al.
Veröffentlicht: (2025)
von: Yoon, Jaesik, et al.
Veröffentlicht: (2025)
LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
von: Xu, Chenkai, et al.
Veröffentlicht: (2025)
von: Xu, Chenkai, et al.
Veröffentlicht: (2025)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO
von: Guan, Wei, et al.
Veröffentlicht: (2025)
von: Guan, Wei, et al.
Veröffentlicht: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
Multiple Queries with Multiple Keys: A Precise Prompt Matching Paradigm for Prompt-based Continual Learning
von: Tu, Dunwei, et al.
Veröffentlicht: (2025)
von: Tu, Dunwei, et al.
Veröffentlicht: (2025)
Order of Magnitude Speedups for LLM Membership Inference
von: Zhang, Rongting, et al.
Veröffentlicht: (2024)
von: Zhang, Rongting, et al.
Veröffentlicht: (2024)
Intermittent Semi-Working Mask: A New Masking Paradigm for LLMs
von: Hu, HaoYuan, et al.
Veröffentlicht: (2024)
von: Hu, HaoYuan, et al.
Veröffentlicht: (2024)
When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
von: Liu, Jinming, et al.
Veröffentlicht: (2025)
von: Liu, Jinming, et al.
Veröffentlicht: (2025)
GSIFN: A Graph-Structured and Interlaced-Masked Multimodal Transformer-based Fusion Network for Multimodal Sentiment Analysis
von: Jin, Yijie
Veröffentlicht: (2024)
von: Jin, Yijie
Veröffentlicht: (2024)
KIEval: Evaluation Metric for Document Key Information Extraction
von: Khang, Minsoo, et al.
Veröffentlicht: (2025)
von: Khang, Minsoo, et al.
Veröffentlicht: (2025)
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
von: Guo, Yijie, et al.
Veröffentlicht: (2025)
von: Guo, Yijie, et al.
Veröffentlicht: (2025)
Contract-Coding: Towards Repo-Level Generation via Structured Symbolic Paradigm
von: Lin, Yi, et al.
Veröffentlicht: (2026)
von: Lin, Yi, et al.
Veröffentlicht: (2026)
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
von: Chaturvedi, Isha, et al.
Veröffentlicht: (2025)
von: Chaturvedi, Isha, et al.
Veröffentlicht: (2025)
Using Sequential Runtime Distributions for the Parallel Speedup Prediction of SAT Local Search
von: Arbelaez, Alejandro, et al.
Veröffentlicht: (2024)
von: Arbelaez, Alejandro, et al.
Veröffentlicht: (2024)
Speedup Chip Yield Analysis by Improved Quantum Bayesian Inference
von: Li, Zi-Ming, et al.
Veröffentlicht: (2025)
von: Li, Zi-Ming, et al.
Veröffentlicht: (2025)
EchoingPixels: Cross-Modal Adaptive Token Reduction for Efficient Audio-Visual LLMs
von: Gong, Chao, et al.
Veröffentlicht: (2025)
von: Gong, Chao, et al.
Veröffentlicht: (2025)
A Universal Identity Backdoor Attack against Speaker Verification based on Siamese Network
von: Zhao, Haodong, et al.
Veröffentlicht: (2023)
von: Zhao, Haodong, et al.
Veröffentlicht: (2023)
The polarization of strongly lensed point-like radio sources
von: Er, Xinzhong
Veröffentlicht: (2025)
von: Er, Xinzhong
Veröffentlicht: (2025)
Quantum Speedups for Multiproposal MCMC
von: Lin, Chin-Yi, et al.
Veröffentlicht: (2023)
von: Lin, Chin-Yi, et al.
Veröffentlicht: (2023)
Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
A Fixed-Volume Variant of Gibbs-Ensemble Monte Carlo Yields Significant Speedup in Binodal Calculation
von: Qin, Sanbo, et al.
Veröffentlicht: (2025)
von: Qin, Sanbo, et al.
Veröffentlicht: (2025)
Unveiling the Deficiencies of Pre-trained Text-and-Layout Models in Real-world Visually-rich Document Information Extraction
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
Parallel Corpora based Translation Resources Extraction
von: Alberto Simões
Veröffentlicht: (2007)
von: Alberto Simões
Veröffentlicht: (2007)
The xAI Lesson: Data Without Structure, Ambition Without Paradigm
von: Guo, Xiangyu
Veröffentlicht: (2026)
von: Guo, Xiangyu
Veröffentlicht: (2026)
Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference
von: Yin, Hao, et al.
Veröffentlicht: (2025)
von: Yin, Hao, et al.
Veröffentlicht: (2025)
OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
von: Shen, Jiyuan, et al.
Veröffentlicht: (2026)
von: Shen, Jiyuan, et al.
Veröffentlicht: (2026)
Observation of Superoscillation Superlattices
von: Ma, Xin, et al.
Veröffentlicht: (2024)
von: Ma, Xin, et al.
Veröffentlicht: (2024)
Semantic-Aware Visual Information Transmission With Key Information Extraction Over Wireless Networks
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
von: Mao, Yanxu, et al.
Veröffentlicht: (2025)
von: Mao, Yanxu, et al.
Veröffentlicht: (2025)
SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
von: Tu, Jinzhe, et al.
Veröffentlicht: (2026)
von: Tu, Jinzhe, et al.
Veröffentlicht: (2026)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
von: Hu, Yijie, et al.
Veröffentlicht: (2025)
von: Hu, Yijie, et al.
Veröffentlicht: (2025)
Federated Q-Learning: Linear Regret Speedup with Low Communication Cost
von: Zheng, Zhong, et al.
Veröffentlicht: (2023)
von: Zheng, Zhong, et al.
Veröffentlicht: (2023)
General LLMs as Instructors for Domain-Specific LLMs: A Sequential Fusion Method to Integrate Extraction and Editing
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
InsightVision: A Comprehensive, Multi-Level Chinese-based Benchmark for Evaluating Implicit Visual Semantics in Large Vision Language Models
von: Yin, Xiaofei, et al.
Veröffentlicht: (2025) -
Keep the General, Inject the Specific: Structured Dialogue Fine-Tuning for Knowledge Injection without Catastrophic Forgetting
von: Hong, Yijie, et al.
Veröffentlicht: (2025) -
UNER: A Unified Prediction Head for Named Entity Recognition in Visually-rich Documents
von: Tu, Yi, et al.
Veröffentlicht: (2024) -
Reinforcement Learning from Denoising Feedback
von: He, Qi, et al.
Veröffentlicht: (2026) -
Fast Monte Carlo Tree Diffusion: 100x Speedup via Parallel Sparse Planning
von: Yoon, Jaesik, et al.
Veröffentlicht: (2025)